I have a short script that's a reduction of my application written using "pandas>=0.25.3" which has been upgraded to "pandas==1.1.5" the latest version of our code. In this version of Pandas, the default engine does not parse xlsx so we've added the engine="openpyxl". However, there's a new issue. The read_excel no longer seems to respect the names argument and has strange behavior.
import pandas
filename = "... .xlsx"
names = ["foo", "bar", "baz"]
data_frame = pandas.read_excel(
filename,
header=None,
names=names,
engine="openpyxl",
skiprows=3,
sheet_name=0,
)
print(data_frame.iloc[3])
Running the script with the new Pandas I get this output:
foo NaN
bar NaN
baz NaN
Name: (FxK2,SMin, 2066.125), dtype: float64
But previously in pandas-0.25.3 parsing with the xlrd engine by default I got what I expected which was this:
foo FxK2
bar SMin
baz 2066.125
The names field gave names to the columns and then I could reference data_frame.iloc[0].baz and get 2066.125. Now for some reason, then entire thing ends up in the optional name field for the data frame.
How can I get the behavior I was used to and is this potentially a bug or just new interface I'm not used to? The pandas-1.1.5 certainly seem to reference the names argument in the same way as I was used to using it.