I'm working with a set of data which has almost 60 columns (Text/Address/Numbers). After processing the data using Pandas, I have to export it to xlsx format.
This is how I'm generating the output:
with pd.ExcelWriter("output.xlsx", engine='xlsxwriter') as writer:
df.to_excel(writer, sheet_name="sheet", index=False)
I have also tried this method:
df.to_excel('output.xlsx', index=False, engine='xlsxwriter')
What I have noticed is, generating xlsx is remarkably slower than a format like csv. And as the number of records grows, the time of generating the xlsx file significantly increases.
Is this the normal and expected behavior of .to_excel or there is something wrong here? is there any way to debug and solve this problem?
I have to be able to generate xlsx files for ~300K to ~600K records in matter of seconds, but as you can see it takes me around 6 minutes to generate an excel file for about 500K records.
The hardware that I'm using to generate these files has 16 Core of CPU, and 64 GB of memory.
