Pandas sparse data export to csv - speed explanation

Viewed 206

I am trying to export subsets of a Pandas dataframe that is made up of columns of type pd.SparseDtype("float32", np.nan) to csv. I have noticed some differences in speed between writing directly to csv vs using the sparse.to_dense() and then writing to csv. Can anyone explain to me whats going on here?

an example:

Create some data.

df = pd.DataFrame(columns = ['metric_' + str(x) for x in range(0,10000)], index = [x for x in range(0,10000)])
df = df.astype(pd.SparseDtype("float32", np.nan))

with out going to dense first

%%timeit
df2 = df.iloc[0:200,:]
df2.to_csv('small_csv2.csv')
1min 8s ± 4.91 s per loop (mean ± std. dev. of 7 runs, 1 loop each)

going to dense before writing to csv

%%timeit
df3 = df.iloc[0:200,:].sparse.to_dense()
df3.to_csv('small_csv3.csv')
4.04 s ± 127 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)

Many thanks

0 Answers
Related