I'm turning pandas dataframes into parquet files. For this I'm using dask, to help to partition the generated files.
my_dask_df = dask.dataframe.from_pandas(my_pandas_df, npartitions=1)
my_dask_df = my_dask_df.repartition(partition_size="100MB")
dask.dataframe.to_parquet(my_dask_df, destination_dir, name_function=lambda
i: f"{dataset_name}-{dataset_version}-part-{i}.parquet")
As you can see in the small snippet, I specifying my_dask_df.repartition(partition_size="100MB"). To my expectation this should create partitions around 100MB.
Now, I've had now a few datasets, which are below 100mb or maybe just over that. The partitioned files turn out to be only 5mb and I'm ending up with n 5mb large/small files.
Why is this?