Batch by batch convert parquet to arrow with categorical values: Arrow IPC files only support a single non-delta dictionary for a given field across

Viewed 110

I have a large parquet file with categorical/dictionary values in it's schema (dictionary<values=string, indices=int32, ordered=0>) and I'm trying to convert the parquet file into pyarrow IPC format using RecordBatchFileWriter, but I'm getting this error:

pyarrow.lib.ArrowInvalid: Dictionary replacement detected when writing IPC file format. Arrow IPC files only support a single non-delta dictionary for a given field across all batches.

Since I'm reading batch by batch Arrow IPC can't figure out the dictionary values (which makes perfect sense). But how can I work around this? I can't read the dataset into ram (too large) but I'm happy to read the dictionary columns separately one by one and pre-compute the categorical - but I can't figure out how to do that.

My best guess would be to somehow pre-compute the categorical values based on the parquet file, but how?

>>> parquet_file = pq.ParquetFile('data.parq')

>>> local = fs.LocalFileSystem()
>>> with local.open_output_stream("data.arrow") as file:
>>> with pa.RecordBatchFileWriter(file, parquet_file.schema_arrow) as writer:
...     for record_batch in parquet_file.iter_batches():
...         writer.write_batch(record_batch)
...

Traceback (most recent call last):
  File "arrow-playground.py", line 47, in <module>
    writer.write_batch(record_batch)
  File "pyarrow/ipc.pxi", line 483, in pyarrow.lib._CRecordBatchWriter.write_batch
  File "pyarrow/error.pxi", line 100, in pyarrow.lib.check_status
pyarrow.lib.ArrowInvalid: Dictionary replacement detected when writing IPC file format. Arrow IPC files only support a single non-delta dictionary for a given field across all batches.

Regards, Niklas

0 Answers
Related