I have a large parquet file with categorical/dictionary values in it's schema (dictionary<values=string, indices=int32, ordered=0>) and I'm trying to convert the parquet file into pyarrow IPC format using RecordBatchFileWriter, but I'm getting this error:
pyarrow.lib.ArrowInvalid: Dictionary replacement detected when writing IPC file format. Arrow IPC files only support a single non-delta dictionary for a given field across all batches.
Since I'm reading batch by batch Arrow IPC can't figure out the dictionary values (which makes perfect sense). But how can I work around this? I can't read the dataset into ram (too large) but I'm happy to read the dictionary columns separately one by one and pre-compute the categorical - but I can't figure out how to do that.
My best guess would be to somehow pre-compute the categorical values based on the parquet file, but how?
>>> parquet_file = pq.ParquetFile('data.parq')
>>> local = fs.LocalFileSystem()
>>> with local.open_output_stream("data.arrow") as file:
>>> with pa.RecordBatchFileWriter(file, parquet_file.schema_arrow) as writer:
... for record_batch in parquet_file.iter_batches():
... writer.write_batch(record_batch)
...
Traceback (most recent call last):
File "arrow-playground.py", line 47, in <module>
writer.write_batch(record_batch)
File "pyarrow/ipc.pxi", line 483, in pyarrow.lib._CRecordBatchWriter.write_batch
File "pyarrow/error.pxi", line 100, in pyarrow.lib.check_status
pyarrow.lib.ArrowInvalid: Dictionary replacement detected when writing IPC file format. Arrow IPC files only support a single non-delta dictionary for a given field across all batches.
Regards, Niklas
- I've read everything I could find, such as https://www.mail-archive.com/user@arrow.apache.org/msg01316.html and https://arrow.apache.org/docs/format/Columnar.html#dictionary-encoded-layout and https://lists.apache.org/thread/ygzc8gj5nv7453wwg4k0g0d9pl51qs2l