I am using pyspark and I want to read/write parquet data with uuids in it, which I'd prefer to save as the parquet UUID LogicalType (which is a 16-bytes fixed array).
See https://github.com/apache/parquet-format/blob/master/LogicalTypes.md
How can I do so in pyspark?
I was wondering whether I should try to extend class pyspark.sql.types.DataType and translate between bytes and uuid.UUID, however it is not clear to me how spark would then write that as the UUID logical type in parquet.