Parquet `write_table` introduces keys of data type to the data when writing to output file

Viewed 335

I have an issue when writing data to parquet files. I tried with different pyarrow versions (both 2.0 and 3.0) but the results look the same.

Examples of how my data looks like:

test_data = {
    'dogs': [
        {'dog': 'frankie'},
        {'dog': 'ricky'}
    ]
}

other_test_data = {
    'dogs': [
        {'dog': 'rory'},
        {'dog': 'marko'}
    ]
}

Then, I reformat them to look like this:

dog_data = {
    'dogs': [
        [{
            'dog': 'frankie'
        }, {
            'dog': 'ricky'
        }],
        [{
            'dog': 'rory'
        }, {
            'dog': 'marko'
        }]
    ]
}

I define the schema:

dog_fields = [
    pa.field('dog', pa.string(), nullable=True)
]

dog_schema = pa.schema([
        ('dogs', pa.list_(pa.struct(dog_fields)))
    ])

I convert them to pyarrow.Table using: pq_table = pa.Table.from_pydict(mapping=dog_data, schema=dog_schema)

Finally, I write to a file: pq.write_table(pq_table, 'dog_data.parquet')

What I see in the file is this, additional keys called list and item:

{
    "dogs": {
        "list": [{
            "item": {
                "dog": "frankie"
            }
        }, {
            "item": {
                "dog": "ricky"
            }
        }]
    }
}

Can anyone explain please why the types of the data fields are added as keys to the data?

Is there a way around it?


EDIT

This is how I get the data with the list and item fields. I install the package with brew install parquet-tools, and then run:

parquet-tools cat --json dog_data.parquet

The reason I chose to load the file like this is that I wanted to inspect what the contents are. The need came from the broken schema I was seeing when loading the data from parquet files to BigQuery. BigQuery doesn't understand the structure of the data and interprets the schema as following:

enter image description here

Annoying .list and .item things are added there.

3 Answers

@christinabo the issue here is that the canonical way of representing a lists in Parquet is to use three groups for a single list. The outer group has the name the user specifies. The two inner groups are meant to be called "list" and "element". pyarrow uses "item" instead of "element" by default. So BQ is ignoring the fact that the nested groups are meant to be one logical type.

The behavior is somewhat controllable with Enable List Inference Parameter

(full disclosure I work for BQ and on Arrow)

How do your get the the dictionary with the additional list/item?

As far as I can tell, converting your data to arrow.Table, saving it to parquet and reloading it yields the same results:

table = pa.Table.from_pydict(mapping=dog_data, schema=dog_schema)
pq.write_table(table, 'dog_data.parquet')
loaded_table = pq.read_table('dog_data.parquet')

print(loaded_table.to_pydict() == dog_data)
>>> True
print (loaded_table.to_pydict())
>>> {'dogs': [[{'dog': 'frankie'}, {'dog': 'ricky'}], [{'dog': 'rory'}, {'dog': 'marko'}]]}
Related