PyArrow get Metadata from file in S3

Viewed 163

I want to get Parquet file statistics (such as Min/Max) from file in S3 using PyArrow. I am able to fetch it using

pq.ParquetDataset(s3_path, filesystem=s3)

and get the statistics if I download and read it using:

ParquetFile(full_path).metadata.row_group(0).column(col_idx).statistics

hope there is a way to achieve it without download the whole file.

Thanks

1 Answers

I came to this post looking for a similar answer few days ago. In the end I found a simple solution that works for me.

import pyarrow.parquet as pq
from pyarrow import fs

s3_files = fs.S3FileSystem(access_key) # whatever need to connect to s3

# fetch the dataset
dataset = pq.ParquetDataset(s3_path, filesystem=s3_files)

metadata = {}
for fragment in dataset.fragments:
    meta = fragment.metadata
    metadata[fragment.path] = meta
    print(meta)

The metadata is store in the dictionary where the keys are the path to the fragment in the s3 and the values is the metadata of that particular fragment.

to acces the statistics just use

meta.row_group(0).column(col_idx).statistics

something like this will be printed for every fragment

<pyarrow._parquet.FileMetaData object at 0x7fb5a045b5e0>
  created_by: parquet-cpp-arrow version 8.0.0
  num_columns: 6
  num_rows: 10
  num_row_groups: 1
  format_version: 1.0
  serialized_size: 3673
Related