How to parse json in parquet file using apache/arrow

Viewed 436

I am using Apache arrow for Go to read parquet files. The schema of my parquet file is:

time_stamp: int64
file_name:  byte_array
offset:     int32
meta_data:  byte_array

That information is printed by fmt.Println(rdr.MetaData().Schema). Although it says column metadata is a byte array, it's actually a json string like following:

{
    "dataType": "left", 
    "features": [
        {
            "feature_name": "dHash", 
            "feature_val": "0000011000000111000001110010011100011111000101110000010100000101"
        }
    ], 
    "pipelineVersion": "0.0"
}

So how can I parse this information into a struct? I have found following methods to read a parquet file, but there seems to be no parameter for schema:

mem := memory.NewCheckedAllocator(memory.DefaultAllocator)
filename := "parquet file path"

rdr, _ := file.OpenParquetFile(filename, false, file.WithReadProps(parquet.NewReaderProperties(mem)))
arrowRdr, _ := pqarrow.NewFileReader(rdr, pqarrow.ArrowReadProperties{}, mem)
tbl, _ := arrowRdr.ReadTable(context.Background())
defer tbl.Release()

chunk0 := tbl.Column(0).Data().Chunk(0)
fmt.Println(chunk0)

And there is no example in official doc at all. Thank you in advance.

1 Answers

That information is printed by fmt.Println(rdr.MetaData().Schema). Although it says column metadata is a byte array, it's actually a json string like following:

If the json payloads has been stored as a json byte array/string in parquet, then you'll have to parse it and convert it to a struct manually. There are some helper functions that can handle json data, but it doesn't look like they are exposed in go.

If you want parquet to handle it automatically as a struct, you'd have to store it as a struct when writing the file.

Related