Read Parquet File with illegal characters (Apache-Avro)

Viewed 451

I have some Parquet files written in Python using PyArrow. Now I want to read them using a Java program. I tried the following, using Apache Avro:

import java.io.IOException;

import org.apache.avro.generic.GenericRecord;
import org.apache.avro.Schema;
import org.apache.avro.SchemaBuilder;
import org.apache.hadoop.conf.Configuration;
import org.apache.hadoop.fs.Path;
import org.apache.parquet.avro.AvroParquetReader;
import org.apache.parquet.avro.AvroReadSupport;
import org.apache.parquet.hadoop.ParquetReader;

public class Main {
    
    private static Path path = new Path("D:\\pathToFile\\review.parquet");

    public static void main(String[] args) throws IllegalArgumentException {
        try {
            Configuration conf = new Configuration();

            Schema schema = SchemaBuilder.record("lineitem")
                    .fields()
                        .name("reviewID")
                        .aliases("review_id$str")
                        .type().stringType()
                        .noDefault()
                    .endRecord();
                conf.set(AvroReadSupport.AVRO_REQUESTED_PROJECTION, schema.toString());           
                
            ParquetReader<GenericRecord> reader = AvroParquetReader.<GenericRecord>builder(path)
                .withConf(conf)
                .build();
            
            GenericRecord r;
            while (null != (r = reader.read())) {
                
                r.getSchema().getField("reviewID").addAlias("review_id$str");
           
                Object review_id = r.get("review_id$str");
                String review_id_str = review_id != null ? ("'" + review_id.toString() + "'") : "-";
                System.out.println("review_id: " + review_id_str);
        
            }
        } catch (IOException e) {
            System.out.println("Error reading parquet file.");
            e.printStackTrace();
        }
    }
}

My Parquet File contains columns whose name contain the symbols [, ], ., \ and $. (In this case, the Parquet file contains a column review_id$str, whose values I want to read). However, these characters are invalid in Avro (see: https://avro.apache.org/docs/current/spec.html#names). Therefore, I tried to use Aliases (see: http://avro.apache.org/docs/current/spec.html#Aliases). Even though now I don't get any "Invalid Character Errors", I am still unable to get the values, i.e. nothing is getting printed even though the column contains values.

It only prints:

review_id: -
review_id: -
review_id: -
review_id: -
...

And expected would be:

review_id: Q1sbwvVQXV2734tPgoKj4Q
review_id: GJXCdrto3ASJOqKeVWPi6Q
review_id: 2TzJjDVDEuAW6MR5Vuc1ug
review_id: yi0R0Ugj_xUx_Nek0-_Qig
...

Am I using the Aliases wrong? Is it even possible to use aliases in this situation? If so, please explain me how I can fix it. Thank you.

Update 2021: In the end, I decided not to use Java for this task. I stuck to my solution in Python using PyArrow which works perfectly fine.

0 Answers
Related