ParquetDecodingException using pyspark

Viewed 296

I saved a parquet file, then loaded it and tried to join it with another dataframe. I used the regular read/write parquet files. Then I got the following error:

Caused by: org.apache.spark.SparkException: Job aborted due to stage 
failure: Task 925 in stage 37.0 failed 1 times, most recent failure: Lost 
task 925.0 in stage 37.0 (TID 6376, localhost, executor driver): 
org.apache.parquet.io.ParquetDecodingException: Can not read value at 
113652 in block 0 in file 
file:/my_parquet_path/part-00031-b9b3442d-8459-4591-956c-9ef2299095cd-c000.snappy.parquet

Caused by: org.apache.parquet.io.ParquetDecodingException: could not read 
page Page [bytes.size=1048635, valueCount=30917, uncompressedSize=1048635] 
in col [my_field_name, list, element] optional binary element (UTF8)

java.io.IOException: FAILED_TO_UNCOMPRESS(5)

Any idea why does it happen? How the filter the problematic rows?

0 Answers
Related