I saved a parquet file, then loaded it and tried to join it with another dataframe. I used the regular read/write parquet files. Then I got the following error:
Caused by: org.apache.spark.SparkException: Job aborted due to stage
failure: Task 925 in stage 37.0 failed 1 times, most recent failure: Lost
task 925.0 in stage 37.0 (TID 6376, localhost, executor driver):
org.apache.parquet.io.ParquetDecodingException: Can not read value at
113652 in block 0 in file
file:/my_parquet_path/part-00031-b9b3442d-8459-4591-956c-9ef2299095cd-c000.snappy.parquet
Caused by: org.apache.parquet.io.ParquetDecodingException: could not read
page Page [bytes.size=1048635, valueCount=30917, uncompressedSize=1048635]
in col [my_field_name, list, element] optional binary element (UTF8)
java.io.IOException: FAILED_TO_UNCOMPRESS(5)
Any idea why does it happen? How the filter the problematic rows?