Reading json file causing corrupt_record in pyspark

Viewed 62

I have been trying to read an open source json file available here in the zip format(https://healthy.kaiserpermanente.org/pricing/innetwork/co/2022-08-01_KPIC_CO-COMMERCIAL_in-network-rates.zip). The zipped file size is 50MB and unzipped file size is about 700MB. When I try to read this file, I encountered ['_corrupt_record'] issue.

df = (spark.read.format("json")
    .option("multiline", "true")
    .load(file_path)
)

Based on this blog https://medium.com/@sasidharan-r/how-to-handle-corrupt-or-bad-record-in-apache-spark-custom-logic-pyspark-aws-430ddec9bb41, I also tried the following method:

df = (spark.read.format("json")
    .schema(schema_in_network)
    .option("multiline", "true").option("mode", "PERMISSIVE")
    .option("columnNameOfCorruptRecord", "_corrupt_record")
    .load(file_path)
)

In addition, based on the Databrick website, I attempted the following:

df = (spark.read.format("json")
    .option("multiline", "true").option("mode", "PERMISSIVE")
    .option("rescuedDataColumn", "_rescued_data")
    .load(file_path)
)

None of the above methods rectified the issue. I am not sure if the problem exists in the json file or my approach. In either case, I want to read the content ignoring corrupted value, if any. Thank you in advance for your help.

PS: I am new to Spark.

1 Answers

Try this:

df = (spark.read.format("json")
        .option("multiline", "true").option("mode", "DROPMALFORMED")
        .option("rescuedDataColumn", "_rescued_data").load(file_path)
)

Mode DROPMALFORMED ignores the whole corrupted records. Also try to read doc for JSON files

Related