Distinguish between null and missing field when reading json into pyspark dataframe

Viewed 20

I'm reading in a set of jsonl files with a schema to a dataframe:

df = spark.read.option("mode", "FAILFAST").json(
    f"{base_input_uri}/*.jsonl*", schema)

and I want to distinguish between fields that are present in the jsonl with an explicit value of null (which I want to maintain as null in the dataframe) and fields that are missing entirely from a record. At the moment, both become None in the dataframe, and in this case, we need to know which are coming in explicitly as nulls.

I've tried setting spark.sql.jsonGenerator.ignoreNullFields to false in the spark configuration to no avail.

I also tried:

df = spark.read.option("mode", "FAILFAST").json(
    f"{base_input_uri}/*.jsonl*", schema, ignoreNullFields=False)

but that resulted in

TypeError: DataFrameReader.json() got an unexpected keyword argument 'ignoreNullFields'
0 Answers
Related