Spark 2.4.7 ignoring null fields while writing to JSON

Viewed 97

I am trying to read the hive table & write spark DF to JSON file. The file is ignoring null fields.

val dataframe = spark.sql("select * from parquet_hive_table")

  dataframe.write.mode(SaveMode.Overwrite).format("json")
  .save("<output_path>")

I tried setting ignoreNullFields to false but I understand it works only for Spark 3+ version.

I am using spark 2.4.7 as GCP dataproc is having a dependency issue for the spark 3 version.

I cannot use dataframe.na.fill("") because it will only take care for string datatypes. I have other datatype fields having nulls in it.

I referred to this solution: Retain keys with null values while writing JSON in spark

but I don't have a schema with me. Need to know a solution that uses Dataframe only.

Can anyone guide me to solve this error?

0 Answers
Related