I am trying to read the hive table & write spark DF to JSON file. The file is ignoring null fields.
val dataframe = spark.sql("select * from parquet_hive_table")
dataframe.write.mode(SaveMode.Overwrite).format("json")
.save("<output_path>")
I tried setting ignoreNullFields to false but I understand it works only for Spark 3+ version.
I am using spark 2.4.7 as GCP dataproc is having a dependency issue for the spark 3 version.
I cannot use dataframe.na.fill("") because it will only take care for string datatypes. I have other datatype fields having nulls in it.
I referred to this solution: Retain keys with null values while writing JSON in spark
but I don't have a schema with me. Need to know a solution that uses Dataframe only.
Can anyone guide me to solve this error?