How to write NullType field in parquet format in Pyspark?

Viewed 56

I am reading a json file and inferring the schema through spark. One of the field is arr: [] , so when I am trying to write this json object into parquet format, it is throwing an error: An error was encountered: 'Parquet data source does not support array<null> data type.;'. I have reproduced the code resulting in the error(in the example code, I have added the schema, but in the actual code, I am using glue dynamic frame):

data = {"key":"val","arr":[]}
glueContext = GlueContext(SparkContext.getOrCreate())

schema = StructType([
    StructField("key", StringType()),
    StructField("arr", ArrayType(NullType()))

])
df = spark.createDataFrame([data], schema)
ddf= DynamicFrame.fromDF(df, glueContext, 'glue_df')
ddf.toDF().write.parquet("/home/file1")

For now, since there is no value in the arr field, the element inside it is inferred as NullType() but it won't be the case everytime since later on it could be StringType() etc.

I want to update the code so that either the field gets dropped or the NullType gets typecasted into StringType. I tried attribute_master = DropNullFields.apply(frame=ddf) but it didn't worked. What could be the workaround?

0 Answers
Related