I would like to write to a parquet file a set of data from the following class:
public class ClassA {
Identifier identifier;
Map<String, Object> info;
}
Keep in mind that "Identifier" is in fact an interface. We have 2 classes implementing the Identifier interface, namely:
public class IdentifierA {
Integer id;
}
public class IdentifierB {
List<Long> ids;
}
Note: ClassA, IdentifierA and IdentifierB all have getters, setters and constructors, I just left them out for simplicity purposes.
I can instantiate sets of data, regardless of what Identifier class I use, and I'm also able to write the data to a text file. The problem appears when I try to write the data to a parquet file as I get the following error:
Exception in thread "main" org.apache.spark.sql.AnalysisException:
Datasource does not support writing empty or nested empty schemas.
Please make sure the data schema has at least one or more column(s).
I also verified the columns and the schema of my data. Here they are:
Schema
StructType(StructField(identifier,StructType(),true),StructField(info,MapType(StringType,StructType(),true),true))
Columns [identifier, info]
I'm guessing the error appears either because I'm using an interface within ClassA or because of the Object type used as a value for the map parameter from ClassA.
Do you guys have any idea of how I could resolve this issue?