I'm pulling a dataset from a Cassandra table using Spark Java. This table has a UDT column PhoneNumbers, and sometimes this column is null.
| account_id | phone_numbers |
|---|---|
| 1 | {mobile_phone: {phone_number: 1234567890, date_added: 2022-01-01}, home_phone: {phone_number: 1234567891, date_added: 2022-01-02}} |
| 2 | null |
CassandraJavaRDD<PhoneNumbers> cassandraJavaRDD = CassandraJavaUtil.javaFunctions(javaSparkContext)
.cassandraTable(
keyspace,
tableName,
mapRowTo(PhoneNumbers.class))
.select(columns);
Dataset<PhoneNumbers> tableAsDataset = sparkSession.createDataset(cassandraJavaRDD.rdd(), Encoders.bean(PhoneNumbers.class));
tableAsDataset.show();
In Spark 2, this works fine and can handle null values in a UDT column (Spark 2.4.8 and Spark Cassandra Connector 2.5.1).
However, after upgrading to Spark 3 (Spark 3.2.2 and Spark Cassandra Connector 3.2.0), I get the following error:
2022-08-03 14:28:18 ERROR Throwable:84 - Caused by: com.datastax.spark.connector.types.TypeConversionException: Cannot convert object null to example.PhoneNumbers
How can I read these null UDT entries? Is there some property that needs to be set in Spark 3 to allow null, or a way to fill null values with an empty object, or ignore rows with a null value altogether?