PySpark Incorrectly Inferring Parquet Date

Viewed 40

I have an AWS Glue (2.0, Spark 2.4, Python3) job that is bringing in s3 stored Parquet files via the create_dynamic_frame.from_options() function. The source data for the parquet files is a MySQL table that houses infrequently populated data, UDF's, where out of 600 rows only two might have a value for the column. Spark is inferring these mostly NULL columns to be integers, where the true datatype is date (stored in Parquet as the number of days from unix date).

What I am wondering is, is there a way to force spark to look at all the data in the column, something like the ratio sampling that you can do for an RDD with either a dynamic frame or data frame? Is there an alternate but good way to go about this? This is a dynamic situation where I need to use the same script for migrating 100+ databases and the UDF's will differ by database, so it is not feasible to hardcode the data types.

Here is the code for creating the dynamic frame:

dyf_full = glueContext.create_dynamic_frame.from_options(connection_type='s3',
                 connection_options={'path': path, 'recurse': True, 'exclusions': exclusion_string},
                 format='parquet',
                 transformation_ctx='df_full')

and here is the inferred schema, where date field 2 is infrequently populated and date_field is more than 95% populated:

root
|-- CASE_ID: decimal
|-- DATE_FIELD: date
|-- DATE_FIELD_2: int

Please let me know if there is something else I can provide that would be beneficial. I am pretty stumped on this one.

0 Answers
Related