My PySpark computes a DataFrame that I want to insert into a BigQuery table (from a dataproc cluster). On the BigQuery side, the partition field is REQUIRED. On the DataFrame side, the partition field inferred is not REQUIRED, that is why I make a schema defining this field as REQUIRED :
StructField("date_part",DateType(),False)
So, I create a new DF with the new schema and when I show this DF, I see as expected :
date_part: date (nullable = false)
But my PySpark ended like that :
Caused by: com.google.cloud.spark.bigquery.repackaged.com.google.cloud.bigquery.BigQueryException: Provided Schema does not match Table xyz$20211115. Field date_part has changed mode from REQUIRED to NULLABLE
Is there something I missed ?
Update :
- I am using Spark 3.0 image
- And spark-bigquery-latest_2.12.jar connector