Adding Labels to Big Query Table from Pyspark job on Dataproc using Spark BQ Connector

Viewed 68

I am trying to use py-spark on google dataproc cluster to run a spark job and writing results to a Big Query table.

Spark Bigquery Connector Documentation - https://github.com/GoogleCloudDataproc/spark-bigquery-connector

The requirement is during the creation of the table, there are certain labels that should be present on the big query table.

The spark bq connector does not provide any provision to add labels for write operation

df.write.format("bigquery") \
    .mode("overwrite") \
    .option("temporaryGcsBucket", "tempdataprocbqpath") \
    .option("createDisposition", "CREATE_IF_NEEDED") \
    .save("abc.tg_dataset_1.test_table_with_labels")

The above command creates bigquery load job in background that loads the table with the data. Having checked further, the big query load job syntax itself does not support addition of labels in contrast to big query - query job.

Is there any plan to support the below

  1. Support for labels in big query load job
  2. Support for labels in write operation of spark bq connector.

Since there is no provision to add labels during the load/write operation, the current workaround used is to have the table created with schema/labels before the pyspark job

0 Answers
Related