I'm working on Data Solution using GCP Dataproc and Airflow.
While creating dataproc Cluster via Airflow its easy using DataprocCreateClusterOperator, but the challenges are when dataproc cluster creation is getting failed due to valid reason like - IP exhaustion in cluster zone etc etc..
We tried 2 approach for this -
Using Custom Operator inherited from DataprocCreateClusterOperator : Using this we are able to achieve cluster creation but we are unable to check retry_number of the task (due to unavailability of provider_context=True parameter in Custom Operator which helps for task_instance object)
Calling DataprocCreateClusterOperator using Python Operator (which accepts provider_context=True and can check task retry_number) : In this case python operator executes successfully but doesn't create dataproc cluster.
There are other ways like branching in Airflow, but this will have following issue -
Code redundancy
Having multiple dataproc config individually for first and subsequent retries. With PythonOperator and CustomOperator, Config can be updated programatically.
Is there any other way to create dataproc cluster with different config if creation failed in first task run.