Airflow job fails after detecting successful Dataflow job as a zombie.
I run an hourly Dataflow job that's triggered by an external Airflow instance using the python DataflowTemplateOperator. A couple of times a week, Dataflow becomes completely unresponsive to status pings. When I've caught the error in real-time and tried looking at the status of the Dataflow job in the GCP UI, the page won't load despite my having a network connection and being able to look at other pages on the GCP site. After a few minutes, everything returns to normal working order. This seems to happen towards the end of a job's run or when workers are shutting down. The Dataflow jobs don't fail, and don't report any errors. Airflow thinks they've failed because, when Dataflow becomes unresponsive, Airflow assumes the jobs are zombies. I needed a fast solution and just increased my number of retries, but I would like to understand the problem and find a better solution.