My team and I are on Airflow v2.1.0 using the Celery executor with Redis. Recently we’ve noticed some jobs are occasionally running until we kick them (many hours, sometimes days—basically until someone notices). There doesn’t seem to be a particular pattern that we’ve noticed yet.
We also use DataDog and the statsd provider to collect and monitor metrics produces by Airflow. Ideally we could setup a DataDog monitor for this but there doesn’t appear to be an obvious metric for this situation.
How can we detect and alarm on stuck jobs like this?