How can Airflow be used to actively control the DAGs given the RAM and CPU constraint

Viewed 1088

I got much familiar with airflow's programming features by trying out lot of samples.. What's keeping me from digging further is how it can perform its job without overloading the CPU or RAM, is there a way to control the Load so that it won't run out of resources

I know one way to reduce the load when scheduler does its job of 'scheduling and picking out the files more often' by changing the values for the following fields min_file_process_interval and scheduler_heartbeat_sec to a minute interval or so. Though it reduces the constant CPU hike, but when the interval passes(i.e., after a minute), It suddenly goes back to sucking ~95% of the CPU as it does during the startup.. How do you reduce that as well so it never consumes more than 70% of the CPU at least ?

EDITED:

Also, when the scheduler_heartbeat interval passes, i see all my python scripts execute once again.. is this the way it works? i thought it will pick up the new DAG if any after the interval otherwise wouldn't do anything.

1 Answers

There are a few techniques you can use to control the number of processes running on airflow.

  1. Use Pools. You can assign pools in the dag setup, or you can just add it to your operator so that the random dag creator has that detail hidden from them.
  2. For backfilling tasks I think there is a parameter concurrency and max_active_runs which are defined when you initialize a DAG
  3. Distribute your compute if you are using CeleryExecutor. You can have the CeleryExecutor execute on remote machines.[Didn't try this myself, but I have heard success stories with this.]

These are the ones I have used. You'll have to be smart about the allocation to control CPU spikes and memory issues.

Related