AWS Batch unmanaged compute env cant schedule jobs on a autoscaler enabled ECS cluster

Viewed 25

We have the following setup:

  1. AWS batch job with an unmanaged compute environment (we need this because we use GPU and custom disk sizes, and managed environment with Launch templates are not working. But that should be a different bug report).
  2. The automatically created ECS cluster is enabled with a EC2 Autoscaling group Capacity provider.
  3. The Autoscaling group has a target tracking dynamic scaling policy which tracks the metric AWS/ECS/ManagedScaling > CapacityProviderReservation.

The min size of the ASG is 0, so during idle periods there are no instances running on the ASG. When we schedule a Batch job at this point, the job becomes stuck in RUNNABLE state until we somehow manually trigger a scale up of the ASG. I think the reason is because Batch sees the ECS cluster and thinks well there are no resources to schedule the job, so I will keep the job in RUNNABLE state. But only if the Batch starts a Task on ECS, then ECS CapacityProviderReservation metric can change and it will trigger an scale up. One can see how it leads to a dead end. Now, we have verified all the related permissions and present and there are no capacity issues. If we scale up the ASG the job gets scheduled as expected. The scale down after the job finishes also works fine and the ASG/ECS cluster returns to 0 instances after.

Have anyone else face a similar issue and found a solution?

0 Answers
Related