I'm trying to use Batch for large-scale parallelised job execution, with Docker containers. I would like to process thousands of tasks simultaneously.
I have everything up and running. My compute environment is configured with a max vCPUs of 2048. Each task is configured to use a single vCPU, and 2GB of RAM. I am using an array job with 1,000 array elements (for now).
Problem is: when I create a new job, concurrency seems to be extremely limited. When I look at the cluster in ECS, "pending tasks" seems to constantly hover around 50 (it might not ever go above 50), and "running tasks" doesn't go far above 30. Even though each individual task only takes ~10 seconds to complete, the entire batch takes ~20 minutes.
This isn't what I expected. With the above settings, I thought Batch would process all 1,000 tasks at the same time.
I originally thought the problem might have been caused by my use of a public subnet (all Fargate containers had public IPs). I changed to use a private subnet (with NAT gateway), but it didn't help.
Does anyone know what I'm doing wrong?
Thanks!