Vertex AI endpoint doesn't scale up / down

Viewed 89

I've been deploying some custom trained models to Vertex AI, but lately, the feature to auto-scale has not been working properly on the later endpoints. Basically, despite the traffic, the endpoint doesn't auto-scale.

I have an older endpoint which works as intended, so I deployed the same model to a different endpoint with the same configuration (same machine specs, same GPU, min 1 machine, max 3 machines, 60% threshold to auto-scale), created it's own task queue and then proceeded to send the same requests to both endpoints at the same time.

The older endpoint worked as intended, scaling up and down depending on the incoming traffic. The newer one, on the other hand, stayed stuck at one machine the entire time.

I can force it to scale up if I lower the threshold to 15-20%, and it does scale up as the requests come in. However, it does not scale down once it has finished processing the requests and it stays with all the machines on even when there has not been any traffic for hours.

So, what may be preventing the newer endpoint to scale up as the traffic increases, given that the older endpoint does scale up and down as intended with the same traffic? And perhaps more importantly, what prevents it to scale down if I force it to scale up?

1 Answers

I can't explain the difference in the two endpoints.

Autoscaling with Vertex AI endpoints is (by default) based on the cpu utilization across all cores of the machine type you've specified. The default threshold of 60% represents 60% on all cores. For a 4 core machine, that means you need 240% utilization to trigger autoscaling.

I'm currently debugging a similar issue and I suspect torchserve is being single threaded and only using one core of a 4 core machine. Torchserve has some options that may help it not be single threaded, but I haven't finished testing these. I'm hopeful that default_workers_per_model will help in my case. Full docs here: https://pytorch.org/serve/configuration.html

You should also check your quota usage for whatever GPU you're requesting to be sure that's not the issue.

I haven't solved this issue yet, but hopefully some of the info above helps.

Related