I'm running kubernetes on AKS with both staging and production clusters having only 2 vCPU/ 8GB per node. I'm constantly in a situation where new pods can't be scheduled due to insufficient CPU or memory.
Allocated resources:
(Total limits may be over 100 percent, i.e., overcommitted.)
Resource Requests Limits
-------- -------- ------
cpu 1869m (98%) 25800m (1357%)
memory 3561Mi (66%) 18484Mi (344%)
ephemeral-storage 0 (0%) 0 (0%)
hugepages-1Gi 0 (0%) 0 (0%)
hugepages-2Mi 0 (0%) 0 (0%)
I understand that VPA can optimize the allocation per deployment. There are 4 deployments for kube-prometheus-stack (helm chart), 2 for open-telemetry-collector (one for the "operator" and one for the collector), and there are even more deployments for the rest of the monitoring stack (zipkin, elasticsearch, etc).
It seems unweildy to create individual VPA objects/manifests for every single deployment. That's a lot of yaml. Is this really the recommended practice? Is there a better way?
For a staging environment, some of the deployments have only 1 replica. VPA only updates resource requests for deployments that have at least 2 replicas. If I increase to 2, ironically that will consume (obviously) more requests, possibly more than the cluster can handle. VPA doesn't seem to be useful in this scenario.