I'm running an application in the following mode
- Trained a ML model in tensorflow
- Created an API using Fast API and wrapped it around the ML model for inferencing.
- Created a Dockerfile and containerized the whole application
- Pushed the image to ECR
- Created an EKS environment (2 Nodes - 1 GPU, 1 CPU)
- Deployed to EKS as a k8s pod running the above image as a container inside the pod.
- Enabled HPA (Horizontal Pod Autoscaling) to achieve scaling
We're able to achieve a high QPS (Query Per Second) ~15 QPS using the above architecture but we need to scale it to 50 for a client. One approach is to add more nodes and scale the application using Node AutoScaling by AWS.
I was also looking into AWS Lambda functions, although intriguing, I can't really use them since one of my pod i.e API end-point needs to have a GPU enabled for fast inferencing.
I was wondering if there is a better approach of dealing this.