We have SaaS application and we have thousands of customers. When our customers website get traffic then we also get same traffic as we are tracking activities of our customer's website visitors.
We couldn't get at which time we get sudden spike and all of our servers got down when we get sudden request spike due to traffic in our customer's website. To handle this we have configured to scale when our CPU or memory usage go beyond 60%. Which means we are paying 40% extra cost for unused resource. If we set it as 90% then our all servers became unresponsive due to sudden load and resource usage.
Instead of scale at 60%, we want to utilise at least 90% of resource we are paying for. Is there any better way to do scaling in cost effective way?
Note : We are using AWS ElasticBeanstalk and also GoogleCloud's Kubernetes Engine services.