I'm new to Prometheus and I got confused about CPU usage metrics. Here are my two cents about CPU usage. (Please correct me if I'm wrong.)
container_cpu_usage_seconds_total=container_cpu_user_seconds_total+container_cpu_system_seconds_total"cores" =
container_spec_cpu_quota/container_spec_cpu_periodThis is the pod used CPU time in 1 second:
rate(container_cpu_usage_seconds_total{image!=""}[1m]) by (pod, namespace)This is the pod allowed CPU time in 1 second:
(sum(container_spec_cpu_quota{image!=""}/100000) by (pod, namespace))So the utilizations is like:
rate(container_cpu_usage_seconds_total{image!=""}[1m]) by (pod, namespace) / (sum(container_spec_cpu_quota{image!=""}/100000) by (pod, namespace)) *100But what about those pods don't have
spec.cpu.limits? I presume:rate(container_cpu_usage_seconds_total{image!=""}[1m]) by (pod, namespace) / CORES_ON_NODE *100