Disk latency causing CPU spikes on EC2 instance

Viewed 1587

We are having an interesting issue where we are seeing a CPU spike on our EC2 instance and at the same time we are seeing a spike in disk latency. Here is the pattern for CPU spike

  1. CPU spike from 50% to 100% within 30 seconds
  2. It stays at 100% utilization for two minutes
  3. CPU utilization is dropped from 100 to almost 0 in 10 seconds. At the same time almost disk latency is also back to normal

This issue has happened on different AWS ec2 instances a couple of times over a week and still happening. In all cases we are seeing CPU spike along with disk latency with CPU spike having a similar pattern as above.

We had put process monitoring tools to check if any particular process was occupying the CPU. That tool revealed that each of process on the ec2 instance starts taking approx twice the CPU. For eg our app server CPU utilization increases from .75% to 1.5 . Similar observation for Nginx and other processes. There was no single process occupying more than 8% CPU. We studied our traffic pattern and there is nothing unusual which can cause this. So the question is

  1. Can increase in disk latency cause the CPU spike pattern as above or in general can disk latency result in CPU spike
2 Answers

Here is my bet: you are running t2 / t3 machines which are burstable instances. You can access 30% of the CPU all the time, and a credit system create a fair usage predictable mode for the 70% remaining. You earn credit by running the instance, you lose credit by going over 30% CPU usage.

You are running out of credits and then AWS reduce your access to CPU. The system goes smooth again when credits are added to your balance.

t2 and t3 doesn't have the system credit system, you can find details here: CPU Credits and baseline

You have two solutions:

  • Take a bigger instance, so you will have more credits per hour and better baseline or another family like c5, m5, r5 etc...
  • Take an unlimited mode option for your t3 instances

I would suggest faster storage. cpu aims to add up to 100%. limiting is working in this strange way that it simulates usage for "unknown" reason. Reasons can be one of those:

  • idle time (notice here this is what you consider FREE cpu, thats why I say it adds up to 100%)
  • user time (normal usage)
  • system time (system usage)
  • iowait (your case, cpu waiting for HDD/SSD to answer)
  • nice time (low priority processes that were not included in user time)
  • interupt time (external device "talk" time - could be your case if you have many usb devices etc - rather unlikely)
  • softirq (queued work from a processed interrupt - see above)
  • steal time (case that Clement is describing)

I would suggest ensuring which one is your case

you can try below to get the info:
$ sudo apt-get install sysstat
$ mpstat -P ALL 1

From here there is 2 options for you :)

  1. EBS allows you to run IO optimized volume called "IO1" (mid price - mid speed)
  2. Change the machine and use one in "Nitro System" (provides bare metal capabilities - that is: as if you had actual NVMe connected directly - max possible speed)
m5.2xlarge  8   37  32 GiB  EBS Only    $0.384 per Hour
m5d.2xlarge 8   37  32 GiB  1 x 300 NVMe SSD    $0.452 per Hour

Source: Instances built on the Nitro System

Related