Problem with EC2 t3.micro being unresponsive (Ubuntu 20.04)

Viewed 143

For a few weeks I've been trying to find out the reason for random downtimes of an EC2 (t3 micro, has a root volume and and additional EBS volume, both gp3).

In EC2 logs (journalctl) there are some errors from awsagent failing to get instance-id from metadata. Shortly after this message the instance becomes unresponsive (that's looking at the time of this message in the logs and the time of outages).

Errors reference Publishers/ArsenalPublisher and Core/MetaDataClient. Stopping/starting the instance brings it back to life. During the downtimes, CPU usage goes up to 60-65%, Read throughput on the root volume spikes too (there is not much traffic to the instance in general, though), the applications on the server are unavailable, and it's impossible to SSH into it. Logs of CPU usage by process during the downtime show that no process is using more than 1% of the CPU.

Sometimes the instance still shows as healthy, sometimes 1 out of 2 health checks fails. On the server (which is Ubuntu 20.04 LTS) there is a Wordpress site and a Mediawiki install. There is a mySQL database, and data has been moved to the additional EBS volume. Any idea what might be going on?

0 Answers
Related