i'm having problems with nginx as reverse proxy to tornado (python server) instances.
We have 3 servers (dns round robin) with the same configuration, and the problem only exists on one, several times each day. It seems random.
The problem is a very slow response times (2,3s to 2minutes!) on proxy pass requests. Most of requests get an answer under one second.
We have a cron job that 'pings' this server every minute. ICMP pings are always OK, static resources get are always OK, but we have 5-50/day requests to tornado via nginx that are really slow. The nginx is in front of 8 instances of tornado servers with a mongo DB. Each nginx server gets 100-200 requests/minute.
When looking at the app logs, the response times on tornado side are never above one second, the problem really seems to be from nginx-tornado interface.
System monitoring (disk, memory, cpu) is always OK and <30% use. The server houses nginx, tornado and mongodb and is 32gb RAM with 8 threads. Under Ubuntu 18LTS.
I will post the nginx config i need to anonymize, we only have a site-available config that is not crazy.
How can i go further on diagnostic ? I think there is a stacking/enqueue of requests somewhere.
some info:
# ss -lt
State Recv-Q Send-Q Local Address:Port Peer Address:Port
LISTEN 0 128 0.0.0.0:8002 0.0.0.0:*
LISTEN 0 128 0.0.0.0:8003 0.0.0.0:*
LISTEN 0 128 0.0.0.0:8004 0.0.0.0:*
LISTEN 0 128 0.0.0.0:8005 0.0.0.0:*
LISTEN 0 128 0.0.0.0:8006 0.0.0.0:*
LISTEN 0 128 0.0.0.0:8007 0.0.0.0:*
....
~# netstat -s | grep -i LISTEN
2129596 times the listen queue of a socket overflowed
2138743 SYNs to LISTEN sockets dropped
And every day theses numbers increase.
Any idea on where the problem is, and how to resolve this ? Thank you!!