How to debug the reason of an unclosed CLOSE_WAIT connections? (tcpdump etc.)

Viewed 970

We have Java-application and Nginx as a reverse-proxy installed on the same host. Periodically, we got CLOSE_WAIT connections hanging long time:

$ ss -n4t | head
State      Recv-Q Send-Q  Local Address:Port    Peer Address:Port
CLOSE-WAIT 1      0           127.0.0.1:8180       127.0.0.1:36599
CLOSE-WAIT 1      0           127.0.0.1:8180       127.0.0.1:36467
CLOSE-WAIT 1      0           127.0.0.1:8180       127.0.0.1:36154

And the number of them is increasing, then the application is having issues. The internet says:

Your server is failing to detect client disconnects, or ignoring them, and not closing the socket.

Nginx says:

2020/06/28 04:59:15 [error] 65506#0: *31719640 recv() failed (104: Connection reset by peer) while reading response header from upstream, client: 55.55.55.55, server: app.mycompany.com, request: "POST /url/url HTTP/1.0", upstream: "http://127.0.0.1:8180/url/url/provider", host: "app.mycompany.com"

Okay, tcpdump during normal behavior shows:

2020-06-08 06:58:23.073139 IP 127.0.0.1.8180 > 127.0.0.1.57786: Flags [P.], seq 1:738, ack 1211, win 1365, options [nop,nop,TS val 2780380992 ecr 2780380974], length 737
2020-06-08 06:58:23.073233 IP 127.0.0.1.8180 > 127.0.0.1.57786: Flags [F.], seq 738, ack 1211, win 1365, options [nop,nop,TS val 2780380992 ecr 2780380992], length 0
2020-06-08 06:58:23.073302 IP 127.0.0.1.57786 > 127.0.0.1.8180: Flags [F.], seq 1211, ack 739, win 353, options [nop,nop,TS val 2780380992 ecr 2780380992], length 0

[F.] from both sides means that connection had been closed correctly (based on the diagram).

Next, tcpdump during abnormal behavior (when CLOSE_WAIT borns) shows:

2020-06-08 06:55:30.282015 IP 127.0.0.1.8180 > 127.0.0.1.57160: Flags [.], ack 1380, win 1365, options [nop,nop,TS val 2780208201 ecr 2780208201], length 0
2020-06-08 06:57:10.279006 IP 127.0.0.1.57160 > 127.0.0.1.8180: Flags [F.], seq 1380, ack 1, win 342, options [nop,nop,TS val 2780308198 ecr 2780208201], length 0
2020-06-08 06:57:10.318432 IP 127.0.0.1.8180 > 127.0.0.1.57160: Flags [.], ack 1381, win 1365, options [nop,nop,TS val 2780308238 ecr 2780308198], length 0

Here we see only one [F.].

I've read thousands of articles, but still having troubles to answer to the following questions:

What does the tcpdump [F.] flag actually mean? Documentations say: Placeholder, usually used for ACK., but where is a first initial FIN (without ACK) like:

this one? --> 2020-06-08 06:58:23.073233 IP 127.0.0.1.57786 > 127.0.0.1.8180: Flags [F], seq 738, ack 1211, win 1365, options [nop,nop,TS val 2780380992 ecr 2780380992], length 0
2020-06-08 06:58:23.073233 IP 127.0.0.1.8180 > 127.0.0.1.57786: Flags [F.], seq 738, ack 1211, win 1365, options [nop,nop,TS val 2780380992 ecr 2780380992], length 0

My question is, tcpdump combines FIN and FIN/ACK into [F.] flag? If so, during normal behavior we see the following order of actions:

  1. The client sends FIN to the server: omitted by tcpdump
  2. The server sends FIN/ACK to the client: 127.0.0.1.8180 > 127.0.0.1.57786: Flags [F.]
  3. The server sends FIN to the client: omitted by tcpdump
  4. The client sends FIN/ACK to the server: 127.0.0.1.57786 > 127.0.0.1.8180: Flags [F.]

If so, during abnormal behavior the client simply doesn't send us FIN, consequently, our server doesn't send FIN/ACK which is [F.]. Right?

The second concern is Nginx. As I understand on netstat:

CLOSE-WAIT 1      0           127.0.0.1:8180       127.0.0.1:36467

The connection is going from localhost to localhost. Is it because of the Nginx reverse-proxy? That's why I don't see real clients IPs? If so, could it be issue between Java-application and Nginx? I'm asking because I'd like to debug/monitor HTTP traffic including request and response headers and message body from localhost to localhost:

tcpdump -A -s 0 'tcp port 8180 and (((ip[2:2] - ((ip[0]&0xf)<<2)) - ((tcp[12]&0xf0)>>2)) != 0)' -i lo

Cause we have SSL-termination on Nginx, I believe I will see what's going on inside of server/client flow. The third concern is, is it correct way/are there any more ways to understand the root reason clearly why we're getting `CLOSE_WAIT'?

0 Answers
Related