RabbitMQ randomly disconnecting application consumers in a Kubernetes/Istio environment

Viewed 168

Issue:

My company has recently moved workers from Heroku to Kubernetes. We previously used a Heroku-managed add-on (CloudAMQP) for our RabbitMQ brokers. This worked perfectly and we never saw issues with dropped consumer connections.

Now that our workloads live in Kubernetes deployments on separate nodegroups, we are seeing daily dropped consumer connections, causing our messages to not be processed by our applications living in Kubernetes. Our new RabbitMQ brokers live in CloudAMQP but are not managed Heroku add-ons.

Errors on the consumer side just indicate a Unexpected disconnect. No additional details. No errors on the Istio envoy proxy level that is evident. We do not have a Istio Egress, so no destination rules set here. No errors on the RabbitMQ server that is evident.

Remediation Attempts:

  1. Read all StackOverflow/GitHub issues for the Unexpected errors we are seeing. Nothing we have found has remediated the issue.

  2. Our first attempt to remediate was to change the heartbeat to 0 (disabling heartbeats) on our RabbitMQ server and consumer. This did not fix anything, connections still randomly dropping. CloudAMQP also suggests disabling this, because they rely heavily on TCP keepalive.

  3. Created a message that just logs on the consumer every five minutes. To keep the connection active. This has been a bandaid fix for whatever the real issue is. This is not perfect, but we have seen a reduction of disconnects.

What we think the issue is:

We have researched why this might be happening and are honing in on network TCP keepalive settings either within Kubernetes or on our Istio envoy proxy's outbound connection settings.

Any ideas on how we can troubleshoot this further, or what we might be missing here to diagnose?

Thanks!

0 Answers
Related