Network Partition in RabbitMQ

Viewed 179

I am trying to analyze how RabbitMQ Partition theorum(pause_minority,pause_if_all_down,autoheal) work? I have reproduced the network partition in three-node clusters in GCP but I was unable to conclude which node will be stopped when there is a network jitter in between them.

I have used perftest for creating a production-like environment and used iptables concept in order to bring partition between two nodes. I created ten queues with replication factor 2(master and one slave) and was using min_master for the uniform distribution of queues. Publishing rate was 1000/sec(100/sec for each queue) Consumption rate was 1000/sec(100/sec for each queue)

Please see the test results here for Pause_Minority

Explanation: Taking the first row, 8,9, and 6 are the connections(before partition) on Node A, B, and C respectively. I have blacklisted Node A with B and the results are Node A and Node B stopped running and connections are transferred to Node C. I got different results for rows 2 and 3(Please See the Linked Image)

Please see the test results here pause_if_all_down

Note:0 Connections means (no publishing and consumption)

Please see the test results here autoheal

For Pause Minority I read this article in which the author has explained the master slave architecture but I was unable to get the resutls as per the blog. I am also attaching the link for the google sheet where I have shared the resuts of my test in detail Other Artciles that I have read are as listed:

https://www.rabbitmq.com/partitions.html https://docs.vmware.com/en/VMware-Tanzu-RabbitMQ-for-Kubernetes/1.2/tanzu-rmq/GUID-partitions.html

Can anyone explain to me how partiton theorums decide which node will be stopped in case of partition? Does it decide on the basis of a number of queues, or no of connections, or anything else?

0 Answers
Related