Avoiding a Kafka Unclean Leader Election

Viewed 163

Context: 6 node kafka cluster. High volume of reads/consumers & writes.

Requirement/desire: Very little/0 downtime

Current state: Every once in a while (so far maybe once every 6 months/year) an unclean leader election occurs

Literature/References:

--Unclean leader election description: https://www.datadoghq.com/blog/kafka-at-datadog/#unclean-leader-elections-to-enable-or-not-to-enable

--Possible solution (two kafka clusters): https://www.datadoghq.com/blog/kafka-at-datadog/#unclean-leader-elections-to-enable-or-not-to-enable

--Second possible solution on https://www.datadoghq.com/blog/kafka-at-datadog/#unclean-leader-elections-to-enable-or-not-to-enable is to turn off unclean leader election config, but this is not acceptable for us, because instead of having an unclean leader election, when the config is off a partition might become inaccessible (quote: "If it cannot elect a new leader, Kafka will halt all reads and writes to that partition"). And we require all data to be accessible

--the new kafka doesn't require zookeeper right? So is this a non issue just by upgrading kafka?

Question: Is there something inherent to our kafka cluster that can cause the unclean leader election that we can avoid? From the literature above datadog seems to suggest just running two kafka clusters (presumably with some dedup layer coordinating between the two fo them?) but of course we would like to avoid the overhead that comes with that. It causes about 10-20 minutes of downtime while the unclean leader election resolves itself

0 Answers
Related