In a KRaft controller quorum, how is a metadata write committed, how many controller failures can the cluster tolerate, and what happens when quorum is lost?
answer
- commit = majority of voters + high-watermark advances
- tolerate floor((N-1)/2): 3→1, 5→2
- epoch fences stale leaders, no split-brain
- lost majority = metadata frozen, data plane degrades gracefully
- kafka-metadata-quorum.sh describe --status
basics
~20 sA metadata record is committed once a majority of controller voters have persisted it. With N voters the cluster tolerates floor((N-1)/2) failures (so 3 tolerate 1, 5 tolerate 2). If a majority is lost, no new metadata can be committed and the controller becomes read-only until quorum returns.
solid answer
~50 sKRaft uses Raft over the controller voters. The active controller (Raft leader) appends a metadata record to __cluster_metadata and replicates it to followers; the record is committed once a majority of voters (including the leader) have it durably — the high-watermark advances and the record is then applied. Fault tolerance is majority-based: with N voters you tolerate floor((N-1)/2) failures, so 3 voters survive 1 loss and 5 survive 2. Losing the majority means no record can reach quorum: leader election and metadata writes stall. Brokers keep serving existing data using their last-known metadata (so data plane degrades gracefully), but you cannot create topics, change leadership, or apply config/ACL changes until enough voters return. KRaft also uses follower fetch with leader epochs to fence stale leaders, and unclean states are avoided because only committed (majority-acknowledged) records are ever applied — there is no unclean controller election.
go deeper
Know that controllers vote and a majority is needed; 3 controllers tolerate 1 failure.
Explain majority commit, the floor((N-1)/2) formula, and that losing quorum freezes metadata changes.
Detail high-watermark advance, epoch-based fencing, graceful data-plane degradation, and recovery.
Reason about voter sizing trade-offs, partition-isolation scenarios, in-sync vs voter-set distinction, and why KRaft has no unclean controller election.
## Consensus model The controller quorum runs **Raft**. The nodes in `controller.quorum.voters` are **voters**; one is elected **leader** (the **active controller**) for a **term/epoch**, the rest are **followers**. KRaft's variant is **pull-based**: followers *fetch* records from the leader (mirroring how brokers fetch from partition leaders), rather than the leader pushing. ## How a write commits 1. A metadata change (create topic, leadership change, config update) is appended by the **active controller** to `__cluster_metadata` at the next offset, tagged with the current **leader epoch**. 2. Followers fetch and persist the record to their own log. 3. Once a **majority** of voters (the leader counts) have durably written it, the record is **committed**: the **high-watermark** advances past its offset. 4. Only committed records are **applied** to the in-memory metadata image and become visible to brokers. This majority rule is the heart of Raft: a committed record is guaranteed to be present on any future leader, because any new leader must be elected by a majority and majorities overlap. ## Fault tolerance math With **N** voters, a majority is `floor(N/2) + 1`, so the cluster tolerates **`floor((N-1)/2)`** simultaneous voter failures: - 3 voters → majority 2 → tolerate **1** failure. - 5 voters → majority 3 → tolerate **2** failures. - 7 voters → majority 4 → tolerate **3** (rarely worth the extra latency). This is why production runs **3 or 5** controllers (odd numbers; an even count adds a node without adding tolerance). ## Leader election and fencing If the leader fails, followers that miss heartbeats start an election, increment the **epoch**, and a candidate that gets a majority of votes becomes the new leader. The monotonically increasing **epoch** **fences** stale leaders: a deposed leader's writes are rejected because they carry an old epoch. Because elections require a majority, **two leaders cannot both commit** — no split-brain. ## Losing quorum If fewer than a majority of voters are alive (e.g. 2 of 3 controllers down): - **No new metadata can be committed** — there is no majority to acknowledge writes, and no new leader can be elected. - The metadata layer is effectively **read-only/frozen**. - **Brokers continue serving existing partitions** from their cached metadata image, so producers/consumers on already-elected leaders keep working for a while — the data plane degrades gracefully rather than hard-failing. - But you **cannot** create/delete topics, reassign partitions, elect new partition leaders, or change configs/ACLs. If a broker partition leader then also fails, no new leader can be appointed for it, so that partition becomes unavailable. - Recovery: bring back enough voters to restore the majority; the quorum re-forms and resumes committing. ## No unclean controller election Unlike unclean *leader* election for data partitions (which can lose data), KRaft never promotes a controller that lacks committed metadata — the Raft majority/epoch rules guarantee the new leader holds all committed records. There is no "unclean" controller mode that sacrifices metadata consistency. ## Inspecting the quorum `kafka-metadata-quorum.sh describe --status` shows the leader, voters, observers, high-watermark, and each replica's lag — the go-to for diagnosing quorum health. ## Edge cases - A lagging follower that has not caught up still counts toward the *voter set* but only an *in-sync majority* can commit; persistent lag shrinks effective fault tolerance. - Network partitions that isolate the leader from the majority cause it to step down (it cannot advance the high-watermark), and the majority side elects a new leader. - With dynamic quorums (KIP-853), changing the voter set itself is a committed metadata operation, so it too requires quorum.
- If you lose the controller quorum but brokers stay up, can producers and consumers still work?Partly. Brokers serve existing partitions from cached metadata, so traffic on already-elected leaders continues. But no new partition leaders, topics, or config changes can be committed; if a partition leader then fails, that partition cannot get a new leader and becomes unavailable.
- Why is there no 'unclean controller election' in KRaft the way there is unclean leader election for data partitions?Raft only elects a leader that a majority votes for, and majorities overlap, so any new leader already holds every committed metadata record. The epoch fences stale leaders. Thus metadata consistency is never sacrificed; there is no unclean mode.
- How many controllers should a production cluster run and why?Typically 3 (tolerates 1 failure) or 5 (tolerates 2). Odd numbers maximize fault tolerance per node since commits need a majority; more than 5 rarely pays off because every commit waits on a larger majority, adding latency.
saying these in an interview costs you the question
- Saying a write commits when ANY one follower has it (it needs a majority)
- Claiming 4 voters tolerate 2 failures (4 still tolerate only 1; majority is 3)
- Asserting losing quorum immediately kills all producer/consumer traffic (existing partitions keep serving from cached metadata)
- Describing an unclean controller election that loses metadata (KRaft has none)
- Confusing the metadata high-watermark with data-partition ISR semantics