How many controllers (voters) should a KRaft controller quorum have, and how many failures can each size tolerate?
answer
- majority = floor(N/2)+1
- 3 voters → tolerate 1; 5 → tolerate 2
- odd numbers only
- 1 voter = dev, zero tolerance
- controller.quorum.voters
basics
~10 sUse an odd number, usually 3 or 5 voters. A quorum needs a majority alive to work. 3 voters tolerate 1 failure; 5 voters tolerate 2 failures.
solid answer
~40 sA KRaft controller quorum is the set of voter nodes (process.roles includes 'controller') that run the Raft metadata log. It stays available only while a strict majority of voters are reachable and agreeing. Standard sizes are 3 voters (majority 2, tolerates 1 failure) or 5 voters (majority 3, tolerates 2 failures). You always pick an odd number: with N voters the majority is floor(N/2)+1, so failure tolerance is floor((N-1)/2). 3 is the common default for most clusters; 5 is used when you need to survive two simultaneous controller losses (for example, spread across availability zones). A single-voter quorum has no fault tolerance and is only for dev. The quorum is configured via controller.quorum.voters (static) or, in newer releases, dynamically with kafka-metadata-quorum.
go deeper
Memorize: odd voters; 3 tolerates 1 failure, 5 tolerates 2; majority must be alive.
Derive failure tolerance from floor((N-1)/2) and know the controller.quorum.voters config and the kafka-metadata-quorum describe command.
Discuss AZ placement, the latency cost of larger quorums, and when 5 over 3 is justified.
Reason about failure-domain mapping, dynamic quorums (KIP-853), and capacity/latency trade-offs of quorum size across a fleet.
## What a controller quorum is In KRaft mode (KRaft = Kafka Raft, the replacement for ZooKeeper), Kafka's cluster metadata — topics, partitions, broker registrations, ACLs, leader assignments — is stored in an internal replicated log called the metadata log (`__cluster_metadata`). The nodes that replicate and vote on this log are the **controllers**, also called **voters**. A node is a voter when its `process.roles` includes `controller`. The set of voters is the **quorum**. They run the Raft consensus protocol: one voter is the **active controller** (the Raft leader) and the others are followers that replicate its log. ## The majority rule Raft requires a **strict majority** of voters to agree before any metadata write is committed. For a quorum of N voters: - majority = floor(N/2) + 1 - failures tolerated = N − majority = floor((N−1)/2) So: | Voters (N) | Majority needed | Failures tolerated | |-----------|-----------------|--------------------| | 1 | 1 | 0 | | 3 | 2 | 1 | | 5 | 3 | 2 | | 7 | 4 | 3 | If fewer than a majority are alive, the quorum cannot elect a leader or commit writes — metadata becomes read-only/unavailable until enough voters return. Note brokers can keep serving existing produce/consume traffic for a while using their cached metadata, but you cannot create topics, reassign partitions, or recover from broker leadership changes. ## Why these specific sizes - **1 voter**: zero fault tolerance. Dev/test only. - **3 voters**: the standard production default. Survives one controller failure (rolling restart, node crash). Majority = 2. - **5 voters**: survives two simultaneous failures. Used for higher resilience or to spread voters across 3 availability zones so losing one AZ (which could take out two voters) still leaves a majority. - **7+**: rarely needed; more voters means every commit must be acknowledged by more nodes, increasing metadata write latency, and the marginal availability gain shrinks. ## Configuration Statically, voters are listed in `controller.quorum.voters` as `id@host:port` entries. Newer Kafka (KIP-853) supports **dynamic** quorums where you add/remove voters at runtime via `kafka-metadata-quorum` and `controller.quorum.bootstrap.servers`. You inspect the quorum with `kafka-metadata-quorum.sh --describe`. ## Key takeaway Pick an odd number — 3 or 5 — sized to the number of simultaneous controller failures you must survive, and place voters in separate failure domains.
- Why pick 5 voters instead of 3?To tolerate two simultaneous controller failures, or to spread voters across 3 availability zones so losing one AZ (which may host two voters) still leaves a majority of 3.
- What happens to a 3-voter quorum if two controllers die at once?Majority (2) is lost, so no new leader can be elected and no metadata writes commit. The cluster metadata is effectively frozen/read-only until a voter recovers; brokers keep serving from cached metadata but admin operations stall.
saying these in an interview costs you the question
- Saying a quorum of N tolerates N-1 failures (it only tolerates floor((N-1)/2)).
- Recommending an even number like 4 or 6 'for more safety'.
- Claiming the quorum needs ALL voters alive to function (it needs a majority).
- Confusing voter count with broker count — controllers are a separate role.