skip to content

How many controllers (voters) should a KRaft controller quorum have, and how many failures can each size tolerate?

level: juniorimportance: must knowfreq 70%

answer

  1. majority = floor(N/2)+1
  2. 3 voters → tolerate 1; 5 → tolerate 2
  3. odd numbers only
  4. 1 voter = dev, zero tolerance
  5. controller.quorum.voters

basics

~10 s

Use an odd number, usually 3 or 5 voters. A quorum needs a majority alive to work. 3 voters tolerate 1 failure; 5 voters tolerate 2 failures.

solid answer

~40 s

A KRaft controller quorum is the set of voter nodes (process.roles includes 'controller') that run the Raft metadata log. It stays available only while a strict majority of voters are reachable and agreeing. Standard sizes are 3 voters (majority 2, tolerates 1 failure) or 5 voters (majority 3, tolerates 2 failures). You always pick an odd number: with N voters the majority is floor(N/2)+1, so failure tolerance is floor((N-1)/2). 3 is the common default for most clusters; 5 is used when you need to survive two simultaneous controller losses (for example, spread across availability zones). A single-voter quorum has no fault tolerance and is only for dev. The quorum is configured via controller.quorum.voters (static) or, in newer releases, dynamically with kafka-metadata-quorum.

go deeper

for a junior

Memorize: odd voters; 3 tolerates 1 failure, 5 tolerates 2; majority must be alive.

for a middle

Derive failure tolerance from floor((N-1)/2) and know the controller.quorum.voters config and the kafka-metadata-quorum describe command.

for a senior

Discuss AZ placement, the latency cost of larger quorums, and when 5 over 3 is justified.

for a principal

Reason about failure-domain mapping, dynamic quorums (KIP-853), and capacity/latency trade-offs of quorum size across a fleet.

## What a controller quorum is In KRaft mode (KRaft = Kafka Raft, the replacement for ZooKeeper), Kafka's cluster metadata — topics, partitions, broker registrations, ACLs, leader assignments — is stored in an internal replicated log called the metadata log (`__cluster_metadata`). The nodes that replicate and vote on this log are the **controllers**, also called **voters**. A node is a voter when its `process.roles` includes `controller`. The set of voters is the **quorum**. They run the Raft consensus protocol: one voter is the **active controller** (the Raft leader) and the others are followers that replicate its log. ## The majority rule Raft requires a **strict majority** of voters to agree before any metadata write is committed. For a quorum of N voters: - majority = floor(N/2) + 1 - failures tolerated = N − majority = floor((N−1)/2) So: | Voters (N) | Majority needed | Failures tolerated | |-----------|-----------------|--------------------| | 1 | 1 | 0 | | 3 | 2 | 1 | | 5 | 3 | 2 | | 7 | 4 | 3 | If fewer than a majority are alive, the quorum cannot elect a leader or commit writes — metadata becomes read-only/unavailable until enough voters return. Note brokers can keep serving existing produce/consume traffic for a while using their cached metadata, but you cannot create topics, reassign partitions, or recover from broker leadership changes. ## Why these specific sizes - **1 voter**: zero fault tolerance. Dev/test only. - **3 voters**: the standard production default. Survives one controller failure (rolling restart, node crash). Majority = 2. - **5 voters**: survives two simultaneous failures. Used for higher resilience or to spread voters across 3 availability zones so losing one AZ (which could take out two voters) still leaves a majority. - **7+**: rarely needed; more voters means every commit must be acknowledged by more nodes, increasing metadata write latency, and the marginal availability gain shrinks. ## Configuration Statically, voters are listed in `controller.quorum.voters` as `id@host:port` entries. Newer Kafka (KIP-853) supports **dynamic** quorums where you add/remove voters at runtime via `kafka-metadata-quorum` and `controller.quorum.bootstrap.servers`. You inspect the quorum with `kafka-metadata-quorum.sh --describe`. ## Key takeaway Pick an odd number — 3 or 5 — sized to the number of simultaneous controller failures you must survive, and place voters in separate failure domains.

  • Why pick 5 voters instead of 3?
    To tolerate two simultaneous controller failures, or to spread voters across 3 availability zones so losing one AZ (which may host two voters) still leaves a majority of 3.
  • What happens to a 3-voter quorum if two controllers die at once?
    Majority (2) is lost, so no new leader can be elected and no metadata writes commit. The cluster metadata is effectively frozen/read-only until a voter recovers; brokers keep serving from cached metadata but admin operations stall.

saying these in an interview costs you the question

  • Saying a quorum of N tolerates N-1 failures (it only tolerates floor((N-1)/2)).
  • Recommending an even number like 4 or 6 'for more safety'.
  • Claiming the quorum needs ALL voters alive to function (it needs a majority).
  • Confusing voter count with broker count — controllers are a separate role.

context