skip to content

How do you size a KRaft controller quorum, and what fault tolerance does each size give?

level: principalimportance: should knowfreq 45%

answer

  1. majority = floor(N/2)+1
  2. tolerate floor((N-1)/2)
  3. 3->1 failure, 5->2 failures
  4. odd counts only (4 wastes a node)
  5. brokers are observers, KIP-853 one at a time

basics

~20 s

Use an odd number of controllers, typically 3 or 5. A quorum of N tolerates floor((N-1)/2) failures: 3 tolerate 1, 5 tolerate 2. Odd counts give the best fault tolerance per node since a majority must still be reachable to elect a leader and commit.

solid answer

~50 s

A KRaft controller quorum needs a **majority of voters** alive and reachable to elect a leader and commit metadata. For N voters, it tolerates `floor((N-1)/2)` simultaneous failures: 3 voters tolerate 1, 5 voters tolerate 2, 7 voters tolerate 3. Prefer **odd** counts — going from 3 to 4 voters raises the majority from 2 to 3 but tolerance stays at 1, so you pay more for no gain and add commit latency. **3** is the common default; **5** suits larger or multi-AZ deployments wanting to survive two failures. More voters mean a larger majority to ack each commit, increasing latency, so don't over-provision. Spread voters across failure domains (racks/AZs) so a single domain outage can't take a majority. Voter set changes use the dynamic-reconfiguration mechanism (KIP-853) to add/remove controllers safely one at a time. Brokers are observers and don't count toward quorum, so you scale brokers independently.

go deeper

for a junior

Know that 3 controllers tolerate 1 failure and 5 tolerate 2, and you need a majority alive.

for a middle

State the floor((N-1)/2) rule and why odd counts are preferred over even.

for a senior

Discuss AZ/rack placement, commit-latency vs fault-tolerance trade-offs, and brokers as observers.

for a principal

Reason about KIP-853 dynamic reconfiguration (one voter at a time), combined vs isolated mode at scale, and end-to-end availability design across failure domains.

## The core rule KRaft progress (electing a leader, committing a record) requires a **majority of the voters** to be available. For a quorum of N voters: - Majority = `floor(N/2) + 1` - Failures tolerated = `floor((N-1)/2)` | Voters (N) | Majority | Failures tolerated | |---|---|---| | 1 | 1 | 0 | | 3 | 2 | 1 | | 4 | 3 | 1 | | 5 | 3 | 2 | | 7 | 4 | 3 | ## Why odd numbers Notice 3 and 4 both tolerate exactly **1** failure, but 4 needs **3** acks per commit instead of 2. The extra node adds cost and commit latency without improving fault tolerance. So odd counts (3, 5, 7) are the sweet spots; even counts waste a node. ## Choosing 3 vs 5 - **3 voters**: tolerates 1 failure. Lowest commit latency (only 2 of 3 must ack). Good default for most clusters. - **5 voters**: tolerates 2 failures. Useful when you want to survive losing two nodes or an entire AZ in a 3-AZ layout while still tolerating one more failure. Costs a larger majority (3 acks) and thus slightly higher metadata-commit latency. - **7+**: rare; the bigger majority hurts latency, and metadata write throughput is usually not the bottleneck. ## Failure-domain placement Fault tolerance is only real if failures are independent. Place voters in **distinct racks/availability zones** so that no single failure domain holds a majority. With 3 voters across 3 AZs, losing one AZ leaves 2 — still a majority. With 3 voters where 2 share an AZ, losing that AZ kills the quorum. ## Brokers don't count Controllers are voters; brokers are **observers** that fetch the committed log. You can have hundreds of brokers with a 3- or 5-node controller quorum. Broker count never changes quorum math or commit latency. ## Dynamic voter changes (KIP-853) Earlier KRaft required a static voter set defined at format time (`controller.quorum.voters`). **KIP-853 (KRaft dynamic quorums)** adds online add/remove of voters via `AddVoter`/`RemoveVoter`, configured with `controller.quorum.bootstrap.servers`. Always change membership **one voter at a time** so the old and new majorities overlap and you never split the quorum. ## Operational pitfalls - **Single controller (N=1)**: zero fault tolerance — any controller restart loses the active controller until it returns; only for dev. - **Even voter counts**: avoid; pure waste. - **Co-locating voters in one AZ**: defeats the purpose; a single AZ outage takes the cluster down. - **Over-sizing**: a 7-node quorum to 'be safe' raises commit latency and gains little; 3 or 5 almost always suffices. - **Combined mode at scale**: `process.roles=broker,controller` is fine for small/dev clusters but production usually isolates controllers so heavy broker load can't starve the metadata quorum.

  • Why not run 4 controllers for 'extra safety' over 3?
    4 voters still tolerate only 1 failure (majority is 3), the same as 3, but require an extra ack on every commit — higher latency and an extra node for zero gain in fault tolerance. Odd counts are strictly better.
  • Can you change the controller voter set without downtime, and what's the rule?
    Yes, via KIP-853 dynamic quorum reconfiguration (AddVoter/RemoveVoter). Change one voter at a time so the old and new majorities overlap, preventing a quorum split.
  • If you have 3 controllers and want to survive an entire AZ outage in a 3-AZ region, how should they be placed?
    One controller per AZ. Losing any one AZ leaves 2 of 3 — still a majority — so the quorum keeps making progress.

saying these in an interview costs you the question

  • Saying more voters always means more fault tolerance (4 == 3 tolerance)
  • Recommending even-sized quorums
  • Counting brokers toward quorum/commit
  • Claiming you can swap the whole voter set at once safely
  • Putting a majority of voters in one AZ

context