When the active KRaft controller fails, how does the quorum elect a new active controller and resume serving metadata?
answer
- fetch timeout → candidate → bump epoch
- win needs majority of votes
- vote only if candidate log >= mine
- one vote per epoch → one leader
- KRaft failover hot, ZK was cold reload
basics
~20 sFollowers stop hearing heartbeats from the dead leader, a candidate starts a new election term, and whichever candidate gets votes from a majority of voters becomes the new active controller and continues the metadata log.
solid answer
~50 sKRaft uses Raft leader election. The active controller periodically sends Fetch/heartbeat-style replication to followers. When a follower's election timeout (controller.quorum.election.timeout.ms / fetch.timeout.ms) elapses with no contact, it becomes a candidate, bumps the epoch (Raft term), votes for itself, and requests votes. A candidate wins only if a majority of voters grant it their vote, and a voter will only grant it if the candidate's log is at least as up-to-date as its own. The winner becomes active controller for the new epoch, replays/continues the committed metadata log, and brokers reconnect to it. Because a quorum needs a majority and each voter votes once per epoch, at most one leader can win an epoch — preventing two active controllers. Typical failover is sub-second to a few seconds. The new controller is in-sync because election requires an up-to-date log, so no committed metadata is lost.
go deeper
Know that if the active controller dies, the others elect a replacement automatically.
Name the election-timeout configs and that a majority vote elects the new leader.
Explain the epoch bump, the log-freshness vote condition, and why that prevents data loss and dual leaders.
Contrast KRaft hot failover with ZK cold reload, reason about timeout tuning vs false elections, and the consistency guarantees of the freshness rule.
## The roles In a KRaft quorum, voters are in one of three Raft states: **leader** (the single active controller), **follower** (replicating from the leader), or **candidate** (trying to become leader during an election). Exactly one leader exists per epoch when the quorum is healthy. ## Detecting the failure The active controller continuously serves metadata-log records to followers. In KRaft, followers actually *pull* via a Fetch protocol (rather than the leader pushing), but functionally there is a liveness signal. Each follower runs an **election timeout**. Relevant configs: - `controller.quorum.fetch.timeout.ms` — how long a follower waits without a successful fetch from the leader before suspecting it is dead. - `controller.quorum.election.timeout.ms` — how long a candidate waits to win an election before retrying with a new epoch. - `controller.quorum.election.backoff.max.ms` — randomized backoff to avoid repeated split votes. When a follower's fetch timeout elapses with no contact, it concludes the leader is gone. ## The election 1. The follower transitions to **candidate**, increments the **epoch** (Raft's monotonically increasing term number), and votes for itself. 2. It sends `Vote` requests to all other voters. 3. A voter grants its vote only if: (a) it has not already voted in this epoch, and (b) the candidate's log is **at least as up-to-date** as the voter's own (compared by last-entry epoch then offset). This 'log freshness' check is what guarantees the new leader has all committed metadata. 4. If the candidate collects votes from a **majority** of voters, it becomes the new **leader** for that epoch. 5. If no candidate wins (e.g. a split vote), each backs off a randomized interval and retries with a higher epoch, until one wins. Because each voter casts at most one vote per epoch and a win needs a majority, two different candidates cannot both win the same epoch — so there is never more than one active controller. This is the structural guarantee against split-brain. ## Resuming service The new leader already holds every committed metadata record (guaranteed by the freshness check). It may have a few uncommitted records from the old leader that it now commits or truncates per Raft rules. It then begins accepting metadata writes under the new epoch. Brokers, which were fetching metadata, discover the new leader (via the quorum endpoints / bootstrap servers) and resume fetching. Producers and consumers see at most a brief metadata stall. ## Timing Failover is typically sub-second to a few seconds — far faster than the old ZooKeeper-based controller, which had to reload all metadata from ZK on failover (the 'controller failover' cold-start problem KRaft was designed to fix). KRaft's metadata is already a replicated log on every voter, so the new controller is hot. ## Edge cases - **No majority available**: if too many voters are down, no election can succeed; the quorum stays leaderless and metadata writes block. - **Repeated split votes**: randomized election backoff breaks the symmetry so elections converge. - **Stale candidate**: a voter with a lagging log cannot win, because peers refuse to vote for a less-up-to-date log — this is what prevents data loss on failover.
- Why can't a controller with a lagging metadata log win the election?Voters only grant their vote to a candidate whose log is at least as up-to-date as their own (by last-entry epoch then offset). A lagging candidate is refused by up-to-date voters, so it can't reach a majority — this prevents committed metadata loss.
- Why is KRaft failover faster than the old ZooKeeper controller failover?With ZooKeeper, the newly elected controller had to read the entire cluster metadata from ZK on failover (cold start). In KRaft the metadata is already a replicated log present on every voter, so the new controller is hot and resumes almost immediately.
saying these in an interview costs you the question
- Saying the new controller reloads all metadata from disk/ZK on failover (that was the ZK design KRaft removed).
- Claiming any follower can become leader regardless of log state (the freshness check forbids stale leaders).
- Describing election without a majority requirement, which would permit two leaders.
- Confusing the Raft epoch with the partition leader epoch on data topics (related concept, different log).