skip to content

How does leader election work in KRaft, and what is a leader epoch?

level: middleimportance: must knowfreq 65%

answer

  1. epoch = Raft term, monotonic
  2. candidate bumps epoch, self-votes, Vote RPC
  3. majority of voters wins
  4. up-to-date log rule (lastEpoch/lastOffset)
  5. BeginQuorumEpoch announces leader

basics

~20 s

When voters detect no leader, a candidate increases the epoch (a term counter), votes for itself, and sends Vote requests to peers. If a majority grants votes, it becomes leader for that epoch. The epoch is a monotonically increasing number that totally orders leadership periods.

solid answer

~50 s

A KRaft voter that stops hearing from the leader (its election timeout elapses) becomes a candidate: it increments the **epoch** (Raft's term), votes for itself, and sends **Vote** RPCs to the other voters. A voter grants its vote only if the candidate's epoch is at least as new and the candidate's log is at least as up-to-date (last offset and last epoch >= its own) — the **up-to-date log rule** that guarantees no committed record is lost. If the candidate collects votes from a **majority of voters**, it becomes leader for that epoch and announces itself via **BeginQuorumEpoch**. The **leader epoch** is a monotonically increasing integer stamped on every record and on RPCs; it totally orders leadership terms and lets nodes reject stale leaders. Randomized election timeouts reduce the chance of split votes; if a vote splits, the epoch bumps again and a new round runs.

go deeper

for a junior

Know that an epoch is a term counter and that a leader needs a majority of votes.

for a middle

Walk through candidate -> Vote RPC -> majority -> BeginQuorumEpoch, and state the up-to-date-log rule.

for a senior

Explain the majority-overlap safety argument and randomized timeouts for split-vote avoidance.

for a principal

Discuss KRaft's pre-vote/fencing extensions, persistence of votes for crash safety, and quorum-size trade-offs.

## Key vocabulary - **Voter**: a controller node that can vote and be elected. - **Epoch (a.k.a. term in classic Raft)**: a monotonically increasing integer that labels each leadership period. Every record in the log and many RPCs carry the epoch of the leader that produced/sent them. - **Candidate**: a voter currently trying to get elected. - **Election timeout**: a randomized timer; if a voter hears nothing from a valid leader before it fires, it starts an election. ## The election sequence 1. **Trigger**: A voter's election timeout fires (it has not received a valid Fetch response/heartbeat from the current leader, or it just started up with no leader). 2. **Become candidate**: It increments the epoch to `currentEpoch + 1`, votes for itself, and persists its vote (so it cannot vote twice in the same epoch even after a crash). 3. **Request votes**: It sends a **Vote** RPC to every other voter, including its `lastOffset` and `lastEpoch` (the offset/epoch of the final record in its own log). 4. **Granting a vote**: A peer grants the vote only if BOTH: (a) the candidate's epoch is not older than the peer's, and (b) the candidate's log is **at least as up-to-date** — its `lastEpoch` is higher, or equal-and-its `lastOffset >= peer's`. The peer records that it voted in this epoch so it won't vote again. 5. **Win**: With votes from a **strict majority** of voters (e.g., 2 of 3), the candidate becomes **leader** for that epoch. 6. **Announce**: The new leader sends **BeginQuorumEpoch** to the other voters so they recognize it and start fetching from it. ## Why the up-to-date-log rule matters This rule is the heart of Raft's safety: a candidate can only win if it already holds every **committed** record. Because a committed record lives on a majority, and a winning candidate needs a majority's votes, those two majorities overlap in at least one voter — guaranteeing the new leader is not missing committed data. This is the **Election Safety / Leader Completeness** property. ## Split votes and liveness If two candidates start simultaneously, votes can split with no majority. Each voter votes at most once per epoch, so neither wins; timeouts fire again, the epoch bumps, and a fresh round runs. **Randomized** timeouts make simultaneous candidacies unlikely, so the cluster converges quickly in practice. ## KRaft-specific details - KRaft adds a **Pre-Vote / pre-candidate** style check (and a leader-fencing mechanism) to avoid disruptions from a partitioned node that keeps bumping the epoch — a node checks it could win before actually incrementing the epoch and forcing a new election. - Epochs also appear in **Fetch** responses; a follower learns about a newer epoch and steps down/re-routes accordingly. - A node that sees a higher epoch than its own always steps down to follower and adopts that epoch. ## Common pitfalls - Thinking the highest broker id or lowest node id wins — election is by majority vote constrained by log freshness, not by id. - Forgetting that votes are persisted, so a crashed-and-restarted voter still honors its earlier vote in that epoch.

  • Why is the 'log at least as up-to-date' check required before granting a vote?
    Because committed records live on a majority and a winner needs a majority's votes; the two majorities overlap, so the winner already has every committed record — no committed data is lost across the leadership change.
  • What stops a partitioned voter from repeatedly disrupting a healthy leader by bumping the epoch?
    KRaft uses a pre-vote / candidate-check (and leader fencing): a node verifies it could actually win a majority before incrementing the epoch and triggering a new election, so a lone partitioned node cannot force re-elections.

saying these in an interview costs you the question

  • Saying the node with the highest/lowest id automatically becomes leader
  • Claiming a candidate can win without a majority
  • Saying epochs can stay the same or decrease across elections
  • Granting a vote based only on epoch while ignoring log freshness

context