How does leader election work in KRaft, and what is a leader epoch?
answer
- epoch = Raft term, monotonic
- candidate bumps epoch, self-votes, Vote RPC
- majority of voters wins
- up-to-date log rule (lastEpoch/lastOffset)
- BeginQuorumEpoch announces leader
basics
~20 sWhen voters detect no leader, a candidate increases the epoch (a term counter), votes for itself, and sends Vote requests to peers. If a majority grants votes, it becomes leader for that epoch. The epoch is a monotonically increasing number that totally orders leadership periods.
solid answer
~50 sA KRaft voter that stops hearing from the leader (its election timeout elapses) becomes a candidate: it increments the **epoch** (Raft's term), votes for itself, and sends **Vote** RPCs to the other voters. A voter grants its vote only if the candidate's epoch is at least as new and the candidate's log is at least as up-to-date (last offset and last epoch >= its own) — the **up-to-date log rule** that guarantees no committed record is lost. If the candidate collects votes from a **majority of voters**, it becomes leader for that epoch and announces itself via **BeginQuorumEpoch**. The **leader epoch** is a monotonically increasing integer stamped on every record and on RPCs; it totally orders leadership terms and lets nodes reject stale leaders. Randomized election timeouts reduce the chance of split votes; if a vote splits, the epoch bumps again and a new round runs.
go deeper
Know that an epoch is a term counter and that a leader needs a majority of votes.
Walk through candidate -> Vote RPC -> majority -> BeginQuorumEpoch, and state the up-to-date-log rule.
Explain the majority-overlap safety argument and randomized timeouts for split-vote avoidance.
Discuss KRaft's pre-vote/fencing extensions, persistence of votes for crash safety, and quorum-size trade-offs.
## Key vocabulary - **Voter**: a controller node that can vote and be elected. - **Epoch (a.k.a. term in classic Raft)**: a monotonically increasing integer that labels each leadership period. Every record in the log and many RPCs carry the epoch of the leader that produced/sent them. - **Candidate**: a voter currently trying to get elected. - **Election timeout**: a randomized timer; if a voter hears nothing from a valid leader before it fires, it starts an election. ## The election sequence 1. **Trigger**: A voter's election timeout fires (it has not received a valid Fetch response/heartbeat from the current leader, or it just started up with no leader). 2. **Become candidate**: It increments the epoch to `currentEpoch + 1`, votes for itself, and persists its vote (so it cannot vote twice in the same epoch even after a crash). 3. **Request votes**: It sends a **Vote** RPC to every other voter, including its `lastOffset` and `lastEpoch` (the offset/epoch of the final record in its own log). 4. **Granting a vote**: A peer grants the vote only if BOTH: (a) the candidate's epoch is not older than the peer's, and (b) the candidate's log is **at least as up-to-date** — its `lastEpoch` is higher, or equal-and-its `lastOffset >= peer's`. The peer records that it voted in this epoch so it won't vote again. 5. **Win**: With votes from a **strict majority** of voters (e.g., 2 of 3), the candidate becomes **leader** for that epoch. 6. **Announce**: The new leader sends **BeginQuorumEpoch** to the other voters so they recognize it and start fetching from it. ## Why the up-to-date-log rule matters This rule is the heart of Raft's safety: a candidate can only win if it already holds every **committed** record. Because a committed record lives on a majority, and a winning candidate needs a majority's votes, those two majorities overlap in at least one voter — guaranteeing the new leader is not missing committed data. This is the **Election Safety / Leader Completeness** property. ## Split votes and liveness If two candidates start simultaneously, votes can split with no majority. Each voter votes at most once per epoch, so neither wins; timeouts fire again, the epoch bumps, and a fresh round runs. **Randomized** timeouts make simultaneous candidacies unlikely, so the cluster converges quickly in practice. ## KRaft-specific details - KRaft adds a **Pre-Vote / pre-candidate** style check (and a leader-fencing mechanism) to avoid disruptions from a partitioned node that keeps bumping the epoch — a node checks it could win before actually incrementing the epoch and forcing a new election. - Epochs also appear in **Fetch** responses; a follower learns about a newer epoch and steps down/re-routes accordingly. - A node that sees a higher epoch than its own always steps down to follower and adopts that epoch. ## Common pitfalls - Thinking the highest broker id or lowest node id wins — election is by majority vote constrained by log freshness, not by id. - Forgetting that votes are persisted, so a crashed-and-restarted voter still honors its earlier vote in that epoch.
- Why is the 'log at least as up-to-date' check required before granting a vote?Because committed records live on a majority and a winner needs a majority's votes; the two majorities overlap, so the winner already has every committed record — no committed data is lost across the leadership change.
- What stops a partitioned voter from repeatedly disrupting a healthy leader by bumping the epoch?KRaft uses a pre-vote / candidate-check (and leader fencing): a node verifies it could actually win a majority before incrementing the epoch and triggering a new election, so a lone partitioned node cannot force re-elections.
saying these in an interview costs you the question
- Saying the node with the highest/lowest id automatically becomes leader
- Claiming a candidate can win without a majority
- Saying epochs can stay the same or decrease across elections
- Granting a vote based only on epoch while ignoring log freshness