skip to content

How do leader epochs and fencing prevent split-brain when a deposed KRaft controller (or a partition leader) comes back without realizing it was replaced?

level: seniorimportance: must knowfreq 50%

answer

  1. epoch = monotonic fencing token
  2. new leader → strictly higher epoch
  3. stale epoch request → FENCED_LEADER_EPOCH
  4. majority elects one; fencing kills zombie
  5. OffsetForLeaderEpoch → truncate divergence (KIP-101/279)

basics

~20 s

Every leadership term has a number called the epoch that only increases. When a new leader is elected the epoch goes up. A stale old leader still on the lower epoch is rejected (fenced) by everyone, so it can't commit anything or be obeyed.

solid answer

~50 s

Split-brain is when two nodes both believe they are leader and both act. KRaft prevents it with monotonically increasing **epochs** (Raft terms for the controller; **leader epochs** for partitions). When a new controller wins an election, it does so under a strictly higher epoch. Voters and brokers stamp the current epoch on requests. Any request carrying an epoch lower than what a node has already seen is **fenced** — rejected with a 'fenced leader epoch' style error. So an old, isolated controller that thinks it's still active is on a stale epoch: its replication is refused by the majority, it can never commit, and brokers ignore it. The same mechanism guards partition data: a returning partition leader on an older leader epoch is fenced, and consumers/replicas use epochs (plus the OffsetForLeaderEpoch API) to detect and truncate divergent records. Combined with majority quorum, fencing guarantees at most one effective leader.

go deeper

for a junior

Know the word: each leadership term has an ever-increasing epoch number, and old leaders on a lower number get rejected.

for a middle

Explain that a new leader has a higher epoch and stale-epoch requests are fenced (rejected), so a zombie leader is harmless.

for a senior

Walk the partition scenario end to end and connect controller epochs to partition leader epochs and OffsetForLeaderEpoch truncation.

for a principal

Frame epoch as a fencing token, argue why majority + fencing are jointly necessary, and tie to KIP-101/279 divergence fixes and broker fencing.

## What split-brain is **Split-brain** is the failure mode where, after a network partition or a slow/paused process, two nodes both believe they hold leadership and both accept writes — producing two divergent histories that later conflict. A correct distributed system must make this impossible or at least make the stale leader harmless. ## Epochs: the fencing token KRaft (and Kafka generally) uses a **monotonically increasing epoch number** as a fencing token. Two related uses: - **Controller / Raft epoch (term)**: incremented on every controller election. Only one leader exists per epoch. - **Leader epoch (partition level)**: each partition leadership term has an epoch, recorded in the log and the leader-epoch checkpoint. The invariant: epochs **only increase**. A new leader always has a strictly higher epoch than any predecessor. ## How fencing works Every participant remembers the highest epoch it has observed. When it receives a request: - If the request's epoch is **lower** than what it has seen, it **rejects** the request — Kafka surfaces errors like `FENCED_LEADER_EPOCH` / `NOT_LEADER_OR_FOLLOWER` / a stale-epoch vote rejection. The stale sender is told it is no longer leader and steps down. - If the epoch is **equal or higher**, it processes (and advances its own remembered epoch). ### The classic split-brain scenario, defused 1. Controller A is active at epoch 7. A network partition isolates A from the majority. 2. The majority can't reach A, times out, and elects controller B at epoch 8 (the freshness + majority rules ensure B is up-to-date and unique). 3. A, isolated, still thinks it's leader at epoch 7. It tries to replicate metadata. 4. Every voter A can reach (if any) sees epoch 7 < 8 and **rejects** it. A can never get a majority to commit anything — even if it weren't partitioned, its epoch is stale. 5. When A rejoins, it learns epoch 8 exists, recognizes it is deposed, and reverts to follower, truncating any uncommitted records past the divergence point. A was never able to do damage: it couldn't reach a majority (no commits) and its stale epoch got it fenced the moment it touched a node that had seen epoch 8. ## Partition-level fencing and truncation The same idea protects topic data. Followers and consumers use the **leader epoch** plus the `OffsetForLeaderEpoch` API to find the exact offset where their log diverges from the current leader's, then **truncate** the divergent suffix. This is what replaced the older, lossy high-watermark-only truncation (KIP-101 / KIP-279) and prevents log divergence when a former leader returns. ## Why epochs + majority together are required - **Majority quorum** ensures only one leader can be *elected* per epoch (you can't get two majorities from disjoint sets). - **Epoch fencing** ensures a leader from an *old* epoch is *neutralized* even if it never noticed it lost leadership (the 'zombie leader' / GC-pause / paused-VM case). Neither alone is enough: majority stops a second concurrent election; fencing stops a stale survivor. Together they give 'at most one effective leader' — the formal split-brain guarantee. ## Practical notes - Fencing also covers brokers: a broker that misses heartbeats is moved to a **fenced** state by the controller and must re-register before serving, so a zombie broker can't masquerade as in-sync. - The fencing token pattern here is the same one Martin Kleppmann describes for safe locking — a monotonic number checked at the resource — applied to leadership.

  • Why isn't majority quorum alone enough to stop split-brain?
    Majority prevents two leaders from being elected concurrently, but a previously-elected leader that was paused (GC, VM stall, partition) might wake up still thinking it's leader. Epoch fencing neutralizes that zombie by rejecting its stale-epoch requests; you need both mechanisms.
  • What error does a stale leader's request produce, and what does it do next?
    It gets fenced — errors like FENCED_LEADER_EPOCH or NOT_LEADER_OR_FOLLOWER. On seeing a higher epoch, the stale node recognizes it's deposed, steps down to follower, and truncates any uncommitted records past the divergence point.
  • How does a follower know exactly which records to truncate when it rejoins?
    It uses the leader epoch and the OffsetForLeaderEpoch API to find the offset where its log diverges from the new leader's, then truncates everything after that point (KIP-101/KIP-279).

saying these in an interview costs you the question

  • Saying epochs can reset or decrease — they are strictly monotonic.
  • Claiming a stale leader can still commit writes if it's only briefly partitioned (it can't reach a majority and is fenced).
  • Confusing fencing (rejecting stale leaders) with simple leader election.
  • Believing high-watermark truncation alone prevents divergence (leader-epoch truncation / KIP-101 was needed precisely because it didn't).

context