Explain how leader epochs and the OffsetsForLeaderEpoch request (KIP-101, completed by KIP-279) let a follower truncate to the correct divergence point instead of to the high watermark.
answer
- epoch = monotonic int, bumped per election
- stamped in batch header + leader-epoch-checkpoint file
- OffsetsForLeaderEpoch → end offset of an epoch
- truncate to divergence point, not HW
- KIP-279 fixes unknown-epoch follower case
basics
~20 sEach leader is assigned a monotonically increasing leader epoch, and the log records which epoch produced each offset range (the leader-epoch cache). On becoming a follower, the broker asks the leader, via OffsetsForLeaderEpoch, for the end offset of its last known epoch, and truncates exactly there — the true point where the logs diverge — instead of guessing with the HW.
solid answer
~50 sA leader epoch is a monotonically increasing integer bumped on every leader election; it is stamped into each record batch and summarized in a per-partition leader-epoch checkpoint file mapping epoch → start offset. After a leader change, a follower doesn't truncate to its HW. Instead it sends an OffsetsForLeaderEpoch request naming its latest leader epoch; the leader replies with the end offset of that epoch in the leader's log (the offset where the next epoch begins, or the leader's LEO if it shares that epoch). The follower truncates to that returned offset — the precise point up to which the two logs are guaranteed identical — then re-fetches. This eliminates both data loss and log divergence because the decision is based on actual log lineage, not the lagging commit boundary. KIP-279 extended this so a follower querying a leader whose epoch the follower has never seen gets the start offset of the first epoch larger than the requested one, fixing a follower-to-follower divergence case KIP-101 missed.
go deeper
Know that each leader has an epoch number used to truncate correctly.
Explain the epoch cache and that followers query the leader for an epoch's end offset.
Detail the OffsetsForLeaderEpoch recovery flow and how it removes divergence.
Explain the KIP-279 refinement and how epoch lineage underpins compaction/tiered-storage/follower-fetch.
## What a leader epoch is A **leader epoch** is a 32-bit, monotonically increasing integer (sometimes called the 'leader generation'). It is incremented **every time a new leader is elected** for a partition (managed by the controller). Crucially, the epoch is **persisted into the log**: every record batch header carries the `partitionLeaderEpoch` of the leader that wrote it, and each replica maintains a **leader-epoch cache / checkpoint file** (`leader-epoch-checkpoint`) that stores a compact list of `(epoch, startOffset)` pairs — the first offset at which each epoch begins. So the log is no longer an anonymous stream of offsets; it is tagged with *who* (which leader generation) wrote each range. This is the lineage information the HW lacked. ## The OffsetsForLeaderEpoch request KIP-101 added a new API, **OffsetsForLeaderEpoch** (a.k.a. `OFFSET_FOR_LEADER_EPOCH`). A follower asks: 'For leader epoch E, what is the end offset of that epoch in *your* log?' The leader answers using its own epoch cache: - If the leader has epoch E and a later epoch E+1 starting at offset X, it returns X — the offset just past where epoch E's records end. - If the leader's current epoch *is* E (they share the latest epoch), it returns the leader's LEO. That returned offset is the **divergence point**: everything below it is guaranteed identical between the two logs. ## Follower recovery flow (the replacement for HW truncation) 1. Broker becomes follower of a new leader after an election. 2. It looks up the **largest leader epoch in its own log** and the LEO for that epoch. 3. It sends OffsetsForLeaderEpoch(epoch = its latest epoch) to the leader. 4. The leader returns the end offset of that epoch in the leader's log. 5. The follower **truncates to min(returned offset, its own LEO)** — the exact point up to which lineage matches — then begins normal fetching from there. Re-running Scenario B from the old bug: after the double crash, follower A would query 'where does epoch (old) end on leader B?' B answers with the offset where its new epoch began (offset 2), so A truncates away its conflicting m2 at offset 2 and re-fetches B's m2'. The logs converge — no divergence. ## KIP-279: the follower-to-follower gap KIP-101 truncation could still diverge in a chain of rapid leader changes where a follower requests an epoch the new leader **never had** (it skipped that epoch). KIP-279 changed the leader's response rule: if the requested epoch is unknown to the leader, return the **start offset of the first epoch greater than** the requested one (rather than an undefined/too-large answer). This guarantees correctness across all sequences of leader changes. ## Related details - The leader-epoch cache is also used by **log compaction**, **tiered storage**, and **fetch-from-follower** to reason about offsets safely. - Producers/clients also use leader epoch information in the metadata and Fetch/Produce paths (KIP-320) to detect stale leaders and avoid out-of-order or lost data on metadata staleness. - Epoch-based truncation is independent of the HW for *recovery*, but the HW is still what bounds *consumer visibility* — the two mechanisms coexist.
- Where is the leader epoch physically stored so it survives restarts?In two places: inside every record batch header (partitionLeaderEpoch) and in the per-partition leader-epoch-checkpoint file that maps each epoch to its starting offset. Both are on disk, so lineage survives broker restarts.
- Does epoch-based truncation replace the high watermark entirely?No. It replaces the HW only for the *recovery/truncation decision*. The HW still defines the consumer-visible commit boundary and is still computed as min(ISR LEOs). The two coexist.
- What problem did KIP-279 specifically address that KIP-101 left open?A follower-to-follower divergence: when a follower queries a leader for an epoch the leader never had (a skipped epoch in a rapid election chain). KIP-279 made the leader return the start offset of the first epoch greater than the requested one, ensuring a correct truncation point.
saying these in an interview costs you the question
- Saying leader epochs replaced the high watermark for consumer reads (they replaced it only for truncation).
- Claiming the epoch is stored only in memory or only in ZooKeeper.
- Thinking OffsetsForLeaderEpoch returns the HW (it returns an epoch's end offset / divergence point).
- Asserting KIP-101 alone made all leader-change sequences safe (KIP-279 was needed).