skip to content

Before KIP-101, Kafka followers truncated their logs to the high watermark on becoming a follower of a new leader. Describe the data-loss / log-divergence scenario this caused.

level: seniorimportance: must knowfreq 50%

answer

  1. old rule: truncate to own HW on becoming follower
  2. HW lags → stale HW
  3. Scenario A: lose committed record on restart
  4. Scenario B: double crash → same offset, different record
  5. HW can't tell whose epoch wrote an offset

basics

~20 s

On rejoining, a follower truncated everything above its high watermark, assuming records above the HW were uncommitted. But because the follower's HW lags, this could throw away records that were actually committed, or let two replicas keep different records at the same offset — causing data loss or log divergence.

solid answer

~50 s

The old recovery rule was: when a replica becomes a follower, truncate its log down to its own HW, then re-fetch from the leader. The flaw: the HW propagates with a lag, so a replica's persisted HW can be behind the true committed offset. Consider unclean or rapid leader changes: a replica that was actually caught up could come back with a stale HW, truncate committed records away, and then re-replicate different records the new leader has at those offsets. Two scenarios result. (1) Data loss: a committed record is truncated and never re-fetched because the new leader also lacks it. (2) Log divergence: after a double leader change, two replicas end up with *different* records at the *same* offset, silently corrupting the log. KIP-101 replaced HW-based truncation with leader-epoch-based truncation using OffsetsForLeaderEpoch to find the exact correct truncation point.

go deeper

for a junior

Aware that old truncation could lose data after leader changes.

for a middle

Explain why a lagging HW makes truncation-to-HW unsafe.

for a senior

Reproduce the divergence scenario step by step and name KIP-101 as the fix.

for a principal

Articulate that the HW is a commit boundary not a lineage marker, motivating epoch-based recovery.

## The old recovery rule Pre-KIP-101 (before Kafka 0.11), when a broker became a **follower** for a partition (e.g., after a leader change or restart), it would: **truncate its local log down to its own high watermark**, then start fetching from the new leader from that point. The assumption was: 'everything at or below my HW is committed and safe; everything above my HW is uncommitted and might differ from the new leader, so discard it.' This is wrong because, as established, **the HW propagates with a lag** — a follower's persisted HW can be *behind* the genuinely committed offset. ## Scenario A — data loss on fast restart 1. Leader L and follower F both have records up to offset 10; both committed, HW=11 on the leader. But F's persisted HW is still 10 (lag) when it crashes. 2. F restarts. Following the rule, it truncates to its HW=10, discarding record at offset 10. 3. Meanwhile L also crashes before anyone else replicated offset 10 beyond F. When F comes back as leader (or a stale replica does), offset 10 is gone everywhere → **committed data lost**, even though it had been acknowledged to the producer. ## Scenario B — log divergence (the classic KIP-101 example) Two brokers A and B, one partition. 1. A is leader, holds offsets up to 2 (records m0, m1, m2). B is follower but its HW is behind; B has m0, m1. 2. **Both brokers crash.** B restarts **first** and, because `unclean.leader.election` is allowed (or it is the only available replica), B becomes leader with log [m0, m1], LEO=2. 3. A producer writes a **new** record m2' at offset 2 to leader B. Now B's log is [m0, m1, m2']. 4. A restarts and becomes a **follower** of B. Under the old rule A truncates to its HW... but its HW may already be past offset 2, so it **keeps its old m2** and only checks against the HW, not the actual records. A and B now both have offset 2 but A has **m2** while B has **m2'** → the logs have **diverged**: the same offset holds different data on two replicas, which can be served to different consumers. The root cause in both cases: the HW is a *coarse* truncation signal that cannot distinguish 'records from the previous leader that the new leader also has' from 'records from a previous leader that conflict with the new leader.' ## Why the HW can't fix it The HW only tells you a *commit boundary*, not *which leader wrote which offsets*. After a leader change, the only safe question is: 'up to what offset do my log and the new leader's log share the same lineage?' The HW cannot answer that. You need to know, per range of offsets, **which leader epoch produced them**. ## The fix (preview) KIP-101 introduced **leader epochs** stamped into the log and the **OffsetsForLeaderEpoch** API so a follower can ask the leader 'where does your epoch X end?' and truncate to that exact divergence point instead of to the HW. KIP-279 closed a remaining follower-to-follower gap. This eliminated both the data loss and the divergence.

  • Did this bug require unclean leader election to manifest?
    The classic divergence example uses unclean election (a replica with a shorter/older log becoming leader), but the underlying flaw — truncating to a lagging HW — could also cause data loss on clean restarts where a committed record sat above a follower's persisted HW.
  • Why isn't 'just make followers truncate to the leader's HW instead of their own' a fix?
    Because it still can't distinguish records of the same offset that came from different leaders. The HW is a commit boundary, not a lineage marker, so it can't locate the actual point where two logs diverge after a leader change.

saying these in an interview costs you the question

  • Claiming the bug only caused performance issues, not data loss/divergence.
  • Saying truncating to the HW is always safe.
  • Asserting min.insync.replicas alone prevented the divergence (it didn't — it's a leader-lineage problem).
  • Confusing this with consumer offset reset behavior.

context