skip to content

Walk through exactly how the high watermark propagates from the leader to the followers, and explain why a follower's HW always lags the leader's HW by at least one fetch round-trip.

level: middleimportance: should knowfreq 45%

answer

  1. pull model: follower fetches, leader never pushes
  2. Fetch carries follower LEO up, leader HW down
  3. follower HW = min(own LEO, leader HW in response)
  4. one round-trip lag is structural
  5. stale follower HW → old truncation bug

basics

~20 s

Followers fetch from the leader. Each Fetch response carries the leader's current HW. The follower applies it after appending the fetched data. Because the HW arrives in the next response, a follower's HW always trails the leader's by one fetch round-trip.

solid answer

~50 s

Replication is follower-pull. A follower sends a Fetch request stating its fetch offset (= its LEO). The leader returns records starting at that offset and includes its current HW in the response. The follower appends the records (advancing its own LEO) and then sets its local HW to min(its LEO, the leader HW it just received). The key asymmetry: the leader can only advance its HW after it sees the follower's *updated* LEO, which happens on the follower's next Fetch — and the follower only learns the resulting HW on the response after that. So there is an inherent one-round-trip lag: leader appends and advances HW at time T; the follower observes that HW at roughly T+1 fetch cycle. This lag is why followers retain slightly older HW values, which matters during failover and truncation.

go deeper

for a junior

Know that followers fetch and the HW comes back in the response.

for a middle

Trace the two-fetch sequence and articulate the one-round-trip lag.

for a senior

Tie the lag to failover correctness and follower-fetch consumer visibility.

for a principal

Reason about how the lag forced the move from HW-based to epoch-based truncation.

## Replication is pull-based Kafka followers **fetch** from the leader; the leader never pushes. This shapes how the HW moves. ## The fetch loop, step by step Assume leader and follower both start with LEO=100, HW=100. A producer writes 5 records (offsets 100..104), so leader LEO becomes 105. 1. **Follower Fetch #1** arrives asking for offset 100, reporting follower LEO=100. The leader records follower LEO=100. Leader recomputes HW = min(leaderLEO=105, followerLEO=100) = 100. (Unchanged.) The response returns records 100..104 **and** the current leader HW=100. 2. **Follower applies the response:** appends 100..104, so follower LEO becomes 105. It sets follower HW = min(followerLEO=105, leaderHW-from-response=100) = 100. So the follower has the *data* but its HW is still 100. 3. **Follower Fetch #2** arrives asking for offset 105, reporting follower LEO=105. Now the leader learns the follower caught up: HW = min(105, 105) = 105. The leader advances its HW to 105. The response carries leader HW=105. 4. **Follower applies #2's response:** sets follower HW = min(105, 105) = 105. ## Why the lag is structural The leader can only advance its HW *after* a follower's LEO reaches the new offset, and it only learns that LEO on the **next** Fetch. The follower then learns the advanced HW on the **response to that next fetch**. So between 'leader advances HW' and 'follower knows the new HW' there is always at least one full request/response cycle. This is fundamental to the pull model, not a bug. ## Why this lag matters - **Failover:** if the leader dies, a follower that becomes the new leader may have a slightly stale HW. The old HW-based protocol used this stale value to decide truncation — and that is precisely the source of the historical truncation data-loss/divergence bugs that KIP-101 (leader epochs) fixed. - **Consumer visibility:** a follower serving fetches (follower-fetching / rack-aware reads, KIP-392) exposes its own HW, which can be slightly behind the leader's — consumers reading from followers may see a marginally older committed boundary. ## Tunables that affect the round-trip - `replica.fetch.wait.max.ms` / `replica.fetch.min.bytes` control how long the leader holds a follower Fetch before responding (long-poll), which directly affects HW propagation latency. - `replica.lag.time.max.ms` governs when a slow follower is dropped from the ISR so its lag stops holding the HW back.

  • Does the follower advance its HW the instant it appends the new records?
    No. It appends (advancing LEO) but sets HW = min(its LEO, the leader HW carried in that response), which still reflects the leader's *previous* HW. It only catches up on the next fetch response.
  • What stops a single slow follower from freezing the HW forever?
    replica.lag.time.max.ms: if a follower fails to fetch up to the leader's LEO within that window, it is removed from the ISR, so the HW = min(ISR LEOs) no longer includes it and can advance.

saying these in an interview costs you the question

  • Saying the leader pushes data/HW to followers (replication is pull).
  • Claiming the follower's HW updates synchronously with the leader's.
  • Forgetting that the follower reports its LEO inside the Fetch request itself.
  • Thinking the lag is a misconfiguration rather than inherent to the pull protocol.

context