skip to content

High Watermark, Leader Epoch and Log Truncation

How the high watermark advances, why followers truncate, and how leader epochs fixed the old truncation data-loss bug. A deep but popular question for senior Kafka roles.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is the difference between the log-end-offset (LEO) and the high watermark (HW) on a Kafka partition leader, and why does it matter to consumers?

level: juniorimportance: must knowfreq 70%

answer

  1. LEO = next offset to write
  2. HW = min(ISR LEOs) = committed boundary
  3. consumers read below HW only
  4. HW hides un-replicated tail
  5. acks=all waits for HW

basics

~20 s

LEO is the offset just past the last message written to the log. HW is the highest offset that has been replicated to all in-sync replicas. Consumers can only read up to (below) the HW, so they never see un-replicated records.

solid answer

~40 s

Every replica tracks its log-end-offset (LEO): the offset of the next record to be appended, i.e. one past the last byte in its local log. The leader also tracks the high watermark (HW): the minimum LEO across all replicas currently in the ISR (in-sync replica set). The HW is the boundary of 'committed' data — records below it exist on every in-sync replica. Consumers are only allowed to fetch messages up to HW-1; anything between HW and LEO is written but not yet fully replicated and is invisible to consumers. This guarantees the read-your-committed-data property: a consumer never sees a record that could later be lost if the leader fails before that record is replicated.

go deeper

for a junior

Recall the two definitions and that consumers read below the HW.

for a middle

Explain HW = min(ISR LEOs) and how Fetch requests propagate LEO/HW.

for a senior

Connect HW to acks/min.insync.replicas durability guarantees and ISR eviction.

for a principal

Reason about HW lag, the read-committed guarantee, and how the boundary interacts with failover design.

## The two offsets A Kafka partition is an append-only log. Each replica (one leader, several followers) keeps its own copy of that log and tracks two key positions: - **LEO (Log-End-Offset):** the offset that will be assigned to the *next* record appended. If a replica has stored records at offsets 0..99, its LEO is 100. Every replica has its own LEO; the leader's LEO advances first when a producer writes, and followers catch up by fetching. - **HW (High Watermark):** the highest offset that is *committed*, meaning it has been replicated to **all replicas currently in the ISR** (In-Sync Replica set). Concretely, the leader computes `HW = min(LEO of all ISR members)`. ## Why consumers stop at the HW Consumers may only read records with offset **< HW**. Records in the range `[HW, LEO)` are physically present on the leader but have not yet been confirmed on every in-sync replica. If the leader crashed at that moment, a new leader elected from the ISR might not have those records, so they could vanish. By hiding them, Kafka guarantees that **once a consumer sees a record, that record is durable** (survives a leader failure within the ISR). ## How it advances 1. Producer sends a batch to the leader. Leader appends it → leader LEO jumps. 2. Followers send Fetch requests. Each Fetch from a follower also reports that follower's current LEO (its fetch offset). 3. The leader updates its view of each follower's LEO and recomputes `HW = min(ISR LEOs)`. 4. The new HW is sent back to followers piggybacked on the next Fetch response, so followers learn the committed boundary slightly after the leader. ## Edge cases / nuances - With **acks=all** + `min.insync.replicas`, a producer is acknowledged only when the record's offset is below the HW (committed on the required number of in-sync replicas). With **acks=1**, it is acknowledged on leader append, before the HW advances — that record can be lost if the leader fails first. - The HW can never exceed any ISR member's LEO, but it can lag the leader's LEO arbitrarily if followers are slow (they get evicted from the ISR via `replica.lag.time.max.ms` so the HW can keep moving). - Followers also track their own HW (learned from the leader); this matters during truncation and leader election.

  • If a producer uses acks=1, can a consumer ever observe a record that later gets lost?
    No. The consumer still only reads below the HW. acks=1 means the producer is acked at leader append (before HW advances), so the producer can lose data, but the consumer-visible boundary is still the HW — it never exposes the un-replicated tail.
  • How does the leader learn each follower's LEO?
    Followers report their current fetch offset (their LEO) inside every Fetch request they send to the leader. The leader uses those to recompute HW = min(ISR LEOs).

saying these in an interview costs you the question

  • Saying the HW is the last offset written to the leader (that's the LEO).
  • Claiming consumers can read up to the LEO.
  • Confusing HW with the committed *consumer* offset (__consumer_offsets) — different concept.
  • Saying HW = max of replica LEOs (it is the min across the ISR).

context

open as a page

Before KIP-101, Kafka followers truncated their logs to the high watermark on becoming a follower of a new leader. Describe the data-loss / log-divergence scenario this caused.

level: seniorimportance: must knowfreq 50%

basics

~20 s

On rejoining, a follower truncated everything above its high watermark, assuming records above the HW were uncommitted. But because the follower's HW lags, this could throw away records that were actually committed, or let two replicas keep different records at the same offset — causing data loss or log divergence.

open as a page

Explain how leader epochs and the OffsetsForLeaderEpoch request (KIP-101, completed by KIP-279) let a follower truncate to the correct divergence point instead of to the high watermark.

level: seniorimportance: must knowfreq 45%

basics

~20 s

Each leader is assigned a monotonically increasing leader epoch, and the log records which epoch produced each offset range (the leader-epoch cache). On becoming a follower, the broker asks the leader, via OffsetsForLeaderEpoch, for the end offset of its last known epoch, and truncates exactly there — the true point where the logs diverge — instead of guessing with the HW.

open as a page

Walk through exactly how the high watermark propagates from the leader to the followers, and explain why a follower's HW always lags the leader's HW by at least one fetch round-trip.

level: middleimportance: should knowfreq 45%

basics

~20 s

Followers fetch from the leader. Each Fetch response carries the leader's current HW. The follower applies it after appending the fetched data. Because the HW arrives in the next response, a follower's HW always trails the leader's by one fetch round-trip.

open as a page

As a principal engineer, how do acks, min.insync.replicas, the high watermark, and leader epochs together determine whether an acknowledged write can ever be lost? What configuration gives the strongest durability and what are the trade-offs?

level: principalimportance: should knowfreq 35%

basics

~20 s

Strongest durability: acks=all, min.insync.replicas=2 (RF=3), and unclean.leader.election.enable=false. Then a write is acked only after the HW advances past it (committed on >=2 in-sync replicas), and leader-epoch truncation prevents divergence on failover. The trade-off is higher latency and reduced availability when replicas fall behind.

open as a page