skip to content

When is a produced record considered 'committed', and how does that differ from being written at the leader? What role does the high watermark play?

level: seniorimportance: should knowfreq 60%

answer

  1. LEO = next write offset
  2. HW = min LEO across ISR = committed line
  3. consumers read ≤ HW only
  4. acks=1 returns above HW (risky)
  5. leader epoch (KIP-101) for safe truncation

basics

~20 s

A record is 'written at the leader' once the leader appends it to its log. It becomes 'committed' only once all in-sync replicas have replicated it — at which point the high watermark advances. Consumers only read up to the high watermark, so they never see uncommitted records.

solid answer

~50 s

Kafka distinguishes two states. A record is **written at the leader** the moment the leader appends it to its local log (this is enough for `acks=1`). It becomes **committed** only when every replica in the ISR has fetched and appended it; at that point the **high watermark (HW)** — the offset up to which all ISR members are caught up — advances past the record. Only committed records (offsets below the HW) are exposed to consumers; the leader's **log end offset (LEO)** may be ahead of the HW for records not yet fully replicated. `acks=all` makes the producer wait for the commit (HW advance); `acks=1` returns at the leader-written stage, before the commit. This is why `acks=1` records can be lost: they were written at the leader but never committed, and a leader change can truncate them. The HW is what makes Kafka's read-your-committed-writes guarantee hold regardless of `acks`.

go deeper

for a junior

Know that committed means replicated to the in-sync replicas, and consumers only see committed records.

for a middle

Distinguish leader-written from committed and explain why acks=1 records can be lost on leader failover.

for a senior

Explain HW = min LEO across ISR, the HW/LEO gap, and how the HW bounds consumer fetches independent of acks.

for a principal

Discuss leader epochs (KIP-101), HW propagation during elections, and read_committed/LSO for transactional consumers as part of the end-to-end consistency model.

## Two log offsets you must know Every partition replica tracks two key offsets: - **Log End Offset (LEO)**: the offset of the next record to be written — i.e. one past the last record in *that replica's* log. The leader's LEO advances as soon as it appends a new record. - **High Watermark (HW)**: the highest offset that has been replicated to **all** ISR members. The leader computes it as the minimum LEO across the current ISR. Records below the HW are **committed**. ## Written-at-leader vs committed 1. **Written at the leader**: the leader appends the record to its log, advancing its LEO. The record now exists on one broker. `acks=1` returns here. The record is **not yet committed** — it sits between the HW and the LEO. 2. **Committed**: followers in the ISR fetch the record and append it, advancing their own LEOs. Once the slowest ISR member has it, the leader advances the HW past the record. `acks=all` waits for this moment. ## Why consumers never see uncommitted data Consumers are only ever served records **below the high watermark**. Even though the leader's log physically contains records up to its LEO, fetch responses cap at the HW. This guarantees that any record a consumer reads is durable across the ISR — you cannot read a record that might later be truncated. This holds **regardless of the producer's acks setting**; acks only affects when the *producer* learns of success, not what consumers can see. ## Why acks=1 records can vanish Under `acks=1` the producer is told 'success' at the written-at-leader stage. If the leader crashes before followers replicate that record: - A follower (which never received the record) can become the new leader. - The old leader, on rejoining, truncates its log to the new leader's HW, **deleting** the un-replicated record. - The producer believed the write succeeded, but it's gone — silent data loss. Under `acks=all`, the producer only got success after the HW advanced, meaning the record was on all ISR members; a new leader is guaranteed to have it. ## Leader epoch and truncation Modern Kafka uses **leader epochs** (KIP-101) instead of naive HW-based truncation to avoid certain divergence/loss scenarios during leader changes. The principle stands: only committed (≤ HW) records are guaranteed to survive leader elections. ## Interplay with transactions / read_committed For transactional producers, consumers with `isolation.level=read_committed` see records only up to the **Last Stable Offset (LSO)** — committed *and* not part of an open transaction. The HW is still the ceiling; the LSO is an additional, lower bound for transactional reads. ## Summary table | Stage | Offset position | acks that returns here | Durable? | |---|---|---|---| | Written at leader | between HW and LEO | acks=1 | No — can be truncated | | Committed | at/below HW | acks=all | Yes — across ISR |

  • Can a consumer ever read a record that the producer sent with acks=1 but that was never replicated?
    No. Consumers only read up to the high watermark, which only advances once all ISR members have the record. An un-replicated acks=1 record sits above the HW and is invisible to consumers until (and unless) it is committed.
  • What is the difference between the high watermark and the log end offset?
    The LEO is one past the last record in a replica's log (the leader's LEO moves on every append). The HW is the minimum LEO across all ISR members — the committed line. The leader's LEO can be ahead of the HW for not-yet-replicated records.

saying these in an interview costs you the question

  • Saying consumers can read records as soon as the leader writes them (they read up to the HW).
  • Conflating LEO and HW — they differ exactly by the un-replicated tail.
  • Claiming acks changes what consumers can see (it changes producer acknowledgement timing, not consumer visibility).
  • Believing acks=1 records are committed.

context