Explain the append-only log structure of a partition and the role of offsets, log-start offset, and the high watermark.
answer
- Append-only segments, immutable records
- Offset = position assigned by leader, monotonic
- Log-start offset advances with retention/compaction
- HW = max offset replicated to all ISR
- Consumers read only below HW; lag = HW − position
basics
~20 sEach partition is an append-only log: records are added at the tail, each getting a monotonically increasing offset. The log-start offset is the earliest still-retained offset; the high watermark is the highest offset replicated to all in-sync replicas — consumers can only read up to it.
solid answer
~50 sA partition is physically an append-only commit log split into segment files. Producers append at the **log end offset (LEO)**, and each record receives the next monotonically increasing offset. Records are immutable once written. Two boundary markers matter. The **log-start offset** is the oldest offset still on disk — it advances as retention (time/size) or compaction deletes old data, so offsets below it are gone (seeking there throws/auto-resets). The **high watermark (HW)** is the highest offset that has been replicated to **all** in-sync replicas (ISR); the leader exposes only records below the HW to consumers, so a consumer never reads data that isn't yet durably replicated. The gap between HW and LEO is data written to the leader but not yet acknowledged by all ISR members. On `acks=all`, a producer's record is considered committed when the HW advances past it. Consumers track their own position (next offset to fetch) per partition and commit it; lag = HW − committed offset.
go deeper
Know a partition is append-only and each record gets an increasing offset.
Distinguish log-start offset from offset 0 and know retention removes old records.
Explain HW vs LEO, why consumers read only below HW, and how lag and acks=all relate to the HW.
Reason about HW-based durability, follower truncation on leader change, and retention/compaction effects when designing durability and replay guarantees.
## The log A partition is a **commit log**: an ordered, append-only sequence of records persisted to disk. Physically it is a series of **segment files** (e.g. `00000000000000000000.log` plus `.index` files); when the active segment hits `segment.bytes` or `segment.ms`, Kafka rolls a new one. Reads are sequential, which is why Kafka is fast — it leans on the OS page cache and sequential I/O. ## Offsets Every appended record gets the next **offset**: a 64-bit, monotonically increasing, gap-free-on-append integer that is its position in *this* partition. Offsets are assigned by the **leader** at append time, are immutable, and never reused. They are per-partition (partition 0's offset 100 is unrelated to partition 1's offset 100). ## Key offset markers - **Log End Offset (LEO):** the offset that will be assigned to the *next* record appended — i.e., one past the last written record. The leader's LEO is the head of the log. - **Log-Start Offset (a.k.a. earliest):** the lowest offset still retained. It is **not** always 0, because retention (`retention.ms`, `retention.bytes`) and log compaction delete old records, advancing the start. Trying to read below it fails or triggers `auto.offset.reset` (`earliest`/`latest`/`none`). - **High Watermark (HW):** the highest offset that has been successfully replicated to **every replica in the ISR** (in-sync replica set). It is the durability boundary. ## Why the high watermark gates reads Consumers are only allowed to read records **strictly below the HW**. Records between HW and LEO exist on the leader but have not yet been confirmed by all in-sync followers, so they are not guaranteed durable. If the leader crashed, those un-replicated records could be lost; exposing them would let a consumer see data that later disappears. By gating at the HW, Kafka guarantees a consumer never reads a record that isn't already replicated across the ISR. ## Commit semantics With `acks=all`, a produced record is **committed** once the HW moves past it (all ISR members have it). The producer's ack is tied to this advancement. With `acks=1`, only the leader must persist it, so the HW concept still governs consumer visibility but durability is weaker. ## Consumer position and lag Each consumer tracks, per partition, the **next offset to fetch** (its position) and periodically **commits** it (to `__consumer_offsets`). **Consumer lag** for a partition = HW − last committed/consumed offset: how far behind the consumer is. Lag is monitored per TopicPartition. ## Edge cases - **Compacted topics:** offsets become sparse — compaction removes superseded records by key, so the log has gaps, but offsets remain monotonic and meaningful. - **Leader change / truncation:** a new leader may truncate the follower logs down to the HW to keep replicas consistent; un-replicated records above the old HW can be discarded. This is why only HW-and-below is considered safe. - **Retention vs compaction:** retention deletes whole old segments (advancing log-start offset); compaction keeps the latest value per key. ### Mental model ``` Partition log: offset: ... 150 198 201 [retained...][HW=198][....LEO=201] ^log-start ^all ISR have ^leader head; (old gone) up to 197 198-200 not yet fully replicated Consumers may read up to offset 197 (below HW). ```
- Why are consumers prevented from reading records between the high watermark and the log end offset?Those records are on the leader but not yet replicated to all in-sync replicas, so they aren't durably committed. If the leader failed, they could be truncated/lost. Gating reads at the HW ensures a consumer never sees a record that might later disappear.
- Why might the log-start offset of a partition be greater than 0?Retention (time/size) and log compaction delete old records, advancing the earliest retained offset. After old segments are removed, offset 0 no longer exists on disk, so the log-start offset moves forward.
saying these in an interview costs you the question
- Assuming the log always starts at offset 0
- Saying consumers can read up to the LEO (they can only reach the HW)
- Confusing high watermark with consumer committed offset
- Believing records can be edited/inserted in place
- Thinking offsets are reused or reset after retention deletes data