How do transaction markers, the Last Stable Offset (LSO), and read_committed consumers interact?
answer
- LSO = first offset of earliest open txn
- read_committed fetch capped at LSO
- marker advances the LSO
- aborted-transaction index (.txnindex) -> client filters
- stuck txn pins LSO, stalls consumers
basics
~20 sA read_committed consumer can only read up to the Last Stable Offset (LSO) — the offset before the earliest still-open transaction. Markers let the LSO advance: once a COMMIT/ABORT marker lands, the broker knows the outcome and exposes those records (delivering committed, skipping aborted).
solid answer
~40 sFor read_committed, the broker never serves records beyond the **Last Stable Offset (LSO)** — the smallest offset belonging to any still-open (uncommitted) transaction. Records past the LSO are withheld because their outcome is unknown. When the coordinator writes a transaction's COMMIT or ABORT **marker** to a partition, that transaction is no longer open there, so the LSO can advance past its records. The broker also maintains a per-partition **aborted transactions index**; on fetch it returns the data plus the list of aborted PID ranges, and the consumer's Fetcher uses that index to drop records from aborted transactions while delivering committed ones. So markers both unblock the LSO and drive the deliver-vs-skip decision. A long-running or stuck transaction pins the LSO, stalling all read_committed consumers on that partition — a key operational hazard.
go deeper
Know read_committed only shows committed records and waits for the marker.
Define the LSO and explain that a marker advances it and unblocks read_committed consumers.
Explain the aborted-transaction index and client-side filtering, and the hanging-transaction LSO hazard.
Reason about latency/availability trade-offs and capacity impact of retained aborted data.
## Three concepts **read_committed**: a consumer `isolation.level` where only records from *committed* transactions (plus non-transactional records) are delivered; aborted records are hidden. The alternative `read_uncommitted` delivers everything immediately and ignores markers. **Last Stable Offset (LSO)**: a per-partition broker watermark equal to the offset of the *first record of the earliest still-open transaction* (or the high watermark if no transaction is open). It is always <= the high watermark. For read_committed fetches, the broker caps the returned data at the LSO — anything at or after the LSO might still be aborted, so it is not yet 'stable'. **Transaction marker**: the COMMIT/ABORT control batch written by the coordinator that finalizes a transaction's outcome in the partition. ## How they interact 1. A producer begins a transaction and appends records at offsets, say, 100-110. The transaction is open. The LSO is pinned at 100; a read_committed consumer cannot see offsets >= 100 yet, even though the high watermark may be 111. 2. The producer commits. The coordinator writes a **COMMIT marker** at, say, offset 111. Now no open transaction starts at 100, so the **LSO advances** past 111. The read_committed consumer can now fetch 100-110. 3. If instead the producer aborts, an **ABORT marker** is written at 111. The LSO advances the same way, but the broker records the range as aborted in the **aborted-transaction index**. On a read_committed fetch, the broker returns the byte range *and* the aborted PID/offset ranges; the consumer's `Fetcher`/`CompletedFetch` logic discards records whose (PID, offset) fall in an aborted range, delivering nothing for them. ## Mechanics of the abort filtering Aborted records are *not* deleted from the log — they remain physically present until retention/compaction. The broker keeps a `.txnindex` file per segment listing aborted transactions (PID, first offset, last stable offset). The fetch response carries `AbortedTransaction` entries; the consumer filters client-side. This is why aborted data still costs disk and still consumes offsets — markers and aborted records advance offsets even though apps see nothing. ## Operational consequences - A **stuck/hanging transaction** keeps its first offset as the LSO, so read_committed consumers freeze on that partition until the transaction commits, aborts, or times out (`transaction.timeout.ms` on the producer, capped by broker `transaction.max.timeout.ms`). The coordinator eventually aborts timed-out transactions, writing ABORT markers to release the LSO. - **Latency**: read_committed adds end-to-end latency equal to roughly the transaction duration, because nothing is visible until the marker lands. - **read_uncommitted** ignores all of this: it reads up to the high watermark and never filters, so it can see records that later abort. ## Summary Markers are the events that move a partition's LSO forward and tell the broker/consumer which buffered records to deliver or skip. Without a marker, the transaction stays open and its records stay invisible to read_committed consumers.
- Are aborted records physically removed when an ABORT marker is written?No. They stay in the log (subject to normal retention/compaction). The broker records them in an aborted-transaction index and the consumer filters them out client-side on fetch.
- What happens to read_committed consumers if a producer's transaction hangs?The LSO is pinned at that transaction's first offset, so those consumers cannot advance on the partition until the transaction commits, aborts, or is aborted by the coordinator after transaction.max.timeout.ms.
saying these in an interview costs you the question
- Saying read_committed reads up to the high watermark (it reads up to the LSO).
- Claiming aborted records are deleted from the log immediately.
- Thinking the broker filters aborted records server-side (the consumer does it using the returned aborted ranges).
- Confusing LSO with the consumer's committed offset.