skip to content

What is the Last Stable Offset (LSO), and how does it gate reads for a read_committed consumer?

level: middleimportance: must knowfreq 55%

answer

  1. LSO = first offset of an open txn
  2. LSO <= HW <= LEO
  3. read_committed capped at LSO
  4. open txn pins LSO -> stall
  5. transaction.timeout.ms bounds the stall

basics

~20 s

The LSO is the offset of the first record belonging to a still-open transaction. A read_committed consumer is never allowed to read past the LSO, so it stops there until that transaction commits or aborts.

solid answer

~40 s

The Last Stable Offset (LSO) is a per-partition broker-tracked offset equal to the offset of the earliest record that is part of an open (not-yet-resolved) transaction. Equivalently, it is the highest offset for which every prior transaction has reached a terminal commit/abort state. A read_committed consumer's fetches are capped at the LSO: the broker will not return any record at or beyond it, because the outcome of records past the LSO is still undecided. As transactions commit or abort, the LSO advances toward the log end offset (high watermark). If a long-running transaction stays open, the LSO is pinned and read_committed consumers stall, accumulating lag, even though newer committed-looking records may physically exist further along. read_uncommitted consumers ignore the LSO entirely and read up to the high watermark.

go deeper

for a junior

Know the LSO is the read ceiling for read_committed consumers and that an open transaction can stall them.

for a middle

Define LSO precisely, give the LSO<=HW<=LEO ordering, and explain the head-of-line stall.

for a senior

Explain recovery via transaction.timeout.ms and how to diagnose LSO-induced lag.

for a principal

Design around LSO latency — short transactions, monitoring, and choosing isolation level per pipeline stage.

## Offsets you need to know first - **Log End Offset (LEO)**: the offset that will be assigned to the next record appended to a partition replica. - **High Watermark (HW)**: the highest offset that has been replicated to all in-sync replicas; ordinary (read_uncommitted) consumers can read up to the HW. - **Last Stable Offset (LSO)**: the boundary used by read_committed consumers. It is the offset of the **first record that belongs to a transaction that has not yet committed or aborted**. Put differently, all transactional data *below* the LSO has a known, final outcome. The ordering relationship is always: `LSO <= HW <= LEO`. ## Why the LSO exists With transactions, records are written to the log *before* the commit/abort marker. A read_committed consumer must never deliver a record whose transaction might still abort. The cleanest way to enforce that is to refuse to read past the point where the outcome becomes uncertain — that point is the LSO. Everything below the LSO is settled (committed records to deliver, aborted records to skip); everything at or above it is potentially still in-flight. ## How gating works in a fetch When a read_committed consumer issues a `Fetch` request, the broker: 1. Caps the returned data at the LSO (never returns records at or beyond it). 2. Attaches the list of aborted transactions in that fetched range so the consumer can filter aborted records out. 3. Returns the LSO in the fetch response so the consumer can compute lag. ## The stall / head-of-line-blocking effect If a producer opens a transaction and writes record at offset 100 but then hangs (slow, GC pause, crash before commit), the LSO is pinned at 100. Every read_committed consumer on that partition stops at offset 100 — it cannot advance even if offsets 101..1000 are committed work from *other* producers, because those sit above the open transaction's first record. This is a head-of-line blocking property: one stuck transaction blocks read_committed progress for the whole partition. ## How the broker recovers Kafka bounds this with **`transaction.timeout.ms`** (producer-side, capped by broker `transaction.max.timeout.ms`). When a transaction exceeds its timeout, the transaction coordinator aborts it, writes an abort marker, and the LSO advances. So a hung producer cannot pin the LSO forever — it is bounded by the timeout, not held indefinitely. ## Practical signal If you see read_committed consumer lag growing while read_uncommitted consumers on the same partition keep up, suspect a long-running or stuck open transaction holding the LSO back.

  • What is the ordering relationship between LSO, high watermark, and log end offset?
    LSO <= HW <= LEO. The LSO sits at or below the high watermark, which sits at or below the log end offset.
  • What stops a hung producer from pinning the LSO forever?
    transaction.timeout.ms (bounded by the broker's transaction.max.timeout.ms). When exceeded, the transaction coordinator aborts the transaction, writes an abort marker, and the LSO advances.

saying these in an interview costs you the question

  • Confusing the LSO with the high watermark — they coincide only when no transaction is open.
  • Claiming read_committed reads up to the high watermark (it reads up to the LSO).
  • Saying a stuck transaction blocks the LSO indefinitely (transaction.timeout.ms bounds it).

context