skip to content

What is the .txnindex file, what does it record, and how do consumers use it to honor read_committed isolation?

level: principalimportance: nice to knowfreq 18%

answer

  1. .txnindex = aborted txn ranges (producerId, first/last offset)
  2. LSO caps read_committed (vs high watermark for read_uncommitted)
  3. records interleaved; commit/abort decided later via control batch
  4. broker sends aborted list; consumer filters client-side
  5. open txn stalls LSO -> consumer lag

basics

~20 s

The .txnindex is a per-segment file listing aborted transactions as ranges (producer id, first offset, last stable offset). With read_committed isolation, the broker sends this list so consumers can filter out records from aborted transactions and not read past the last stable offset.

solid answer

~50 s

Each segment may have a .txnindex (transaction index) recording AbortedTxn entries: producerId, firstOffset of the aborted transaction, lastOffset, and the last stable offset (LSO) context. It exists because Kafka's transactions interleave records from concurrent producers in the log, and control batches (COMMIT/ABORT markers) are appended out-of-band. For a read_committed consumer, the broker must (1) never expose records beyond the LSO — the offset before the first still-open transaction — and (2) tell the consumer which already-written records belong to aborted transactions so it can drop them. On a fetch, the broker reads the relevant .txnindex entries and returns the list of aborted transactions alongside the data; the consumer uses producerId + offset ranges plus the ABORT control markers to filter aborted records client-side. read_uncommitted consumers ignore all of this and read up to the high watermark. The index is sparse only in that it has one entry per aborted transaction, not per record.

go deeper

for a junior

Awareness only: a .txnindex tracks aborted transactions for transactional topics.

for a middle

Know read_committed must skip aborted records and stop at the LSO.

for a senior

Explain the LSO vs high watermark and broker-sends-list / consumer-filters split.

for a principal

Reason about EOS internals, control batches, hung-transaction lag, and recovery of the index.

## Why a transaction index exists Kafka's idempotent/transactional producers (EOS — exactly-once semantics) let a producer write to multiple partitions atomically. But the log is a single append-only stream: records from different transactions and different producers are **interleaved by arrival**, and the transaction outcome (commit or abort) is decided **later** by the transaction coordinator, which appends a **control batch** (a COMMIT or ABORT marker) into each involved partition. So at the time a record is written, the broker does not yet know if it will be committed or aborted. ## Two concepts the broker must enforce for read_committed 1. **Last Stable Offset (LSO):** the offset of the first record belonging to a transaction that is **still open** (not yet committed/aborted). A `read_committed` consumer must never be served records **at or beyond the LSO**, because their fate is undecided. (Contrast: `read_uncommitted` reads up to the **high watermark**.) 2. **Aborted transactions:** records that were written but later got an ABORT marker. They are physically in the log (Kafka doesn't rewrite the log) but must be **filtered out** for read_committed consumers. ## What the .txnindex stores When a transaction aborts, the broker appends an entry to the segment's `.txnindex`. Each `AbortedTxn` entry holds: the **producerId**, the **firstOffset** of the aborted transaction's data in this segment, the **lastOffset**, and LSO bookkeeping. One entry per aborted transaction (not per record) — so it's tiny. ## The fetch path On a `read_committed` fetch, the broker: - caps returned data at the **LSO**; - reads the `.txnindex` for the fetched range and returns the matching list of **aborted transactions** (producerId + offset range) in the fetch response. The **consumer** then walks the returned record batches and, using the aborted-transaction list together with the ABORT **control markers** it sees in the stream, **discards** records that belong to aborted transactions (matched by producerId and offset range). Committed records and control markers are filtered/handled appropriately; the application only sees committed data. ## Edge cases & implications - **read_uncommitted** consumers receive everything up to the high watermark and ignore the `.txnindex` entirely. - A long-running **open** transaction stalls the LSO, so read_committed consumers can lag even though data is physically present (a real operational gotcha — hung transactions block consumers). - The `.txnindex` only exists for segments that actually contain aborted transactions; non-transactional topics never create one. - Like other index files, it is a per-segment companion and is rebuilt/validated during recovery. - Control markers themselves are not delivered to the application; they are an internal mechanism.

  • What is the Last Stable Offset and how does it differ from the high watermark for read_committed consumers?
    The LSO is the offset of the first record in a still-open transaction; read_committed consumers can only read up to it. The high watermark is the highest fully-replicated offset; read_uncommitted consumers read up to it. LSO <= high watermark.
  • Why does an aborted transaction's data still physically exist in the log rather than being removed?
    Kafka's log is append-only and immutable; it never rewrites it. Aborted records remain on disk and are filtered out at read time for read_committed consumers using the .txnindex and ABORT control markers.
  • Operationally, what happens to read_committed consumers if a transaction stays open for a long time?
    The LSO is pinned at that open transaction's first offset, so read_committed consumers cannot advance past it and appear to lag, even though newer records are physically present and replicated.

saying these in an interview costs you the question

  • Saying aborted records are deleted/rewritten from the log (they stay; they're filtered at read time).
  • Claiming the broker filters aborted records (it sends the list; the consumer filters them).
  • Confusing LSO with the high watermark.
  • Saying read_uncommitted consumers also use the .txnindex (they don't).

context