skip to content

How does a read_committed consumer actually avoid delivering records from aborted transactions?

level: seniorimportance: should knowfreq 35%

answer

  1. .txnindex per segment
  2. abortedTransactions list in FetchResponse
  3. filter by producerId + offset range
  4. control (commit/abort) batches never delivered
  5. aborted data deleted only by retention/compaction

basics

~20 s

The broker keeps an aborted-transaction index per segment and sends the relevant aborted (producerId, firstOffset) entries with each fetch. The consumer reads records below the LSO and drops any whose producerId+offset fall inside an aborted transaction's range.

solid answer

~50 s

Aborted records are physically present in the log, so filtering happens cooperatively between broker and consumer. Each log segment has a `.txnindex` file recording aborted transactions as (producerId, firstOffset, lastOffset/lastStableOffset) entries. On a read_committed Fetch, the broker returns the data batches below the LSO plus the list of aborted transactions that overlap the fetched offset range. The consumer-side Fetcher then walks the returned record batches: each batch carries a producerId and an isTransactional/isControl flag. Using the aborted-transaction list, the consumer skips every batch whose producerId matches an aborted transaction while that transaction is open over those offsets, and it also discards the commit/abort control records themselves. What surfaces to the application is exactly the committed transactional records plus any non-transactional records, in offset order. The broker never rewrites the log to remove aborted data; it stays until normal retention or compaction removes it.

go deeper

for a junior

Just know aborted records exist in the log but the consumer hides them when read_committed.

for a middle

Know the broker sends an aborted-transaction list and the consumer skips those records.

for a senior

Explain .txnindex, the abortedTransactions list, producerId-keyed per-batch filtering, and control-record handling.

for a principal

Reason about storage/latency cost, retention interaction, and recovery/rebuild of the transaction index.

## The core fact: aborted records are not deleted When a transaction aborts, Kafka does **not** go back and erase the records it already wrote. They remain in the log at their original offsets. Hiding them from read_committed consumers is therefore a *read-time filtering* problem, solved by metadata the broker maintains and ships to the consumer. ## What the broker stores: the transaction index (.txnindex) Each log segment has a companion **transaction index** file (`.txnindex`). When a transaction aborts, the broker appends an entry capturing: - the **producerId** (PID) of the aborting producer, - the **first offset** of that producer's data in this transaction, - the offset of the abort marker / last relevant offset, - and the LSO at that point. This lets the broker, for any fetched offset range, quickly enumerate which (producerId, offset-range) spans were aborted. ## What the broker sends on a Fetch For a `read_committed` fetch, the `FetchResponse` for each partition includes: 1. The record batches, but only **up to the LSO**. 2. An **`abortedTransactions`** list: entries of `(producerId, firstOffset)` for transactions that aborted within the returned offset range. 3. The current LSO (so the consumer can compute lag and know the read ceiling). ## What the consumer does (Fetcher filtering) The consumer's fetch path (in `CompletedFetch` / the fetch collector) reconstructs which records to deliver: - It maintains a set of **currently-aborted producerIds**, seeded from the `abortedTransactions` list, advancing through the data in offset order. - Each record **batch** header carries the `producerId`, plus flags: `isTransactional` and `isControl`. - **Control batches** (commit and abort markers Kafka writes to mark transaction boundaries) are never delivered to the application; the consumer consumes them only to know when a producer's aborted span ends. - For each transactional data batch, if its producerId is in the aborted set for that offset region, the whole batch is **skipped**. Otherwise the records are delivered. - Non-transactional batches are delivered unconditionally. ## Why per-batch, per-producerId Filtering granularity is the **batch** (record set), keyed by **producerId**, because a transaction's writes from one producer to one partition form contiguous batches tagged with that PID. Multiple producers can interleave on one partition, so the producerId is what disambiguates whose records to drop. ## Lifecycle and cleanup Aborted records and their control markers persist in the log until normal **retention** (time/size) or **log compaction** eventually removes them. The `.txnindex` is maintained alongside segment files and is rebuilt/validated during recovery. The cost of read_committed is thus mostly: extra metadata in fetches, a small consumer-side filtering pass, and storage of data that will never be delivered until retention reclaims it. ## Edge cases - **Cross-segment transactions**: aborted entries are tracked per segment, so a transaction spanning segments contributes entries to each. - **Consumer seeks**: when a read_committed consumer seeks into the middle of a range, the broker still ships the aborted-transaction list overlapping the new fetch position, so filtering remains correct after a seek. - **No double delivery of markers**: control records are filtered universally, independent of isolation level.

  • Are aborted records physically removed from the log when a transaction aborts?
    No. They remain at their offsets and are filtered at read time. They are only removed later by normal retention or log compaction.
  • What is the filtering granularity and key?
    Filtering is per record batch, keyed by producerId. A batch whose producerId belongs to an aborted transaction over that offset range is skipped wholesale.
  • Are commit/abort control records ever delivered to the application?
    No. Control batches are consumed internally to mark transaction boundaries and are never surfaced to the application under either isolation level.

saying these in an interview costs you the question

  • Claiming the broker rewrites or deletes the log on abort.
  • Saying the consumer filters by offset alone without producerId.
  • Thinking control (commit/abort marker) records are delivered to the application.
  • Assuming filtering is per-record rather than per-batch.

context