skip to content

Delivery Semantics and Transactions

At-most-once, at-least-once and exactly-once in Kafka, plus the transactional producer, read_committed consumers, and consume-transform-produce. Interviewers use this area to separate buzzword answers from real understanding.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

explore

questions

page 2 of 2

Explain transaction.timeout.ms — what it controls, how it interacts with transaction.max.timeout.ms, and what happens on expiry.

level: seniorimportance: should knowfreq 40%

basics

~10 s

transaction.timeout.ms is the producer-side max time an open transaction may stay uncommitted before the broker's transaction coordinator proactively aborts it. It defaults to 60000 ms and cannot exceed the broker's transaction.max.timeout.ms.

open as a page

How does the __transaction_state log let a transaction coordinator survive broker failover, and what does it store?

level: seniorimportance: should knowfreq 40%

basics

~20 s

__transaction_state is an internal, log-compacted, replicated topic keyed by transactional.id. It stores each producer's TransactionMetadata (PID, epoch, state, involved partitions). When a coordinator broker fails, a replica becomes the new partition leader, replays the compacted log to rebuild metadata in memory, and resumes as coordinator.

open as a page

How are transaction markers timestamped, and what timestamp do committed records carry under LogAppendTime vs CreateTime?

level: seniorimportance: should knowfreq 25%

basics

~20 s

Markers are control batches the broker stamps with broker append time. Data records keep their producer CreateTime unless the topic uses message.timestamp.type=LogAppendTime, in which case the leader overwrites every record's timestamp with the broker's append time when it appends the batch.

open as a page

You're scaling a hand-rolled consume-transform-produce service across many instances and partitions. How do you assign transactional.id values, and what are the trade-offs between the pre-KIP-447 per-partition model and the post-KIP-447 per-thread model?

level: principalimportance: should knowfreq 25%

basics

~20 s

On modern brokers (KIP-447), bind one transactional.id to each stable processing thread/instance and rely on consumer generation fencing for rebalance safety — far fewer producers. The legacy model required a deterministic transactional.id per input partition so fencing worked, which exploded the producer count at scale.

open as a page

As a platform architect, how would you reason about end-to-end exactly-once across a pipeline that ingests from an external source, processes in Kafka, and lands in an external sink?

level: principalimportance: should knowfreq 35%

basics

~20 s

Treat the pipeline as three hops: source→Kafka, Kafka→Kafka, Kafka→sink. Only the Kafka-internal hop can be true EOS. The source and sink boundaries are at-least-once unless you add idempotency, dedupe keys, or store-local atomicity. End-to-end exactly-once is engineered at the edges, not granted by Kafka.

open as a page

As an architect, how do you choose among at-most-once, at-least-once (with idempotent consumers), and exactly-once for a given pipeline, and what are the trade-offs?

level: principalimportance: should knowfreq 45%

basics

~20 s

Match the guarantee to the cost of loss vs duplicates. At-most-once for cheap, loss-tolerant data; at-least-once + idempotent consumers as the pragmatic default; exactly-once when both loss and duplicates are unacceptable and effects stay within Kafka. EOS adds latency/throughput cost.

open as a page

When building an exactly-once read-process-write pipeline on Kafka, why is isolation.level=read_committed required on the consuming side, and what breaks if you forget it?

level: principalimportance: should knowfreq 30%

basics

~20 s

Exactly-once needs downstream stages to never act on rolled-back data. read_committed ensures the consumer only sees committed records. With the default read_uncommitted, the consumer processes aborted/uncommitted records, so a rolled-back transaction still leaks downstream, breaking exactly-once.

open as a page

How does Kafka make a multi-partition transactional write atomic? Describe the coordinator, control records, and Last Stable Offset.

level: principalimportance: should knowfreq 30%

basics

~20 s

A transaction coordinator (a broker) tracks the transaction in the __transaction_state log. Records are written to each partition immediately, then the coordinator writes commit or abort control records (markers) to every involved partition. read_committed consumers only read up to the Last Stable Offset, so they never see uncommitted or aborted data.

open as a page

In Kafka Streams exactly-once, how did zombie fencing evolve from per-partition transactional.ids to a single producer per instance, and why?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Original EOS (exactly_once) used one transactional.id (and producer) per input partition, so fencing was per-task. EOS v2 (exactly_once_v2, KIP-447) uses one producer per stream thread/instance and fences via consumer group metadata, drastically cutting producers and connections while keeping correct fencing.

open as a page

What is the WriteTxnMarkers request, and why must marker delivery be idempotent and epoch-fenced?

level: principalimportance: nice to knowfreq 18%

basics

~20 s

WriteTxnMarkers is an inter-broker request the coordinator sends to each partition leader telling it to append a COMMIT/ABORT marker for a given PID and producer epoch. It must be idempotent (retries after a coordinator crash must not double-finalize) and epoch-fenced (a stale producer epoch is rejected to block zombies).

open as a page

showing 31–40 of 40