Delivery Semantics and Transactions
At-most-once, at-least-once and exactly-once in Kafka, plus the transactional producer, read_committed consumers, and consume-transform-produce. Interviewers use this area to separate buzzword answers from real understanding.
part ofApache Kafkaoverview, primer and where to startread it →on this pageshowhide
explore
- At-Most/At-Least/Exactly-Once Defined5 questions
- Transactional Producer API5 questions
- Transaction Coordinator and Zombie Fencing5 questions
- Transaction Markers and Atomic Commit5 questions
- Read-Committed Isolation and LSO5 questions
- Consume-Transform-Produce EOS5 questions
- Idempotent Consumers and Dedup5 questions
- EOS Scope and Limits5 questions
questions
page 2 of 2Explain transaction.timeout.ms — what it controls, how it interacts with transaction.max.timeout.ms, and what happens on expiry.
basics
~10 stransaction.timeout.ms is the producer-side max time an open transaction may stay uncommitted before the broker's transaction coordinator proactively aborts it. It defaults to 60000 ms and cannot exceed the broker's transaction.max.timeout.ms.
How does the __transaction_state log let a transaction coordinator survive broker failover, and what does it store?
basics
~20 s__transaction_state is an internal, log-compacted, replicated topic keyed by transactional.id. It stores each producer's TransactionMetadata (PID, epoch, state, involved partitions). When a coordinator broker fails, a replica becomes the new partition leader, replays the compacted log to rebuild metadata in memory, and resumes as coordinator.
How are transaction markers timestamped, and what timestamp do committed records carry under LogAppendTime vs CreateTime?
basics
~20 sMarkers are control batches the broker stamps with broker append time. Data records keep their producer CreateTime unless the topic uses message.timestamp.type=LogAppendTime, in which case the leader overwrites every record's timestamp with the broker's append time when it appends the batch.
You're scaling a hand-rolled consume-transform-produce service across many instances and partitions. How do you assign transactional.id values, and what are the trade-offs between the pre-KIP-447 per-partition model and the post-KIP-447 per-thread model?
basics
~20 sOn modern brokers (KIP-447), bind one transactional.id to each stable processing thread/instance and rely on consumer generation fencing for rebalance safety — far fewer producers. The legacy model required a deterministic transactional.id per input partition so fencing worked, which exploded the producer count at scale.
As a platform architect, how would you reason about end-to-end exactly-once across a pipeline that ingests from an external source, processes in Kafka, and lands in an external sink?
basics
~20 sTreat the pipeline as three hops: source→Kafka, Kafka→Kafka, Kafka→sink. Only the Kafka-internal hop can be true EOS. The source and sink boundaries are at-least-once unless you add idempotency, dedupe keys, or store-local atomicity. End-to-end exactly-once is engineered at the edges, not granted by Kafka.
As an architect, how do you choose among at-most-once, at-least-once (with idempotent consumers), and exactly-once for a given pipeline, and what are the trade-offs?
basics
~20 sMatch the guarantee to the cost of loss vs duplicates. At-most-once for cheap, loss-tolerant data; at-least-once + idempotent consumers as the pragmatic default; exactly-once when both loss and duplicates are unacceptable and effects stay within Kafka. EOS adds latency/throughput cost.
When building an exactly-once read-process-write pipeline on Kafka, why is isolation.level=read_committed required on the consuming side, and what breaks if you forget it?
basics
~20 sExactly-once needs downstream stages to never act on rolled-back data. read_committed ensures the consumer only sees committed records. With the default read_uncommitted, the consumer processes aborted/uncommitted records, so a rolled-back transaction still leaks downstream, breaking exactly-once.
How does Kafka make a multi-partition transactional write atomic? Describe the coordinator, control records, and Last Stable Offset.
basics
~20 sA transaction coordinator (a broker) tracks the transaction in the __transaction_state log. Records are written to each partition immediately, then the coordinator writes commit or abort control records (markers) to every involved partition. read_committed consumers only read up to the Last Stable Offset, so they never see uncommitted or aborted data.
In Kafka Streams exactly-once, how did zombie fencing evolve from per-partition transactional.ids to a single producer per instance, and why?
basics
~20 sOriginal EOS (exactly_once) used one transactional.id (and producer) per input partition, so fencing was per-task. EOS v2 (exactly_once_v2, KIP-447) uses one producer per stream thread/instance and fences via consumer group metadata, drastically cutting producers and connections while keeping correct fencing.
What is the WriteTxnMarkers request, and why must marker delivery be idempotent and epoch-fenced?
basics
~20 sWriteTxnMarkers is an inter-broker request the coordinator sends to each partition leader telling it to append a COMMIT/ABORT marker for a given PID and producer epoch. It must be idempotent (retries after a coordinator crash must not double-finalize) and epoch-fenced (a stale producer epoch is rejected to block zombies).
showing 31–40 of 40