What is a transaction marker (control batch) in Kafka, and who writes it to a partition?
answer
- COMMIT / ABORT control batch
- written by coordinator, not producer
- isControl bit in batch header
- consumers filter it out
- read_committed waits for the marker
basics
~20 sA transaction marker is a special COMMIT or ABORT record the transaction coordinator writes at the end of a transaction into every partition the transaction touched. It tells consumers whether the preceding transactional records are valid.
solid answer
~40 sWhen a producer runs a transaction, it writes data records to one or more partitions. To finish, the transaction coordinator (a broker-side component backed by the __transaction_state log) writes a transaction marker — a control batch — into each partition that received data. The marker is either COMMIT or ABORT. It is not application data; it carries no key/value payload visible to apps, and consumers filter it out. Its job is to signal, at the partition log level, the final outcome of the producer's transaction so that read_committed consumers know whether to deliver or skip the buffered records. The producer never writes markers itself; only the coordinator does, after the producer calls commitTransaction() or abortTransaction().
go deeper
Know that a marker is a COMMIT/ABORT control record, written by the coordinator, that tells consumers if records are valid.
Explain the EndTxn -> WriteTxnMarkers flow and that markers live in data partitions while state lives in __transaction_state.
Discuss the control-batch header bits (isControl/isTransactional), PID/epoch association, and read_committed buffering semantics.
Reason about why markers are necessary given independent partition logs and the absence of cross-partition locks.
## The problem markers solve Kafka lets a producer write to several partitions inside one *transaction* and have all those writes become visible atomically (all-or-nothing). But Kafka partitions are independent append-only logs on possibly different brokers. There is no shared lock across them. So how does a consumer reading one partition know whether the transactional records it sees were ultimately committed or aborted? That decision might be made after the records were already appended. ## What a marker is The answer is the **transaction marker**, also called a **control batch** or **control record**. It is a special record batch appended to the *data* partition's log, distinguished by a flag in the record batch header (the `isControl`/control bit, with `isTransactional` also set). Its body encodes a **ControlRecordType**: `COMMIT` (value 1) or `ABORT` (value 0). It is associated with the producer's `producerId` (PID) and `producerEpoch`. A marker is *not* application data: it has no meaningful key/value for consumers, it does not increase the offset that application records occupy in a user-visible way beyond consuming one log offset, and the consumer's `KafkaConsumer` silently drops control batches so `poll()` never returns them. ## Who writes it The **transaction coordinator** writes markers. The coordinator is a broker that owns the relevant partition of the internal `__transaction_state` topic (selected by hashing the producer's `transactional.id`). The producer does NOT write the marker. The flow is: 1. Producer sends data to partitions (and registers each partition with the coordinator via `AddPartitionsToTxn`). 2. Producer calls `commitTransaction()` (or `abortTransaction()`), sending an `EndTxn` request to the coordinator. 3. The coordinator durably records the decision (PrepareCommit/PrepareAbort) in `__transaction_state`. 4. The coordinator sends `WriteTxnMarkers` requests to every broker leading a partition the transaction touched; each leader appends the COMMIT or ABORT marker to that partition's log. 5. After all markers are written, the coordinator records CompleteCommit/CompleteAbort. ## Why consumers need it A consumer with `isolation.level=read_committed` buffers transactional records until it sees the matching marker. On COMMIT it delivers them; on ABORT it discards them. A `read_uncommitted` consumer ignores markers and delivers everything. Thus the marker is the single point of truth in the data log for a transaction's outcome.
- Does the producer write the transaction marker?No. The producer signals the end of the transaction via an EndTxn request; the transaction coordinator writes the markers into the data partitions using WriteTxnMarkers requests to the partition leaders.
- What are the two marker types and what do they tell a consumer?COMMIT and ABORT. A read_committed consumer delivers the buffered transactional records on COMMIT and discards them on ABORT.
saying these in an interview costs you the question
- Saying the producer writes the marker itself.
- Claiming markers are normal records visible to application consumers.
- Confusing the marker (in the data partition) with the transaction state stored in __transaction_state.
- Thinking read_uncommitted consumers also wait for markers.