skip to content

Which broker and producer configs govern the transaction coordinator and zombie fencing, and what do they control?

level: middleimportance: should knowfreq 35%

answer

  1. transactional.id = the fencing switch (implies idempotence)
  2. transaction.timeout.ms vs broker transaction.max.timeout.ms cap
  3. state.log.num.partitions (50) = coordinator sharding, immutable
  4. state.log.replication.factor / min.isr = durability (RF 3 prod)
  5. transactional.id.expiration.ms = idle cleanup

basics

~10 s

Producer side: transactional.id (enables fencing) and transaction.timeout.ms. Broker side: transaction.state.log.num.partitions (coordinator sharding), transaction.state.log.replication.factor and min.isr (durability), and transaction.max.timeout.ms (cap on producer timeout).

solid answer

~40 s

On the producer, setting transactional.id is what activates transactions and cross-instance fencing; it also implicitly enables idempotence. transaction.timeout.ms tells the coordinator how long to wait before aborting an idle in-flight transaction. On the broker, transaction.state.log.num.partitions (default 50) defines how transactional.ids shard across coordinators and is effectively immutable once data exists, since it drives the id->partition hash. transaction.state.log.replication.factor (production default 3) and transaction.state.log.min.isr protect coordinator state durability so failover never loses metadata. transaction.max.timeout.ms (default 15 min) caps any producer-requested timeout to prevent a stuck transaction from holding the Last Stable Offset and blocking read_committed consumers indefinitely. transactional.id.expiration.ms controls when an idle id's metadata is expired and eligible for compaction cleanup.

go deeper

for a junior

Know transactional.id turns on transactions/fencing and that there are timeout and replication settings.

for a middle

Map each config to what it controls: sharding, durability, timeout caps, and id expiration.

for a senior

Relate timeout caps to the LSO/read_committed blocking and durability settings to fencing recoverability.

for a principal

Advise on cluster-wide defaults (RF/min.isr), immutability of partition count, and tuning timeouts against EOS processing windows.

## Producer-side configs - **`transactional.id`** (string): the single switch that turns a producer into a *transactional* one. Setting it (a) enables transactions, (b) implicitly forces `enable.idempotence=true`, and (c) is the key used for coordinator assignment and **zombie fencing**. Two producers with the same id are fenced by epoch; two with different ids are independent. - **`transaction.timeout.ms`** (default 60000): the maximum time the coordinator allows a single transaction to remain open before it proactively **aborts** it. Prevents a hung producer from pinning the **Last Stable Offset (LSO)**. - **`enable.idempotence`**: implied `true` under a transactional id; gives per-partition dedup via (PID, epoch, sequence). ## Broker-side configs - **`transaction.state.log.num.partitions`** (default 50): number of partitions of `__transaction_state`, i.e., how many coordinator shards exist and how `transactional.id` hashes to a coordinator. **Effectively immutable** once transactions exist — changing it remaps ids and breaks recovery. - **`transaction.state.log.replication.factor`** (production default 3): replication of coordinator state; higher = more durable failover. - **`transaction.state.log.min.isr`** (default 2 with RF 3): minimum in-sync replicas required to accept writes to `__transaction_state`, protecting against data loss on failover. - **`transaction.max.timeout.ms`** (default 900000 / 15 min): a hard cap; if a producer requests `transaction.timeout.ms` larger than this, the coordinator rejects it. Bounds how long a transaction can block `read_committed` consumers via the LSO. - **`transactional.id.expiration.ms`** (default 7 days): after an id is idle this long, its metadata is expired so compaction can clean it up; the next use re-inits fresh. ## Why these matter for fencing and availability - **Durability (RF/min.isr):** if coordinator state is lost, fencing identity (PID/epoch) and in-flight commit intent are lost — so production must use RF 3. - **Sharding (num.partitions):** balances coordinator load across brokers; too few concentrates load, but it can't be changed later without disruption. - **Timeouts (max/timeout):** the LSO blocks `read_committed` reads past the oldest open transaction; an unbounded timeout could stall consumers, so `transaction.max.timeout.ms` is a safety cap. ## Edge cases - Dev single-broker defaults sometimes set RF/min.isr to 1, which is unsafe for production fencing durability — must be raised before going live. - Setting `transaction.timeout.ms` near or above the consumer's `max.poll.interval.ms` in EOS apps risks aborts during long processing; tune together.

  • Why is transaction.max.timeout.ms a broker-side cap on the producer's transaction.timeout.ms?
    An open transaction holds the Last Stable Offset, blocking read_committed consumers from reading past it. The broker caps the timeout so no single producer can stall consumers indefinitely.
  • What goes wrong if transaction.state.log.replication.factor is 1 in production?
    Losing that single broker loses coordinator state (PID/epoch and commit intent), breaking fencing durability and potentially leaving transactions unrecoverable.

saying these in an interview costs you the question

  • Treating transaction.state.log.num.partitions as freely changeable at runtime
  • Forgetting that setting transactional.id implicitly enables idempotence
  • Confusing producer transaction.timeout.ms with broker transaction.max.timeout.ms
  • Leaving replication factor at the dev default of 1 in production

context