Explain transaction.timeout.ms — what it controls, how it interacts with transaction.max.timeout.ms, and what happens on expiry.
answer
- default 60000 ms, open-transaction cap
- broker ceiling transaction.max.timeout.ms (900000)
- expiry → abort markers + epoch bump → fenced
- protects Last Stable Offset / read_committed
- not delivery.timeout.ms
basics
~10 stransaction.timeout.ms is the producer-side max time an open transaction may stay uncommitted before the broker's transaction coordinator proactively aborts it. It defaults to 60000 ms and cannot exceed the broker's transaction.max.timeout.ms.
solid answer
~40 stransaction.timeout.ms (producer config, default 60000) bounds how long a single transaction may remain open before the transaction coordinator force-aborts it, preventing a stuck producer from blocking read_committed consumers indefinitely (an open transaction holds the Last Stable Offset and stalls those consumers). The producer sends this value to the coordinator; the broker rejects it if it exceeds transaction.max.timeout.ms (broker config, default 900000), throwing during initTransactions. On expiry the coordinator writes abort markers, bumps the producer epoch, and the original producer gets a ProducerFencedException or InvalidProducerEpochException on its next operation — effectively fenced. You tune it above your worst-case process-batch time in consume-process-produce loops, but not so high that a hung producer can block consumers for a long time. It is unrelated to delivery.timeout.ms, which bounds individual send completion.
go deeper
Know it is the max time a transaction can stay open before the broker aborts it (default 60s).
Explain the broker ceiling transaction.max.timeout.ms and that expiry aborts the transaction.
Connect it to the Last Stable Offset, read_committed stalls, epoch bump, and tuning vs batch time.
Reason about blast-radius tradeoffs, interaction with downstream I/O latency, and operational defaults across a fleet.
## What it is `transaction.timeout.ms` is a **producer** configuration (default **60000** = 60s). It tells the **transaction coordinator** (the broker managing transaction state) the maximum time this producer's transaction may stay **open** (begun but not yet committed or aborted) before the coordinator proactively aborts it. ## Why it exists An open transaction blocks `read_committed` consumers. Such consumers can only read up to the **Last Stable Offset (LSO)** — the offset before the earliest still-open transaction. If a producer begins a transaction and then hangs (GC pause, deadlock, network partition, lost process), without a timeout the LSO would never advance and read_committed consumers on those partitions would stall forever. The timeout caps that blast radius: the coordinator eventually aborts the zombie transaction so the LSO can move. ## Interaction with transaction.max.timeout.ms `transaction.max.timeout.ms` is a **broker** config (default **900000** = 15 min) that sets an upper bound on any client's requested timeout. When the producer calls `initTransactions()`, it advertises its `transaction.timeout.ms`. If that exceeds the broker's max, the broker rejects it and the producer fails with an error (e.g. `InvalidTxnTimeoutException`). So the effective ceiling is `min(client transaction.timeout.ms, broker transaction.max.timeout.ms)` — and a client request above the broker max simply fails rather than being silently clamped. ## What happens on expiry When the timer fires, the coordinator: 1. Transitions the transaction to a pending-abort state and writes **abort markers** to all partitions that were added to the transaction. 2. **Bumps the producer epoch**, fencing the original producer. 3. The next API call from the original producer (send/commit) returns `ProducerFencedException` (or `InvalidProducerEpochException` on newer versions) — fatal; the producer must be closed and recreated. The records written before expiry are aborted and never become visible to read_committed consumers. ## Tuning guidance - Set it comfortably **above** the worst-case time to process one batch in a consume-process-produce loop, accounting for slow downstreams and GC. - Don't set it gratuitously high: a genuinely hung producer would then block read_committed consumers that whole time. - It is **distinct** from `delivery.timeout.ms` (per-send delivery deadline) and from consumer `max.poll.interval.ms` (which governs consumer group liveness, not transactions). ## Common pitfall Doing slow external I/O (a database call, an HTTP request) **inside** an open transaction can blow past the timeout, getting your producer fenced mid-batch. Keep transactions short, or raise the timeout deliberately and within the broker max.
- What is the Last Stable Offset and how does transaction.timeout.ms relate to it?The LSO is the highest offset before the earliest still-open transaction; read_committed consumers can't read past it. A hung open transaction pins the LSO; transaction.timeout.ms bounds how long that stall can last before the coordinator aborts and lets the LSO advance.
- What error does the original producer see after its transaction times out?Because the coordinator bumps the producer epoch on abort, the next call returns ProducerFencedException (or InvalidProducerEpochException) — a fatal error requiring the producer to be closed.
saying these in an interview costs you the question
- Confusing transaction.timeout.ms with delivery.timeout.ms or max.poll.interval.ms.
- Thinking a client request above the broker max is silently clamped — it is rejected.
- Assuming expiry just commits or just discards quietly without fencing the producer.
- Doing long blocking external I/O inside an open transaction without raising the timeout.