How does commit.interval.ms behave under EOS, and how do you tune it against the latency/throughput tradeoff?
answer
- EOS default 100 ms; ALO default 30000 ms
- commit = transaction = visibility + offset advance
- smaller = lower latency, more overhead/markers
- larger = throughput, more replay, LSO lag
- must stay under transaction.timeout.ms (60s)
basics
~20 sUnder EOS, commit.interval.ms defaults to 100 ms (vs 30000 ms at-least-once) because each commit ends a transaction, and downstream read_committed consumers only see output after a commit. Smaller interval = lower end-to-end latency but more commit overhead; larger interval = higher throughput but more latency and bigger replay on failure.
solid answer
~50 sIn EOS, a commit boundary is a transaction boundary: output and changelog writes become visible to read_committed consumers only when the transaction commits, and offsets advance only then. So commit.interval.ms directly controls end-to-end latency. Streams therefore defaults it to 100 ms under EOS (vs 30000 ms for at-least-once). Tuning: lowering it cuts latency for downstream consumers but raises overhead — more transaction markers, more __consumer_offsets writes, more fsyncs — hurting throughput. Raising it batches more work per transaction, improving throughput and amortizing commit cost, but increases visibility latency and the amount of work replayed on an abort/crash. Constraints: each transaction must finish within transaction.timeout.ms (default 60s), so commit.interval.ms plus worst-case processing must stay under it. Also watch read_committed LSO lag: consumers can't advance past the last stable offset, so long transactions stall downstream lag metrics.
go deeper
Know EOS lowers commit.interval.ms to 100 ms because commits gate when downstream sees output.
Explain the latency-vs-throughput tradeoff and that a commit is a transaction boundary affecting visibility and offsets.
Tie tuning to transaction.timeout.ms, commit overhead (markers/offsets writes), and LSO-driven downstream lag.
Frame an end-to-end budget across latency SLOs, broker commit load, replay cost, and durability config; advise org-wide defaults.
## Why commit interval matters more under EOS In **at-least-once**, `commit.interval.ms` defaults to **30000 ms** — committing offsets occasionally is fine because output is visible immediately as it's produced. In **EOS**, a commit is a **transaction commit**, and two things hinge on it: 1. **Visibility**: output and changelog records are written during the transaction but a downstream `read_committed` consumer can't see them until the **commit marker** lands. So the commit interval is (roughly) the floor of added end-to-end latency for exactly-once consumers. 2. **Offset advancement**: input offsets only advance at commit. A crash mid-transaction replays everything since the last commit. Because of (1), Streams lowers the EOS default to **100 ms** so latency stays reasonable. ## The tradeoff dials **Lowering commit.interval.ms (e.g. 100 → 30 ms):** - Pros: lower downstream latency; less work replayed on abort. - Cons: more transactions/sec → more commit markers per touched partition, more `__consumer_offsets` writes, more producer flushes/fsyncs, more coordinator round-trips → **lower throughput** and higher broker load. At very small intervals, fixed per-commit cost dominates. **Raising commit.interval.ms (e.g. 100 → 1000+ ms):** - Pros: more records batched per transaction, commit cost amortized → **higher throughput**; fewer markers. - Cons: higher visibility latency for read_committed consumers; more work to reprocess on a crash; bigger in-flight transaction (more memory / larger marker fan-out); risk of approaching `transaction.timeout.ms`. ## Hard constraints - **transaction.timeout.ms** (producer, default 60000 ms; broker bounds it via `transaction.max.timeout.ms`, default 900000): a transaction must commit within this window or the coordinator aborts it. So `commit.interval.ms` + worst-case processing burst per interval must stay safely under it. Streams validates/caps this relationship. - **Last Stable Offset (LSO) lag**: `read_committed` consumers only read up to the LSO — the offset before the earliest still-open transaction. A long-running transaction holds the LSO back, inflating downstream consumer-lag metrics even though nothing is wrong. Long commit intervals make this worse. ## Other forces that trigger commits A commit isn't only time-driven. Streams also commits on **rebalance** (to release state cleanly), on **flush/close**, and can be influenced by buffered-records pressure. So real commit frequency may exceed the configured interval under load or churn. ## Practical tuning approach 1. Start with the EOS default (100 ms). Measure downstream p99 latency and throughput. 2. If throughput-bound and latency budget allows, raise toward 200–1000 ms; verify you stay well under transaction.timeout.ms and watch LSO/consumer lag. 3. If latency-bound, lower toward ~30–50 ms only if brokers can absorb the commit overhead. 4. Keep `transaction.timeout.ms` comfortably above `commit.interval.ms` + max processing time; raise it (and broker `transaction.max.timeout.ms`) if you intentionally use long intervals. 5. Ensure `acks=all` and adequate `min.insync.replicas` so commits are durable — EOS sets idempotence/acks=all automatically. ## Common mistake Assuming the at-least-once 30s default still applies under EOS, then being surprised by either huge downstream latency (if it did) or, in reality, by the much chattier 100 ms default's broker load. Always reason from ‘commit = visibility + offset advance’.
- Why is the EOS default commit interval 100 ms instead of the at-least-once 30000 ms?Because under EOS a commit is a transaction commit that gates output visibility for read_committed consumers and offset advancement; 30s would impose ~30s of downstream latency, so Streams lowers it to 100 ms.
- What can go wrong if you set a very large commit.interval.ms with a default transaction.timeout.ms?The transaction may not commit within transaction.timeout.ms (60s) and the coordinator aborts it, causing repeated reprocessing; you must raise the timeout (and broker max) to match.
- Why might downstream consumer lag look high even though processing is healthy under a long commit interval?read_committed consumers can only read up to the Last Stable Offset; a long open transaction holds the LSO back, inflating reported lag until the transaction commits.
saying these in an interview costs you the question
- Stating the EOS default commit interval is 30000 ms (that's at-least-once).
- Claiming smaller commit intervals always improve throughput.
- Ignoring transaction.timeout.ms as a ceiling on the commit interval.
- Not knowing LSO/read_committed causes apparent lag with long transactions.