skip to content

Why must enable.auto.commit be false in a transactional consume-transform-produce loop, and what breaks if it's left on?

level: middleimportance: must knowfreq 55%

answer

  1. auto-commit = second offset writer
  2. background timer vs transaction
  3. commits polled position regardless of produce
  4. crash → resume past lost output
  5. only sendOffsetsToTransaction moves offsets

basics

~20 s

Auto-commit makes the consumer commit offsets on its own timer, outside the producer transaction. That decouples the offset advance from the output produce, so a crash can advance offsets without the matching output (lost data) — defeating exactly-once. Offsets must be committed only via sendOffsetsToTransaction.

solid answer

~40 s

End-to-end exactly-once requires the offset commit to be atomic with the output produce, which is achieved by committing offsets through producer.sendOffsetsToTransaction inside the transaction. If enable.auto.commit=true, the consumer also commits offsets independently during poll() on the auto.commit.interval.ms timer. Now there are two writers of offsets: the transaction and the background timer. The auto-commit can advance the offset for records whose output transaction hasn't committed (or later aborts). After a crash the consumer resumes past those records, but their output was never durably written — silent data loss. Auto-commit can also commit a stale position that overwrites the transactional offset. So you set enable.auto.commit=false and let only the transaction move offsets. In Kafka Streams with exactly_once_v2 this is enforced internally.

go deeper

for a junior

Know that auto-commit must be off and offsets move only through the transaction.

for a middle

Explain the two-writers race and how it causes silent data loss on crash.

for a senior

Separate auto-commit from read_committed and explain that any consumer-side commit breaks atomicity.

for a principal

Relate to how Kafka Streams enforces this and design loops/configs so the transaction is the sole offset authority.

## What auto-commit does With `enable.auto.commit=true`, the `KafkaConsumer` periodically (every `auto.commit.interval.ms`, default 5000 ms) commits the offsets of records already returned by `poll()`, on the next `poll()` call. It's a convenience for at-least-once consumers that don't want to call `commitSync()` themselves. Crucially, it commits based on *what has been polled*, **not** on whether downstream processing or producing succeeded. ## Why it breaks transactional EOS In the CTP loop, the contract is: **offsets advance only when, and exactly when, the output transaction commits.** That's enforced by routing offsets through `producer.sendOffsetsToTransaction(...)`, which makes the coordinator write them to `__consumer_offsets` together with the output commit markers. Leaving auto-commit on creates a **second, independent offset writer**: - The background auto-commit can fire after `poll()` returns records but *before* their transaction commits — or while a transaction is in flight that later **aborts**. It commits the polled position regardless. - After a crash/restart, the consumer reads its committed offset and resumes *past* those records. But the corresponding output was rolled back or never committed → those input records are **silently dropped** (effectively at-most-once for them). - Conversely the two writers can race and overwrite each other, making the committed offset non-deterministic. Either way, atomicity is gone, so it isn't exactly-once anymore. ## The rule ``` enable.auto.commit = false // consumer never commits on its own // offsets committed ONLY via: producer.sendOffsetsToTransaction(offsets, consumer.groupMetadata()) ``` Do **not** call `commitSync()`/`commitAsync()` either — same problem, different trigger. The transaction is the single source of truth for offset progress. ## Notes - `isolation.level=read_committed` on the *downstream* consumer is a separate setting; it controls visibility of aborted records, not offset committing. Both are needed for a full EOS pipeline but they solve different problems. - Kafka Streams sets and manages all of this when `processing.guarantee=exactly_once_v2`, so you don't toggle auto-commit by hand there.

  • Is calling consumer.commitSync() manually a safe alternative to auto-commit in this loop?
    No. Any consumer-side commit (sync or async) is out-of-band and not atomic with the produce. Offsets must move only via sendOffsetsToTransaction.

saying these in an interview costs you the question

  • Saying auto-commit is fine as long as the interval is small — the race still exists at any interval.
  • Confusing enable.auto.commit (offset committing) with isolation.level=read_committed (downstream visibility).
  • Thinking commitSync() is an acceptable substitute inside the transaction loop.

context