skip to content

Walk through exactly where the duplicate-creating window is in a consumer that does an idempotent write, and how offset-commit choices interact with it.

level: seniorimportance: should knowfreq 50%

answer

  1. window = after side effect, before offset commit
  2. auto.commit: timer-based, can lose AND duplicate
  3. manual commit after work: no loss, smaller window
  4. close window only by storing offset in the sink txn
  5. rebalance redelivers in-flight uncommitted records

basics

~20 s

The window is between performing the side effect and committing the offset: a crash there causes the record to be reprocessed after rebalance/restart. Idempotent writes make that reprocessing harmless. Offset-commit choice (auto vs manual, before vs after) only shifts how often duplicates happen, not whether idempotence is needed.

solid answer

~50 s

The duplicate window opens the moment the side effect succeeds and closes when the offset is committed; a crash inside it means the next consumer resumes from the older committed offset and reprocesses the record. With `enable.auto.commit=true`, commits happen on a timer (`auto.commit.interval.ms`) during `poll()`, so the window can be large and you may even commit records you haven't finished processing — both duplicates and (worse) silent loss are possible. Best practice for idempotent consumers is manual commit (`enable.auto.commit=false`) *after* the side effect, so you never commit ahead of work; this minimizes the duplicate window but cannot eliminate it, because the write and the commit are two separate systems with no shared transaction. Rebalances widen exposure: in-flight uncommitted records get reprocessed by the new owner. The conclusion: tune commits to shrink and bound duplicates, but rely on idempotence — not commit timing — for correctness.

go deeper

for a junior

Know the window is between doing the work and committing the offset, and a crash there reprocesses the record.

for a middle

Contrast auto vs manual commit and know manual-after-work avoids loss and shrinks duplicates.

for a senior

Explain auto-commit's loss risk, why the window can't close without a shared transaction, and rebalance amplification.

for a principal

Decide commit strategy and the store-offset-in-sink pattern across services, balancing duplicate volume, loss risk, and rebalance behavior at scale.

## The anatomy of one record Processing a record has three logical steps: - **(A)** the framework hands you the record; - **(B)** you perform the side effect; - **(C)** you commit the offset (tell Kafka 'I'm past this position'). Kafka's only durable memory of progress is the **committed offset**. There is no transaction spanning B and C — B hits your DB/API, C hits Kafka's `__consumer_offsets` topic. ## Where the window is Between B succeeding and C committing. If the process dies, the JVM is killed, or a **rebalance** revokes the partition in that gap, then the last committed offset still points *before* this record. The consumer that next owns the partition (this one after restart, or another group member) re-serves the record and runs B again. That is the duplicate. Idempotence makes the second B a no-op in terms of observable state. ## Auto-commit (`enable.auto.commit=true`) The client commits asynchronously on a schedule (`auto.commit.interval.ms`, default 5000ms) — specifically, the commit is triggered inside `poll()` and commits the offsets of records returned by the *previous* poll. Two consequences: - (1) the duplicate window can be up to the commit interval wide, so a crash reprocesses everything since the last timer commit; - (2) more dangerously, the timer can commit offsets for records you returned but have *not finished processing* — if you crash after the auto-commit but before finishing, those records are skipped → silent **loss** (at-most-once leakage). Auto-commit is convenient but couples offset progress to wall-clock, not to your work. ## Manual commit after the effect Set `enable.auto.commit=false` and call `commitSync()` / `commitAsync()` *after* B completes for the record (or batch). Now you never commit ahead of work, so loss is avoided, and the window is just B→C for the current batch — much smaller and bounded by batch size. - `commitSync` is durable-before-proceeding but slower; - `commitAsync` is faster but a failed async commit can leave an older offset, slightly widening the window. Either way the window is non-zero. ## Why you can't close it without a transaction B and C are different systems. To make them atomic you'd need the offset stored *in the same transactional store as the side effect* — e.g. write the offset into your DB in the same transaction as the business row (the **'store offset in the sink' pattern**), then on startup seek to that stored offset instead of Kafka's. That collapses the window. Short of that (or Kafka EOS for internal sinks), the window exists and idempotence is what saves you. ## Rebalances amplify it During a consumer-group rebalance, partitions are revoked and reassigned. Any record that was processed (B done) but not yet committed (C pending) at revocation time will be redelivered to the new owner. High rebalance frequency (scaling, deploys, `max.poll.interval.ms` violations from slow processing) therefore increases duplicate volume. `ConsumerRebalanceListener.onPartitionsRevoked` is the hook to commit-on-revoke and shrink this, but it still can't make B+C atomic. ## The senior conclusion Offset-commit strategy is a *duplicate-rate and loss-avoidance* knob: - manual commit after work eliminates loss and minimizes duplicates; - auto-commit is sloppier on both. But no commit strategy gives correctness on its own. Correctness for an external sink comes from **idempotent processing**; commit tuning just keeps the duplicate volume (and reprocessing cost) low.

  • How can auto-commit cause data loss, not just duplicates?
    The timer commits offsets for records returned by the previous poll even if you haven't finished processing them. If you crash after that commit but before finishing, those records are never reprocessed — they're skipped, i.e. lost. Manual commit-after-work avoids this.
  • Is there any way to make the side effect and the offset commit truly atomic without Kafka transactions?
    Yes — store the offset in the same transactional store as the side effect (write business row + offset in one DB transaction) and on startup seek to the stored offset. That collapses the B→C window. It only works when the sink is itself transactional.

saying these in an interview costs you the question

  • Saying manual offset commit eliminates duplicates — it shrinks the window but can't close it without a shared transaction.
  • Believing auto-commit only risks duplicates; it can also silently lose records.
  • Ignoring rebalances as a major source of reprocessing.
  • Claiming commit-after-processing is unnecessary if writes are idempotent — it still matters for loss avoidance and duplicate volume.
  • Thinking committing offsets and writing the side effect are part of one Kafka transaction for an external sink.

context