skip to content

On the consumer side, how do the order of offset-commit vs record-processing produce at-most-once versus at-least-once, and what crash window causes loss or duplicates in each?

level: middleimportance: must knowfreq 70%

answer

  1. commit-first → skip → loss
  2. process-first → replay → dup
  3. window = between the two ops
  4. auto-commit = at-least-once, 5s interval
  5. rebalance counts as restart

basics

~10 s

Commit-before-process = at-most-once: if you crash after committing but before processing, the record is skipped (lost). Process-before-commit = at-least-once: if you crash after processing but before committing, you reprocess it (duplicate).

solid answer

~50 s

The consumer chooses the semantic by ordering two steps. With **commit-before-process**, you advance the committed offset first, then do the work. The danger window is *between commit and the end of processing*: if the consumer dies there, on restart it resumes after the committed offset, so the unprocessed record is skipped — **at-most-once / possible loss**. With **process-before-commit** (the default flow), you do the work first, then commit. The danger window is *between finishing processing and the commit landing*: if the consumer dies there, on restart it re-reads from the last committed offset and processes the record again — **at-least-once / possible duplicates**. Kafka's `enable.auto.commit=true` commits periodically (`auto.commit.interval.ms`) and is effectively at-least-once because records processed since the last auto-commit get replayed after a crash. To control the window precisely you set `enable.auto.commit=false` and call `commitSync()`/`commitAsync()` yourself relative to your processing.

go deeper

for a junior

Remember the two orderings and which one loses vs duplicates.

for a middle

Be able to point to the precise crash window and explain auto-commit's effect.

for a senior

Discuss rebalances, async-commit failure modes, and why ordering alone cannot reach exactly-once.

for a principal

Advise teams on commit strategy + idempotency design vs adopting transactional EOS based on cost and correctness needs.

## The two operations For each batch a consumer does two things: 1. **Process** the record(s) — the business side effect (DB write, downstream call). 2. **Commit the offset** — persist 'I've consumed up to offset N' to the internal `__consumer_offsets` topic so that after a restart or rebalance, consumption resumes from N+1. A crash can happen at any instant. The semantic depends entirely on which of these two you do first, because that decides what gets replayed vs skipped after the crash. ## Commit-before-process → at-most-once Sequence: `commit(offset) ; process(record)`. - Normal path: commit lands, then you process. Fine. - **Loss window**: the consumer crashes AFTER the commit is durable but BEFORE processing completes. On restart, the committed offset already points past this record, so the consumer never re-reads it. The record's work was never done → **lost**. No record is ever seen twice (the commit moved past it before any retry could occur) → no duplicates. Use when reprocessing is worse than dropping (e.g., some metrics/telemetry where a stale duplicate is unacceptable but a rare gap is tolerable). ## Process-before-commit → at-least-once (default) Sequence: `process(record) ; commit(offset)`. - Normal path: process, then commit. Fine. - **Duplicate window**: the consumer crashes AFTER processing completes but BEFORE the commit is durable. On restart, the last committed offset still points at (or before) this record, so it is re-read and **processed again** → **duplicate**. Nothing is ever skipped, because you never advance the offset past unprocessed work → no loss. This is the safe default for most pipelines; pair it with **idempotent processing** so duplicates are harmless. ## Auto-commit nuance `enable.auto.commit=true` commits the offsets of the *last poll()* at intervals of `auto.commit.interval.ms` (default 5s), triggered during `poll()`. It commits offsets for records that were *returned*, on the assumption they were processed. If you crash mid-processing, everything since the last auto-commit replays → at-least-once with a potentially large duplicate window. Setting `enable.auto.commit=false` and committing manually (`commitSync` after each batch) tightens the window and makes the ordering explicit. ## Edge cases - **Rebalances** count as 'restarts': when a partition moves to another consumer, the new owner resumes from the committed offset, so the same commit/process ordering governs duplicates/loss across rebalances too. - **Async commit (`commitAsync`)** can fail silently; a later successful commit usually covers it, but a crash right after a failed async commit widens the duplicate window. - You cannot eliminate BOTH windows with ordering alone — that requires making process+commit atomic via transactions (exactly-once). ## Summary table | Order | Crash window | Result | |---|---|---| | commit → process | after commit, before process | loss (at-most-once) | | process → commit | after process, before commit | duplicate (at-least-once) |

  • Does enable.auto.commit=true give at-most-once or at-least-once?
    At-least-once. Auto-commit happens after records are returned by poll() and assumed processed, but a crash mid-processing replays everything since the last commit interval — duplicates, not loss.
  • How would you minimize the duplicate window in a process-before-commit consumer?
    Disable auto-commit and commitSync() the exact offsets immediately after each batch is fully processed, so the unflushed window is one batch rather than up to auto.commit.interval.ms.

saying these in an interview costs you the question

  • Saying process-before-commit can lose messages (it can only duplicate them).
  • Claiming auto-commit is at-most-once.
  • Forgetting that a rebalance, not just a process crash, triggers replay from the committed offset.
  • Believing tighter commit timing can achieve exactly-once (it only shrinks, never closes, the window).

context