skip to content

When a Kafka consumer processes a batch of 10 records and commits only the offset of the last record after the whole batch succeeds, what happens on a crash after record 6 finishes, and why is committing after every single record not simply 'safer'?

level: middleimportance: must knowfreq 75%

answer

  1. commit = single per-partition cursor, not per-message ack
  2. batch size sets replay-window size
  3. commitAsync can complete out of order
  4. commit inside onPartitionsRevoked before losing a partition
  5. commit granularity is a throughput knob, not a duplicate-elimination tool

basics

~20 s

Committing the offset once per batch is faster but replays the whole batch if a crash happens partway through, since Kafka only tracks one checkpoint per partition, not per message. Committing after every single message shrinks that replay window but is much slower.

solid answer

~40 s

A Kafka offset commit isn't a per-message ack; it's a single per-partition cursor saying 'everything before this point is done.' Committing once per batch is cheap but means a crash mid-batch replays the entire batch, since the cursor hasn't moved past record zero of that batch yet. Committing after every record shrinks the replay window to at most one record but adds a network round trip per record, cutting throughput sharply. The common middle ground is commitAsync on smaller sub-batches plus a final commitSync on shutdown or rebalance. None of these choices eliminates duplicates — they only change how many happen — so the downstream write still has to tolerate replay regardless of batch size.

go deeper

for a junior

Should know that committing less often is faster but replays more records on crash — the basic throughput-vs-safety trade.

for a middle

Should explain that a Kafka offset is a single per-partition cursor, not per-message, and connect batch size directly to replay-window size with a concrete example.

for a senior

Should know the commitAsync out-of-order hazard and the rebalance interaction, and reason about mixing async steady-state commits with a sync shutdown commit.

for a principal

Should design the commit-granularity policy per data-flow based on the cost of replay (idempotent upsert vs expensive external side effect) and set guidance distinguishing 'commit tuning' from 'duplicate elimination,' which requires a separate idempotency layer.

## What commit ordering actually decides Offset commit ordering refers to two related decisions: 1. **At what granularity you commit** — after every single record, after every batch/poll, or on a timer. 2. **In what order commits actually land** relative to the processing they represent, especially when commits are issued asynchronously. In Kafka, a commit doesn't acknowledge an individual message the way a queue ack does; it writes a single cursor value per partition — 'the next offset this consumer group should read from here' — to an internal offsets topic. Because the guarantee is expressed as one monotonically-advancing number per partition rather than one flag per message, committing offset `110` implicitly tells the broker 'everything up through 109 on this partition is done,' even if what actually happened was 'I processed records 100-109 as one batch and I'm only checkpointing once, at the end.' That single design choice is why granularity matters so much: it directly sets the size of the replay window on crash. ## Why a cursor instead of per-message acks This design exists because per-message acknowledgment, as RabbitMQ or JMS do, requires the broker to track delivery state for every in-flight message individually, which doesn't scale to the sustained high-throughput, ordered-log workloads Kafka targets. A single per-partition cursor is cheap to store and cheap to advance, but it forces an all-or-nothing checkpoint over whatever batch of records preceded it. Consumers therefore have to explicitly choose how often to move that cursor, trading throughput against the size of the 'at risk' window. ## The granularity choices - **After every single record.** Committing after every single record minimizes the replay window — at most one record gets reprocessed after a crash — but every commit is itself a network round trip to the offsets topic, so throughput drops sharply, often by an order of magnitude, under `commitSync()`. - **Once per polled batch.** Committing once per polled batch amortizes that cost and is the common production pattern, but it means a crash after processing record 6 of a 10-record batch causes all 10 to be replayed, not just the unfinished 4, because the cursor hasn't moved past record 0 of that batch yet. - **Async steady-state batches with a final `commitSync()`.** Using `commitAsync()` for steady-state batches with a final `commitSync()` on shutdown gets most of the throughput of infrequent commits with a smaller staleness window, but introduces its own hazard: async commits can complete out of order, so a consumer must never let an older, in-flight async commit overwrite a newer one, which client libraries handle by tracking a monotonically increasing sequence and discarding stale commit callbacks. ## Failure modes 1. **Duplicate storms after a restart.** The most common production issue is teams choosing large batch/commit intervals purely for throughput and then being surprised by 'duplicate storms' after any restart or rolling deploy, because every restart replays the entire uncommitted tail of the last batch, not just the record that was mid-flight. 2. **Committing for a reassigned partition.** A second failure mode is committing an offset for a partition the consumer no longer owns after a rebalance: if a `commitAsync()` call issued just before a rebalance completes after the partition has been reassigned, it can commit a stale offset that causes the new owner to skip records it hasn't processed yet, or briefly conflicts with the new owner's own commits — this is why offset commits should be tied to the consumer-group generation and why manual-commit code typically calls a final synchronous commit inside the rebalance-listener's `onPartitionsRevoked` callback before giving up a partition. 3. **Porting the ack model between broker types.** A third, subtler failure is out-of-order manual commits when porting code between broker types: in RabbitMQ, acking message 8 before message 5 is safe per-message, but mistaking RabbitMQ's per-message ack model for Kafka's single-cursor model, or vice versa, can silently change the delivery guarantee. ## A worked pattern A concrete production pattern: an order-processing consumer polls 200 records per batch, writes each order to a database inside the loop, and calls `commitSync()` only once after the whole batch of 200 finishes. Throughput is good because there's one round trip per 200 records instead of 200. But a crash after order #150 replays all 200 orders on restart, so the database write must be an idempotent upsert keyed on order ID regardless of batch size — the batch-size choice only changes how many redundant upserts happen, not whether they can happen at all. Shrinking the batch size or adding a mid-batch async commit every 20 records reduces the blast radius of a crash but never removes the need for the write itself to tolerate replay, which is exactly why commit granularity is a throughput/staleness knob, not a substitute for making the write itself safe to repeat.

  • Why doesn't committing after every single record eliminate duplicates entirely?
    It shrinks the replay window to at most one record, but a crash can still occur after the record is processed and before the commit's network round trip completes, so the last record can still be reprocessed. Per-record commit minimizes but never zeroes out the duplicate window; only a transactional write that atomically bundles the output and the offset removes it.
  • What specifically goes wrong if a consumer commits offsets asynchronously right as a rebalance is happening?
    An in-flight commitAsync issued just before the partition is revoked can land after the partition has already been reassigned, either committing a stale offset that makes the new owner skip unprocessed records or racing against the new owner's own commits. The standard fix is to flush a synchronous commit inside the onPartitionsRevoked callback before releasing the partition.
  • If a team wants to reduce duplicate reprocessing without paying the full cost of per-record commitSync, what's a middle-ground pattern?
    Use commitAsync on a smaller sub-batch interval, such as every 20-50 records instead of every 500, during steady-state processing, paired with a final commitSync on shutdown or rebalance to guarantee the last state is durably flushed. This trades some throughput for a smaller crash-replay window without paying a synchronous round trip on every single record.

Like a librarian who, instead of stamping each returned book back onto the shelf list individually, only updates the master ledger once after reshelving an entire cart of twenty books — if she's called away after shelving twelve, the ledger still shows none of the twenty as done, so all twenty get 're-shelved' by the next librarian, even the twelve that were already fine.

saying these in an interview costs you the question

  • Thinks a Kafka offset commit acknowledges individual messages the way a queue ack does
  • Assumes per-record commitSync fully eliminates duplicates
  • Doesn't know a large batch commit interval widens the replay window on crash, not just reduces overhead
  • Unaware that async commits can complete out of order and needs a monotonic guard
  • Never mentions committing inside onPartitionsRevoked, or an equivalent, before losing a partition on rebalance

context