skip to content

Offset Management and Commit Strategies

Auto-commit versus manual commitSync/commitAsync, what a committed offset actually means, and where offsets are stored. Interviewers ask because commit placement is what decides at-least-once versus at-most-once.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is a committed offset in Kafka, and what value does a consumer actually store when it commits?

level: juniorimportance: must knowfreq 80%

answer

  1. committed = next-to-read = lastProcessed + 1
  2. per (group, topic, partition)
  3. stored in __consumer_offsets
  4. OffsetAndMetadata wraps offset + 1
  5. no commit -> auto.offset.reset

basics

~20 s

A committed offset records how far a consumer group has processed in a partition. It stores the offset of the NEXT message to read (last processed + 1), so a restart resumes from there without re-reading.

solid answer

~30 s

An offset is the position of a record within a partition. A committed offset is the consumer group's saved progress, stored per (group, topic, partition). The committed value is the next-to-read position: if you processed record at offset 10, you commit 11 (lastProcessed + 1). On restart or rebalance, a consumer reads its committed offset and resumes from exactly that point. If no offset is committed (new group), auto.offset.reset (earliest/latest/none) decides where to start. Commits are stored in the internal __consumer_offsets topic. The KafkaConsumer API expresses the commit value as an OffsetAndMetadata wrapping that next-to-read offset plus optional metadata.

go deeper

for a junior

Know that an offset is a record's position and that committing saves progress as last + 1 so restarts resume correctly.

for a middle

Be precise that the value is next-to-read, stored per group/topic/partition in __consumer_offsets, expressed via OffsetAndMetadata.

for a senior

Explain the +1 off-by-one duplicate bug, the relationship to auto.offset.reset, and when each applies.

for a principal

Discuss leader epoch in OffsetAndMetadata for fencing, offset expiry/retention, and designing commit semantics for exactly-once vs at-least-once.

## Offsets, first principles A Kafka **partition** is an append-only log of records. Each record has a monotonically increasing integer **offset** (0, 1, 2, ...) that is its permanent position in that partition. Consumers read sequentially and track a **position** — the offset of the next record they will fetch. ## What 'committing' means A **committed offset** is the durably saved processing progress for a **consumer group** on a specific **(topic, partition)**. It answers: 'if a consumer in this group (re)starts owning this partition, where should it begin?' Committing is how progress survives crashes, restarts, and rebalances. ## The key gotcha: committed = NEXT to read, not last processed The most common interview trap. If you successfully process the record at offset 10, the value you commit is **11**, not 10. The committed offset is the position the consumer should *resume at* — i.e., `lastProcessedOffset + 1`. The high-level `commitSync()`/`commitAsync()` with no arguments handle this automatically by committing the consumer's current `position()` for each assigned partition (which already points one past the last fetched record). When you commit a specific map yourself, you must add 1 manually: ``` new OffsetAndMetadata(record.offset() + 1) ``` ## OffsetAndMetadata Commits are expressed as a `Map<TopicPartition, OffsetAndMetadata>`. `OffsetAndMetadata` wraps the offset (the next-to-read value) plus an optional `metadata` String you can attach (e.g., a trace id or app-level marker) and, in newer clients, leader epoch for fencing. ## Where it's stored Committed offsets live in the internal compacted topic `__consumer_offsets`, keyed by (group, topic, partition). The group coordinator broker owns the partition of that topic for the group. ## No committed offset yet For a brand-new group (or expired offsets), there is nothing to resume from, so `auto.offset.reset` applies: `earliest` (start at the beginning), `latest` (only new records), or `none` (throw). This is distinct from a committed offset and only kicks in when no valid commit exists. ## Why +1 matters operationally Get it wrong by committing `record.offset()` instead of `+1`, and on restart you reprocess the last record of every batch — a subtle, persistent duplicate.

  • You processed up to offset 99. What value do you commit and why?
    100. The committed offset is the next position to read, so it is lastProcessed + 1 = 99 + 1. Committing 99 would cause record 99 to be re-read on restart.
  • What happens if a consumer group has never committed an offset for a partition?
    There is no resume point, so auto.offset.reset decides: earliest starts at the log beginning, latest at the end (new records only), none throws NoOffsetForPartitionException.

saying these in an interview costs you the question

  • Saying the committed offset equals the last processed offset (it's last + 1)
  • Claiming offsets are stored in ZooKeeper (true only for very old Kafka < 0.9; now __consumer_offsets)
  • Confusing position/committed offset with auto.offset.reset (reset only applies when no commit exists)

context

open as a page

Compare automatic offset commits (enable.auto.commit) with manual commits. What are the trade-offs and failure modes?

level: middleimportance: must knowfreq 85%

basics

~20 s

With enable.auto.commit=true, the consumer commits the current position periodically (auto.commit.interval.ms, default 5000ms) during poll(). It's simple but can lose or duplicate records on crash. Manual commits (commitSync/commitAsync after processing) give you control over timing and delivery semantics.

open as a page

You need at-least-once delivery with no data loss in a consumer. How do you order processing and commits, and how do you bound duplicates?

level: seniorimportance: must knowfreq 70%

basics

~20 s

Disable auto-commit, process the records first, then commit. Never commit before the work is durable. Because a crash between processing and commit replays the batch, make handlers idempotent. For exactly-once, use Kafka transactions or commit offsets atomically with your output to an external store.

open as a page

Explain the __consumer_offsets topic: why it's log-compacted, how offsets are keyed, and how the group coordinator uses it.

level: seniorimportance: should knowfreq 45%

basics

~20 s

__consumer_offsets is an internal Kafka topic (50 partitions by default) where consumer commits are stored as messages keyed by (group, topic, partition). It's log-compacted so only the latest committed offset per key is retained. The group coordinator broker reads/writes it to track group progress.

open as a page

How do you commit offsets safely around a consumer group rebalance using ConsumerRebalanceListener?

level: seniorimportance: should knowfreq 60%

basics

~20 s

Register a ConsumerRebalanceListener via subscribe(). In onPartitionsRevoked, commit offsets for partitions you're about to lose before they move to another consumer. onPartitionsAssigned runs when you gain partitions (e.g., to seek to a custom position). This prevents duplicate or lost processing across rebalances.

open as a page