What is a committed offset in Kafka, and what value does a consumer actually store when it commits?
answer
- committed = next-to-read = lastProcessed + 1
- per (group, topic, partition)
- stored in __consumer_offsets
- OffsetAndMetadata wraps offset + 1
- no commit -> auto.offset.reset
basics
~20 sA committed offset records how far a consumer group has processed in a partition. It stores the offset of the NEXT message to read (last processed + 1), so a restart resumes from there without re-reading.
solid answer
~30 sAn offset is the position of a record within a partition. A committed offset is the consumer group's saved progress, stored per (group, topic, partition). The committed value is the next-to-read position: if you processed record at offset 10, you commit 11 (lastProcessed + 1). On restart or rebalance, a consumer reads its committed offset and resumes from exactly that point. If no offset is committed (new group), auto.offset.reset (earliest/latest/none) decides where to start. Commits are stored in the internal __consumer_offsets topic. The KafkaConsumer API expresses the commit value as an OffsetAndMetadata wrapping that next-to-read offset plus optional metadata.
go deeper
Know that an offset is a record's position and that committing saves progress as last + 1 so restarts resume correctly.
Be precise that the value is next-to-read, stored per group/topic/partition in __consumer_offsets, expressed via OffsetAndMetadata.
Explain the +1 off-by-one duplicate bug, the relationship to auto.offset.reset, and when each applies.
Discuss leader epoch in OffsetAndMetadata for fencing, offset expiry/retention, and designing commit semantics for exactly-once vs at-least-once.
## Offsets, first principles A Kafka **partition** is an append-only log of records. Each record has a monotonically increasing integer **offset** (0, 1, 2, ...) that is its permanent position in that partition. Consumers read sequentially and track a **position** — the offset of the next record they will fetch. ## What 'committing' means A **committed offset** is the durably saved processing progress for a **consumer group** on a specific **(topic, partition)**. It answers: 'if a consumer in this group (re)starts owning this partition, where should it begin?' Committing is how progress survives crashes, restarts, and rebalances. ## The key gotcha: committed = NEXT to read, not last processed The most common interview trap. If you successfully process the record at offset 10, the value you commit is **11**, not 10. The committed offset is the position the consumer should *resume at* — i.e., `lastProcessedOffset + 1`. The high-level `commitSync()`/`commitAsync()` with no arguments handle this automatically by committing the consumer's current `position()` for each assigned partition (which already points one past the last fetched record). When you commit a specific map yourself, you must add 1 manually: ``` new OffsetAndMetadata(record.offset() + 1) ``` ## OffsetAndMetadata Commits are expressed as a `Map<TopicPartition, OffsetAndMetadata>`. `OffsetAndMetadata` wraps the offset (the next-to-read value) plus an optional `metadata` String you can attach (e.g., a trace id or app-level marker) and, in newer clients, leader epoch for fencing. ## Where it's stored Committed offsets live in the internal compacted topic `__consumer_offsets`, keyed by (group, topic, partition). The group coordinator broker owns the partition of that topic for the group. ## No committed offset yet For a brand-new group (or expired offsets), there is nothing to resume from, so `auto.offset.reset` applies: `earliest` (start at the beginning), `latest` (only new records), or `none` (throw). This is distinct from a committed offset and only kicks in when no valid commit exists. ## Why +1 matters operationally Get it wrong by committing `record.offset()` instead of `+1`, and on restart you reprocess the last record of every batch — a subtle, persistent duplicate.
- You processed up to offset 99. What value do you commit and why?100. The committed offset is the next position to read, so it is lastProcessed + 1 = 99 + 1. Committing 99 would cause record 99 to be re-read on restart.
- What happens if a consumer group has never committed an offset for a partition?There is no resume point, so auto.offset.reset decides: earliest starts at the log beginning, latest at the end (new records only), none throws NoOffsetForPartitionException.
saying these in an interview costs you the question
- Saying the committed offset equals the last processed offset (it's last + 1)
- Claiming offsets are stored in ZooKeeper (true only for very old Kafka < 0.9; now __consumer_offsets)
- Confusing position/committed offset with auto.offset.reset (reset only applies when no commit exists)