What does the consumer config auto.offset.reset control, and what do its values earliest, latest, and none do?
answer
- only fires when NO valid committed offset
- earliest=oldest, latest=newest, none=throw
- default is latest
- NoOffsetForPartitionException / OffsetOutOfRangeException
- does NOT override existing commit
basics
~10 sauto.offset.reset decides where a consumer starts reading when there is no valid committed offset for a partition. earliest starts at the oldest message, latest at the newest, and none throws an exception.
solid answer
~40 sauto.offset.reset only applies when the consumer has NO valid committed offset for a partition (a brand-new consumer group, or the committed offset has fallen out of retention). It does NOT override an existing committed offset. With earliest, the consumer begins at the smallest available offset (oldest retained record). With latest (the default), it begins at the end, so it only sees records produced after it joins. With none, the consumer throws a NoOffsetForPartitionException, forcing you to handle the missing-offset case explicitly. The same setting also kicks in if a committed/requested offset is out of range (e.g. data was deleted by retention), where it is reported via OffsetOutOfRangeException internally and then reset per this policy.
go deeper
Memorize the three values and that the default is latest; know it only matters when there's no committed offset.
Explain the out-of-range scenario and that an existing commit overrides the reset policy.
Discuss operational implications: latest can silently skip data, earliest can flood reprocessing, none lets you fail fast on missing offsets.
Frame reset policy as a correctness/cost tradeoff and pair it with offset retention, topic retention, and explicit seek/reset tooling in the operational design.
## What problem this solves A Kafka topic partition is an append-only log of records, each addressed by a monotonically increasing integer called an **offset**. A **consumer** reads records sequentially and periodically records how far it has read by **committing an offset** to the internal `__consumer_offsets` topic, keyed by `(group.id, topic, partition)`. On restart or rebalance, the consumer resumes from the committed offset. But sometimes there is **no valid committed offset**: - A brand-new consumer group that has never committed. - The committed offset has expired (offsets have their own retention, `offsets.retention.minutes`, default 7 days). - The committed offset points to data that has been deleted by topic retention, so it is now **out of range**. In these cases the consumer needs a fallback policy. That policy is **`auto.offset.reset`**. ## The three values - **`earliest`** — start at the **smallest** offset still present in the partition (the oldest retained record). Use this when you must process all available history (e.g. building a cache/materialized view). - **`latest`** (the DEFAULT) — start at the **end** of the log. The consumer only sees records produced *after* it positions there. Use this when only new/live data matters and reprocessing old data is wasteful or wrong. - **`none`** — do NOT auto-reset. The consumer throws **`NoOffsetForPartitionException`** (a subclass relevant here is `OffsetOutOfRangeException` when a position is out of range). Use this when missing offsets indicate a real problem you want to detect rather than silently paper over. ## Critical edge cases 1. **It does not override an existing valid committed offset.** A common myth is that setting `earliest` rewinds every restart to the beginning — it does not. If a valid commit exists, the consumer resumes from it; `auto.offset.reset` is ignored. 2. **Out-of-range protection.** If your committed offset is, say, 1000 but retention has deleted everything below offset 5000, that offset is out of range. Internally the broker signals this and the client applies `auto.offset.reset` to recover (jumping to earliest or latest), or throws with `none`. 3. **Default is `latest`** — surprising to newcomers who expect to see old data and instead see nothing until new records arrive. 4. **An explicit `seek()` always wins** over `auto.offset.reset`; the reset policy is only the *fallback* when no position is otherwise established.
- If a consumer group already has committed offsets, does changing auto.offset.reset to earliest make it reprocess from the start?No. auto.offset.reset only applies when there is no valid committed offset. With a valid commit, the consumer resumes from it regardless of the setting. To reprocess you must reset offsets explicitly (e.g. kafka-consumer-groups --reset-offsets or seekToBeginning).
- When would auto.offset.reset=none be useful?When a missing offset signals an operational problem you want surfaced rather than silently handled — e.g. you require offsets to always be managed externally, and an unexpected reset to earliest/latest could cause data duplication or data loss you'd rather fail fast on.
saying these in an interview costs you the question
- Claiming earliest rewinds to the start on every restart even with committed offsets
- Thinking the default is earliest
- Saying none silently starts at latest instead of throwing an exception
- Confusing auto.offset.reset with enable.auto.commit