skip to content

Position, Seek and Offset Reset

Controlling where a consumer starts and moves: auto.offset.reset, seek, and timestamp lookups. Practical replay and backfill knowledge that interviewers probe with 'how would you reprocess yesterday'.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What does the consumer config auto.offset.reset control, and what do its values earliest, latest, and none do?

level: juniorimportance: must knowfreq 78%

answer

  1. only fires when NO valid committed offset
  2. earliest=oldest, latest=newest, none=throw
  3. default is latest
  4. NoOffsetForPartitionException / OffsetOutOfRangeException
  5. does NOT override existing commit

basics

~10 s

auto.offset.reset decides where a consumer starts reading when there is no valid committed offset for a partition. earliest starts at the oldest message, latest at the newest, and none throws an exception.

solid answer

~40 s

auto.offset.reset only applies when the consumer has NO valid committed offset for a partition (a brand-new consumer group, or the committed offset has fallen out of retention). It does NOT override an existing committed offset. With earliest, the consumer begins at the smallest available offset (oldest retained record). With latest (the default), it begins at the end, so it only sees records produced after it joins. With none, the consumer throws a NoOffsetForPartitionException, forcing you to handle the missing-offset case explicitly. The same setting also kicks in if a committed/requested offset is out of range (e.g. data was deleted by retention), where it is reported via OffsetOutOfRangeException internally and then reset per this policy.

go deeper

for a junior

Memorize the three values and that the default is latest; know it only matters when there's no committed offset.

for a middle

Explain the out-of-range scenario and that an existing commit overrides the reset policy.

for a senior

Discuss operational implications: latest can silently skip data, earliest can flood reprocessing, none lets you fail fast on missing offsets.

for a principal

Frame reset policy as a correctness/cost tradeoff and pair it with offset retention, topic retention, and explicit seek/reset tooling in the operational design.

## What problem this solves A Kafka topic partition is an append-only log of records, each addressed by a monotonically increasing integer called an **offset**. A **consumer** reads records sequentially and periodically records how far it has read by **committing an offset** to the internal `__consumer_offsets` topic, keyed by `(group.id, topic, partition)`. On restart or rebalance, the consumer resumes from the committed offset. But sometimes there is **no valid committed offset**: - A brand-new consumer group that has never committed. - The committed offset has expired (offsets have their own retention, `offsets.retention.minutes`, default 7 days). - The committed offset points to data that has been deleted by topic retention, so it is now **out of range**. In these cases the consumer needs a fallback policy. That policy is **`auto.offset.reset`**. ## The three values - **`earliest`** — start at the **smallest** offset still present in the partition (the oldest retained record). Use this when you must process all available history (e.g. building a cache/materialized view). - **`latest`** (the DEFAULT) — start at the **end** of the log. The consumer only sees records produced *after* it positions there. Use this when only new/live data matters and reprocessing old data is wasteful or wrong. - **`none`** — do NOT auto-reset. The consumer throws **`NoOffsetForPartitionException`** (a subclass relevant here is `OffsetOutOfRangeException` when a position is out of range). Use this when missing offsets indicate a real problem you want to detect rather than silently paper over. ## Critical edge cases 1. **It does not override an existing valid committed offset.** A common myth is that setting `earliest` rewinds every restart to the beginning — it does not. If a valid commit exists, the consumer resumes from it; `auto.offset.reset` is ignored. 2. **Out-of-range protection.** If your committed offset is, say, 1000 but retention has deleted everything below offset 5000, that offset is out of range. Internally the broker signals this and the client applies `auto.offset.reset` to recover (jumping to earliest or latest), or throws with `none`. 3. **Default is `latest`** — surprising to newcomers who expect to see old data and instead see nothing until new records arrive. 4. **An explicit `seek()` always wins** over `auto.offset.reset`; the reset policy is only the *fallback* when no position is otherwise established.

  • If a consumer group already has committed offsets, does changing auto.offset.reset to earliest make it reprocess from the start?
    No. auto.offset.reset only applies when there is no valid committed offset. With a valid commit, the consumer resumes from it regardless of the setting. To reprocess you must reset offsets explicitly (e.g. kafka-consumer-groups --reset-offsets or seekToBeginning).
  • When would auto.offset.reset=none be useful?
    When a missing offset signals an operational problem you want surfaced rather than silently handled — e.g. you require offsets to always be managed externally, and an unexpected reset to earliest/latest could cause data duplication or data loss you'd rather fail fast on.

saying these in an interview costs you the question

  • Claiming earliest rewinds to the start on every restart even with committed offsets
  • Thinking the default is earliest
  • Saying none silently starts at latest instead of throwing an exception
  • Confusing auto.offset.reset with enable.auto.commit

context

open as a page

Explain seek(), seekToBeginning(), and seekToEnd() — how do they work and when does the new position take effect?

level: middleimportance: must knowfreq 64%

basics

~10 s

These methods manually set where the consumer reads next. seek() jumps to a specific offset; seekToBeginning() to the oldest; seekToEnd() to the newest. The change takes effect on the next poll(), and overrides auto.offset.reset.

open as a page

What is the difference between position(), committed(), and beginningOffsets()/endOffsets()?

level: middleimportance: should knowfreq 52%

basics

~10 s

position() is where the consumer will read next (in-memory). committed() is the last offset persisted to Kafka for the group. beginningOffsets()/endOffsets() are the oldest and newest (end) offsets currently in each partition.

open as a page

How do you make a consumer start reading from a specific point in time, and what does offsetsForTimes return at the edges?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Use offsetsForTimes(map of partition to timestamp) to look up the first offset whose record timestamp is >= the given time, then seek() to it. If no record has a timestamp at or after the target, it returns null for that partition.

open as a page

Design how you'd reprocess a topic from a specific point for a running consumer group without losing the current state, and discuss the tradeoffs of the available mechanisms.

level: principalimportance: should knowfreq 34%

basics

~10 s

Use the kafka-consumer-groups CLI --reset-offsets (to earliest, a timestamp, or a specific offset) while the group is stopped, or seek() the partitions in code. auto.offset.reset can't do it because committed offsets already exist.

open as a page