Design how you'd reprocess a topic from a specific point for a running consumer group without losing the current state, and discuss the tradeoffs of the available mechanisms.
answer
- auto.offset.reset useless on existing group
- CLI --reset-offsets needs INACTIVE group
- --to-earliest / --to-datetime / --to-offset / --shift-by
- in-code seek via rebalance listener; you own coordination
- separate group.id for zero-blast-radius replay + idempotency
basics
~10 sUse the kafka-consumer-groups CLI --reset-offsets (to earliest, a timestamp, or a specific offset) while the group is stopped, or seek() the partitions in code. auto.offset.reset can't do it because committed offsets already exist.
solid answer
~40 sFor an existing group you can't rely on auto.offset.reset (it only fires when there's NO committed offset). Options: (1) Offline, use kafka-consumer-groups.sh --reset-offsets with --to-earliest, --to-datetime, --to-offset, --shift-by, or --to-current — but the group must be INACTIVE (no live members) or the command refuses, to avoid racing the live position. (2) In-code, register a ConsumerRebalanceListener and seek()/offsetsForTimes()+seek() to reposition on assignment, optionally driven by an external trigger. (3) For surgical replay without disturbing the production group, spin up a SEPARATE group.id (or use assign() with no group) reading the same data into a side pipeline. Tradeoffs: CLI is simple but needs downtime; in-code seek is dynamic but you own the coordination; a separate group isolates blast radius but doubles read load. Combine with idempotent/transactional writes since replay implies reprocessing.
go deeper
Know that you can reset a group's offsets with the kafka-consumer-groups CLI and that auto.offset.reset won't reprocess an existing group.
Use --reset-offsets options and understand the inactive-group requirement and dry-run vs execute.
Choose between CLI, in-code seek, and separate group.id based on downtime and isolation needs, and pair with idempotency.
Architect a replay/backfill strategy weighing blast radius, downtime, read-load, rebalance coordination, time-index accuracy, and exactly-once guarantees.
## Why auto.offset.reset is the wrong tool here `auto.offset.reset` only applies when a partition has **no valid committed offset**. A running production group already has committed offsets, so changing the config does nothing for it. Reprocessing requires explicitly **moving** offsets. ## Mechanism 1 — kafka-consumer-groups.sh --reset-offsets (operational) The CLI repositions a group's committed offsets: - Scope: `--all-topics` or `--topic t` (optionally `:partitions`). - Target: `--to-earliest`, `--to-latest`, `--to-offset N`, `--to-datetime <ISO8601>` (uses the time index, like `offsetsForTimes`), `--shift-by ±N`, `--to-current`, or `--from-file`. - Modes: `--dry-run` (default, prints what *would* change) vs `--execute`. **Hard constraint:** the group must have **no active members** — the tool refuses to reset a live group to avoid racing the running consumers' positions. So this implies a controlled **downtime/maintenance window**: stop consumers, reset, restart. Pro: simple, auditable, no code. Con: requires stopping the group. ## Mechanism 2 — in-code seek / offsetsForTimes (programmatic) Inside the application: on assignment (`ConsumerRebalanceListener.onPartitionsAssigned`) or on an external command, call `seek(tp, target)`, or `offsetsForTimes` + `seek` for time-based replay, or `seekToBeginning`. Pro: dynamic, no restart, fine-grained (per partition / per condition), can be triggered at runtime. Con: you own the coordination (which instance seeks which partitions during a rebalance), and a naive seek without committing can be undone by the next rebalance. Pattern: store the desired replay point externally and re-seek deterministically on every assignment until replay is complete. ## Mechanism 3 — separate group / manual assign (isolation) To replay without touching the production group at all, start a **new `group.id`** (or use `assign()` with no group coordination) that reads the same partitions into a separate sink. Pro: zero blast radius on the live group; you can replay all history while production continues. Con: doubles broker read load and you must wire up a parallel consumer + sink. ## Cross-cutting concern — idempotency Any replay reprocesses records, so downstream effects must tolerate duplicates: idempotent writes (upserts keyed by event id), or Kafka transactions / exactly-once for read-process-write within Kafka. Otherwise replay causes double side effects. ## Choosing - One-off backfill, downtime acceptable -> CLI `--reset-offsets --to-datetime`. - Runtime-triggered, per-partition, no restart -> in-code seek/offsetsForTimes. - Replay while production must keep running untouched -> separate group.id. In all cases verify with `--describe` / `position()`/`committed()` and bound seeks with `beginningOffsets`/`endOffsets`.
- Why does kafka-consumer-groups --reset-offsets refuse to run while the group is active?Because live members are continuously advancing and committing positions; resetting underneath them would race and produce inconsistent or immediately-overwritten offsets. Requiring an inactive group guarantees the reset is the sole writer of those committed offsets.
- If you must replay but can't take downtime, what's your approach?Either in-code seek/offsetsForTimes driven by a runtime trigger inside the running app, or spin up a separate group.id reading the same partitions into a side pipeline so the production group is untouched. Pair with idempotent/transactional downstream writes.
- How do you make replay safe against duplicate side effects?Make downstream writes idempotent (upsert by event key) or use Kafka transactions / exactly-once semantics for read-process-write loops, so reprocessing the same records doesn't double-apply effects.
saying these in an interview costs you the question
- Proposing to flip auto.offset.reset=earliest to reprocess an existing group (no effect)
- Running --reset-offsets against a live group (it refuses; would race)
- Forgetting --execute (dry-run is the default and changes nothing)
- Ignoring idempotency so replay double-applies side effects
- Assuming an in-code seek persists without committing or re-seeking on rebalance