After quarantining poison pills to a DLQ, how do you keep offset progress and replay/reprocessing safe and idempotent?
answer
- quarantine FIRST, then commit offset
- crash-before-commit → at-least-once reprocess (safe)
- key DLQ by orig key/offset; idempotent producer
- deterministic handler → replay reproduces same skips
- EOS: DLQ producer joins the transaction
basics
~20 sOnly commit the offset after the DLQ publish succeeds, so a crash mid-quarantine re-tries instead of skipping silently. Make handling deterministic per record and DLQ writes idempotent/keyed, so replaying the same bad record from an earlier offset produces the same quarantine, not duplicates or new gaps.
solid answer
~50 sSkip-and-advance only moves the consumer past a poison pill once the record is safely quarantined. The ordering invariant is: **publish to the DLQ first, then commit the source offset.** If you commit before the DLQ write and crash, the bad record is lost silently; if you commit after, a crash just reprocesses the same record (at-least-once into the DLQ). Make the DLQ write keyed by the original key/offset so duplicates from a retry are detectable/idempotent downstream. Replay safety means the deserialization handling is **deterministic**: replaying from an earlier offset (offset reset, application-reset tool, or topic re-read) must hit the same poison pill and produce the same skip-to-DLQ — never a thread crash and never new gaps. Avoid handlers with nondeterministic side effects. For exactly-once Streams, the DLQ producer should participate in the same transaction so the skip decision and DLQ write commit atomically with the offset. Pair the whole thing with skip-count metrics so silent loss is observable.
go deeper
Just know bad records can be sent to a DLQ and the consumer moves on.
Understand quarantine-before-commit and that the bad record stays in the log on replay.
Reason about at-least-once DLQ writes, idempotent/keyed DLQ records, and deterministic replay.
Architect EOS-transactional quarantine, DLQ re-drive tooling, and org-wide observability for silent loss.
## What "skip-and-advance" really means Handling a poison pill produces two facts that must stay consistent: 1. The bad record is **preserved** somewhere (the DLQ), so it isn't lost. 2. The consumer **offset advances** past it, so the partition unblocks. The danger is doing these in the wrong order or non-atomically. ## Ordering invariant: quarantine before commit - **Commit-then-quarantine (wrong):** if the process crashes between committing the offset and writing the DLQ, the record is gone forever — **silent data loss**. - **Quarantine-then-commit (right):** if it crashes after the DLQ write but before commit, the consumer simply reprocesses the record on restart and quarantines it again — **at-least-once** into the DLQ. That's recoverable. So: **DLQ publish → (success) → commit offset.** Spring's `DeadLetterPublishingRecoverer` follows this — the offset advances only after the recoverer succeeds. ## Idempotency of the DLQ write Because quarantine-then-commit is at-least-once, the same bad record may be written to the DLQ more than once after a crash. Make this tolerable: - **Key the DLQ record** by the original key (or original topic-partition-offset) so duplicates collide / are dedupable. - Enable **idempotent producer** (`enable.idempotence=true`) so producer retries don't add duplicates within a session. - Downstream DLQ consumers should dedupe on `(original-topic, original-partition, original-offset)` headers. ## Replay safety — determinism Reprocessing happens for many reasons: `kafka-consumer-groups --reset-offsets`, the Streams **application-reset tool**, reading a topic from the beginning into a new pipeline, or restoring from a backup. The poison pill is *still in the log* (Kafka is immutable/append-only). So your handling must be **deterministic**: - The same bytes must always be classified the same way (skip vs fail) and routed the same way. - Avoid handlers whose decision depends on wall-clock time, external mutable state, or random sampling — those make replays diverge, creating *new* gaps or *new* crashes that didn't happen the first time. A deterministic skip-to-DLQ means a full replay reproduces the exact same set of quarantined records — safe and predictable. ## Exactly-once (EOS) considerations In a Streams EOS topology, the offset commit, state-store updates, and output writes are one Kafka transaction. If you quarantine to a DLQ, the **DLQ producer should be transactional and join that transaction** so the skip + DLQ write + offset commit are atomic. Otherwise a failure can split them (record in DLQ but offset not committed → reprocessed, or vice versa), breaking the EOS guarantee around the bad record. ## Observability — make silent loss loud Skipping is invisible by design. Always: - Increment a **counter metric** per skip/quarantine. - Alert on a skip-rate threshold (a spike usually means an upstream producer changed format or schema). - Keep diagnostic headers (exception class, original topic/offset) on the DLQ record for triage and eventual reprocessing. ## Reprocessing the DLQ Fixing the root cause (e.g. correcting the producer or registering the missing schema) lets you **re-drive** the DLQ back through the pipeline. Because original bytes and headers are preserved, a replay job can re-publish to the source topic or feed a corrected consumer — closing the loop without data loss.
- Why is committing the offset before writing to the DLQ dangerous?A crash between the commit and the DLQ write loses the record permanently — silent data loss. Quarantine-then-commit instead only risks an at-least-once duplicate into the DLQ on restart, which is recoverable.
- A teammate's deserialization handler samples 1% of bad records to the DLQ at random and skips the rest. Why is that replay-unsafe?The decision is nondeterministic, so replaying from an earlier offset quarantines a different subset — producing different gaps each run. Replay safety requires the same bytes to always map to the same skip/quarantine outcome.
saying these in an interview costs you the question
- Committing the offset before confirming the DLQ write succeeded.
- Assuming skip-and-advance is exactly-once by default (it's at-least-once into the DLQ; you must dedupe).
- Using nondeterministic skip logic (time/random/external state) that breaks replay reproducibility.
- Forgetting that the poison pill stays in the immutable log, so every replay re-encounters it.
- No metrics/alerts on skip counts, hiding silent data loss.