A background worker consumes events from an at-least-once message queue and updates a search index for each event. Some events get redelivered after worker restarts. How would you use a request-deduplication store to make this consumer effectively idempotent, and what has to go into that store?
answer
- producer-assigned ID or a natural content key
- check-then-work-then-record, or record atomically with the work
- TTL must outlive max redelivery window
- Stripe webhook dedup by event ID is the textbook example
basics
~20 sThe worker keeps a record of every event ID it has already handled. Before processing a new event, it checks that record; if the ID is already there, it skips the event instead of applying it again.
solid answer
~40 sEach event needs a stable, unique identifier - either one assigned by the producer or a natural key derived from the event's content. The consumer keeps a dedup store recording IDs it has successfully processed. On each delivery, it checks the store first; if the ID is present, it skips processing and just acknowledges. If not, it does the work and records the ID as part of the same operation, ideally atomically with the side effect, so a crash between 'do the work' and 'record the ID' doesn't leave a gap. The store needs a retention window long enough to outlive the queue's maximum redelivery delay, and a mechanism that makes the lookup cheap even at high throughput.
go deeper
Should understand the basic check-before-processing idea: keep a list of what's already been handled and skip repeats.
Should describe what identifies an event (producer ID vs natural key), where the dedup record lives, and why 'record after the work, atomically' is the safe ordering.
Should reason explicitly about the TTL-vs-redelivery-window trade-off and use natural idempotency (upserts) as a second line of defense against dedup-store gaps.
Should discuss this pattern across heterogeneous systems (webhooks, stream processors, cron triggers) and connect it to real production incident shapes - e.g., late redelivery after an extended outage exceeding the dedup TTL.
## The general pattern **Request deduplication** is the general pattern that idempotency keys are one specific instance of: given that a delivery mechanism can redeliver the same logical message more than once, the receiver keeps a record of what it has already handled and skips reprocessing anything it recognizes. It applies anywhere a consumer sits behind an at-least-once channel, not just synchronous HTTP APIs guarded by a client-supplied idempotency key: - message queues - event streams - webhook receivers - cron-triggered jobs that might double-fire ## What identifies 'the same event' The first design decision is what identifies 'the same event.' Ideally the producer stamps each message with a globally unique ID at creation time, so redelivery of the identical bytes carries the identical ID. When the producer doesn't provide one, the consumer can sometimes derive a natural key from the event's own content - for example, 'order 4821 shipped' is naturally deduplicable on the tuple `(order_id, event_type)`, since a second 'shipped' event for the same order really is the same fact even if it arrived through a different delivery path. This distinction matters: - a **synthetic ID** dedups exact retransmissions; - a **natural key** can also dedup logically-equivalent events that arrive through different code paths. ## The store, and the order of the two writes The dedup store itself is usually a fast key-value structure - a database table with a unique index on the event ID, or a small TTL'd cache - that the consumer checks before doing any work: 'have I seen this ID?' - **If yes**, acknowledge the message without reprocessing. - **If no**, perform the side effect, then record the ID. The ordering and atomicity of 'do the work' versus 'record the ID' is the same correctness hazard seen with idempotency keys. | Ordering | What a crash in the middle produces | | --- | --- | | **Work first, record after** | If the side effect happens first and the crash lands before the ID is recorded, the next redelivery will redo the work (a duplicate slips through - usually tolerable if the side effect is naturally idempotent, like a search-index upsert). | | **Record first, work after** | If the ID is recorded first and the crash lands before the side effect completes, the event is permanently and silently skipped on redelivery (data loss - usually much worse). | The safer ordering is almost always 'do the work, then record the ID, atomically,' which for a database-backed side effect means writing the ID into the same transaction as the update. ## The retention trade-off The trade-off engineers have to make explicitly is retention: how long does the dedup store need to remember an ID? It must outlive the maximum plausible redelivery window of the upstream system - if the queue can redeliver a message up to 14 days after failure, the dedup record needs at least that long a TTL, or a redelivery arriving on day 15 will be treated as brand new and reapplied. Keeping records forever avoids this risk entirely but grows storage unboundedly and eventually becomes its own operational problem. Most systems pick a TTL comfortably larger than the queue's maximum redelivery window and accept the residual, very-low-probability risk of a late duplicate slipping through - which is why engineers also try to make the underlying side effect naturally idempotent (an upsert keyed by `order_id` rather than an unconditional insert) as a second line of defense, so that even a dedup-store miss doesn't corrupt data. ## Where the pattern already runs A concrete, widely used real-world instance of this pattern is a webhook receiver for a payment provider like **Stripe**: - **Stripe** explicitly documents that webhook events can be delivered more than once, and recommends the receiver store the event's ID and check it before acting on the event, exactly the dedup-store pattern described above. - **Another** is **Kafka Streams'** internal state-store checkpointing, which effectively deduplicates reprocessed records after a consumer rebalance by tracking committed offsets alongside the derived state, so a replay from the last committed offset doesn't double-apply updates already reflected in the state store.
- What's the risk if the consumer records the event ID as 'processed' before actually finishing the side effect?If the process crashes after recording the ID but before completing the side effect, the event is permanently marked as done even though it was never actually applied - a silent data-loss bug that's hard to detect because nothing errors, the event is just quietly missing. This is why the record-as-processed step should come after the side effect completes, ideally in the same atomic transaction.
- Why might you prefer a natural content-based key over a producer-assigned unique ID for deduplication?A natural key catches logically duplicate events even when they arrive through different delivery paths that don't share a synthetic ID - for example, a manually triggered replay from an admin tool and an automatic redelivery from the queue would have different message IDs but the same natural key, so only the natural key catches both as the same fact.
- If your dedup store's TTL is shorter than the queue's maximum redelivery window, what's the actual failure mode, and is it usually severe?A very late redelivery, arriving after the dedup record has expired, gets treated as brand new and reprocessed - producing a duplicate side effect. Severity depends entirely on whether the side effect is naturally idempotent downstream, in which case the duplicate is harmless, versus a strictly additive operation like 'increment a balance,' where it silently corrupts data.
It's like a bouncer with a guest list checking off names at the door: even if the same person tries to walk in through three different entrances because of a mix-up, the bouncer only lets the name through once, because the checklist remembers who's already inside.
saying these in an interview costs you the question
- Records the event as processed before performing the side effect
- No retention/TTL strategy tied to the upstream system's actual redelivery window
- Assumes producer-assigned IDs always exist without a fallback for natural-key dedup
- Treats a dedup-store hit as an error condition instead of a normal skip-and-acknowledge