skip to content

In a notification service consuming order events at least once, how do you stop a redelivered event from texting the same user twice?

level: middleimportance: must knowfreq 72%

answer

  1. same logical message, same key
  2. what stays stable across redelivery
  3. claim before the side effect
  4. unique constraint, not check-then-insert
  5. timeout: sent or not?

basics

~20 s

Derive a deterministic idempotency key from the event id, recipient, channel and template, claim it in a send ledger with a unique constraint before calling the provider, skip keys already marked sent, and pass the same key to providers that accept one.

solid answer

~50 s

At-least-once delivery means the same `OrderShipped` event can arrive twice, so the send step itself must be idempotent. I build a **deterministic key**, for example a hash of `event_id + recipient + channel + template`, and insert it into a **send ledger** with a unique constraint as `PENDING` before calling the provider. A redelivery that finds `SENT` is acknowledged and dropped; one that finds a fresh `PENDING` backs off because another worker owns it. The hard case is a provider call that **timed out**: the SMS may or may not have gone out. If the provider accepts a client idempotency key I resend with the same key; otherwise I query status by my reference, or decide per message class whether a rare duplicate or a rare miss is worse. The key must never include a processing timestamp or random id, or every retry looks new.

code

pseudocode · 18 lines
pseudocode
key = sha256(event_id + recipient_id + channel + template_id)
if not ledger.insert_if_absent(key, state = PENDING, claimed_at = now):
    row = ledger.get(key)
    if row.state in (SENT, FAILED):
        ack(event)
        return
    if now - row.claimed_at < claim_timeout:
        return  // another worker owns it; broker redelivers later
    if not ledger.reclaim(key, row.claimed_at, now):
        return  // lost the takeover race
result = provider.send(message, idempotency_key = key)
if result.ok:
    ledger.mark(key, SENT, result.provider_id)
    ack(event)
else if result.permanent:
    ledger.mark(key, FAILED, result.reason)
    ack(event)
// transient: stay PENDING, retry later with the same key

go deeper

for a junior

Remember that brokers can deliver the same event twice and that a stable key recorded before sending is what stops the second SMS.

for a middle

Explain how the key is composed, why the insert must be atomic, what each ledger state means on redelivery, and why the key never includes attempt-specific values.

for a senior

Name the ambiguous timeout window and handle it per message class; tune claim and ledger retention timeouts against real provider and broker behaviour, and measure duplicates.

for a principal

Decide where duplicate protection is a platform guarantee versus a per-sender concern, and which message classes accept a rare duplicate versus a rare miss.

## Why duplicates happen A notification service usually consumes events from a message broker that guarantees **at-least-once** delivery: if a worker crashes, times out or fails to acknowledge, the broker hands the same event out again. Retries inside the worker add more chances. Without protection, a redelivered 'order shipped' event produces a second SMS, and a redelivered login event produces a second one-time code. Duplicates cost money on paid channels, annoy users and erode confidence in the product. The general theory of idempotent consumers is a separate topic; here the question is how to apply it to **sending messages**, where the side effect happens at an external provider you do not control. ## Designing the key The **idempotency key** must be identical for every retry of the same logical message and different for every genuinely new message. A good composition: - **event id** — the producer's unique id for the business event, stable across redeliveries; - **recipient id** — one event can notify both buyer and seller; - **channel** — push, email and SMS for the same event are distinct messages, and a fallback email must not be blocked by the push record; - **template or purpose** — one event can legitimately trigger two different messages. Things that must **not** be in the key: the time the worker picked the event up, a random id generated per attempt, or the provider's message id (which only exists after the send). A user who taps 'resend code' is a new request and should carry a new event id, so the key does not block intentional repeats. ## The claim-then-send flow 1. Compute the key and try to insert a ledger row in state `PENDING` with a claim time. A unique constraint on the key makes the insert atomic across workers. 2. If the insert fails, read the existing row. `SENT` or `FAILED` means the work is done: acknowledge the event. `PENDING` with a recent claim means another worker is sending: leave the event for redelivery. `PENDING` with an old claim means that worker probably died: take over with a conditional update on the old claim time. 3. Call the provider, passing the key as the provider's idempotency parameter when it supports one. 4. On success, mark `SENT` and store the provider message id; on a permanent rejection, mark `FAILED`; on a transient failure, leave `PENDING` so a later attempt retries with the **same key**. ```pseudocode key = sha256(event_id + recipient_id + channel + template_id) if not ledger.insert_if_absent(key, state = PENDING, claimed_at = now): row = ledger.get(key) if row.state in (SENT, FAILED): ack(event); return if now - row.claimed_at < claim_timeout: return if not ledger.reclaim(key, row.claimed_at, now): return result = provider.send(message, idempotency_key = key) if result.ok: ledger.mark(key, SENT, result.provider_id); ack(event) else if result.permanent: ledger.mark(key, FAILED, result.reason); ack(event) ``` ## The ambiguous window The ledger cannot close one gap on its own: the worker called the provider, the provider sent the SMS, and the response was lost to a timeout or a crash. The ledger still says `PENDING`. The options differ in cost: | Option | Effect | Needs | |---|---|---| | Resend with the same provider idempotency key | provider drops the repeat | provider support for client keys | | Look up status by your own reference | send only if the provider has no record | a provider lookup by client reference | | Resend blindly | rare duplicate | nothing | | Give up and mark unknown | rare missed message | nothing | Providers differ: some accept a client idempotency key or reference, others do not. When neither mechanism exists, choose **per message class**. A duplicate one-time code is mildly confusing but harmless, so resending is usually right; a duplicate marketing SMS costs money and goodwill, so giving up is often better. ## Operational details - Keep ledger rows at least as long as the broker's longest possible redelivery delay plus retry horizon; expiring them early silently reopens duplicates. - The claim timeout must exceed the provider call timeout, or a slow but healthy worker gets its message taken over and sent twice. - Count duplicates explicitly, for example by recording when a redelivery finds `SENT`, so a regression is visible. - The same ledger doubles as the source of truth for support and for delivery-status callbacks. The core idea to explain in an interview: the broker cannot give you exactly-once sending, so you make the send step idempotent with a stable key, and you name the one window that remains ambiguous and how you handle it.

  • Why must the claim timeout be longer than the provider call timeout?
    If a healthy worker is still waiting on a slow provider call when its claim looks stale, a second worker reclaims the key and sends again, so both calls may succeed. Setting the claim timeout comfortably above the call timeout plus processing time means a claim only looks stale once its worker has certainly given up or died.
  • A user taps 'resend code' twice. Should the idempotency key suppress the second code?
    No. Each tap is a new user request, so it should produce a new event id and a new key. Idempotency protects against the system repeating one request, not against the user asking again. Abuse of repeated requests is handled by per-user and per-number rate limits, which are a separate control from deduplication.
  • How long should send-ledger entries be kept?
    At least as long as an event can be redelivered or retried: the broker's maximum redelivery delay, the retry horizon and any replay window used during incident recovery. Expiring keys earlier silently reopens duplicates. Many teams keep them longer anyway because the same records serve support lookups and delivery-status callbacks.

saying these in an interview costs you the question

  • The message broker guarantees exactly-once delivery, so no dedup is needed.
  • A random id per attempt is a fine idempotency key.
  • Check whether the key exists, then insert it, as two separate steps.
  • The provider's message id can be used to detect duplicates before sending.
  • One key per event is enough even when several channels are sent.
  • A timed-out provider call means the message was not sent.