skip to content

An operation like 'increment a user's loyalty-points balance by 10' or 'send a welcome email' is not naturally idempotent - running it twice produces a different (or duplicated) outcome than running it once. Given that the underlying operation can be delivered more than once, what general design techniques make it safe to apply repeatedly?

level: seniorimportance: should knowfreq 50%

answer

  1. relative delta -> absolute keyed transaction (ledger pattern)
  2. conditional/versioned writes for 'set' operations
  3. dedup gate for unrepeatable side effects (email, SMS)
  4. natural idempotency self-heals; dedup gates don't

basics

~20 s

Instead of blindly repeating the action, redesign it so repeats do nothing extra: turn 'add 10 points' into 'record this specific +10 transaction with an ID', or check a 'have I already emailed this user for this event' flag before sending.

solid answer

~40 s

There are two broad families of fix. First, make the operation naturally idempotent by changing what it means to apply it: replace a relative update ('add 10') with an absolute, keyed one ('record this specific +10 transaction with ID T123', then derive the balance from all recorded transactions, or upsert keyed on T123 so re-applying it is a no-op). Second, when the operation truly can't be made naturally idempotent (an email send has an external, unrepeatable side effect), wrap it in a dedup gate: record 'sent welcome email for user X, event Y' before or atomically with the send, and check that record before sending again. Both techniques boil down to attaching a stable identity to the intended effect and making the system check that identity before re-applying it.

go deeper

for a junior

Should recognize, when prompted, that 'add 10' run twice gives a wrong answer, unlike 'set to X'.

for a middle

Should propose at least one concrete technique - either a transaction-ledger/upsert approach or a dedup-gate check before the side effect.

for a senior

Should articulate both techniques, know when each applies, and explain why natural idempotency self-heals in a way dedup gates don't.

for a principal

Should reason about picking dedup-key granularity correctly, cite a concrete real-world pattern or platform guidance, and discuss the residual risk when no stable event identity exists at all.

## Why some operations resist repetition Not every operation behaves like an idempotent `PUT` out of the box. - Some are **relative rather than absolute**: increment, append, decrement. - Some have side effects entirely outside the system's control (sending an email, firing a push notification), where 'running it twice' doesn't just risk a wrong number in a database - it risks an observable, irreversible duplicate action the user actually sees. Making these safe under at-least-once delivery requires deliberate design, not just wrapping them in a generic retry loop. ## The preferred technique: change the shape of the operation The first and generally preferred technique is to make the operation naturally idempotent by changing its shape from relative to absolute, and giving each intended effect a stable identity. Take 'add 10 loyalty points': instead of executing an unconditional increment on every delivery (which duplicates on redelivery, since running it N times adds 10N points), the system instead records a transaction with a stable ID derived from the triggering event: - an order ID - a message offset - anything unique to 'this specific reason for +10 points.' If that record is re-attempted with the same ID, a unique constraint on the transaction ID makes the second write a no-op, and the user's actual balance is computed as a derived sum, or maintained incrementally but only advanced once per unique transaction ID via an upsert. This pattern - record discrete, uniquely-identified facts rather than mutate a running total directly - is the same idea behind **event sourcing** and behind idempotent financial ledgers: you never mutate the balance in place from an unguarded delta, you append an identified entry and derive or upsert the total. ## The 'set' variant: conditional and versioned writes A closely related technique for operations that are naturally 'set' rather than 'add' is to use **conditional/versioned writes**: an update guarded by an expected-version check (optimistic concurrency control), or a write that's naturally absolute ('set the resource's shipping status to shipped' rather than 'advance the shipping status by one step'). Re-applying an absolute set is safe by construction, because applying it twice leaves the same end state as applying it once - the literal definition of idempotence. ## The fallback technique: gate the effect behind a dedup record The second technique, needed when the side effect genuinely cannot be made naturally idempotent - sending an email is the canonical example, since there is no way to upsert an email that's already left an SMTP relay - is to gate the operation behind an explicit **dedup record**, structurally identical to the idempotency-key pattern used for HTTP APIs. Before sending, the system checks (or atomically claims, using the same claim-pattern used for concurrent duplicate requests) a record keyed by something like `(user_id, event_type, event_id)` - 'welcome email for user U, triggered by signup event S.' - If that record already exists, skip the send. - If not, atomically claim it, send the email, and mark it sent. This doesn't make sending inherently repeatable; it prevents the system from attempting to repeat it in the first place, by remembering that the intended effect already happened. ## Which one to reach for The trade-off between these two techniques is where correctness lives. | Technique | Behavior when the bookkeeping is lost | | --- | --- | | **Natural idempotency** - the transaction-ledger / absolute-write approach | Generally more robust, because it tolerates gaps in the dedup bookkeeping: even if a 'processed' marker is somehow lost, re-deriving state from uniquely-identified facts self-heals, since replaying the same fact twice is harmless by construction. | | **Dedup-gated operations** - the email-send approach | Strictly dependent on the gate itself being correct and durable; if the 'already sent' record is lost, corrupted, or has too short a TTL, the side effect will genuinely repeat, with no self-healing possible after the fact - you cannot un-send an email. | This is why, wherever the operation allows it, engineers prefer to redesign toward natural idempotency rather than relying purely on a dedup gate, and reserve the dedup-gate approach for operations where no natural-idempotency redesign is possible. ## The same advice from a platform A concrete real-world instance: **AWS's** own guidance for Lambda functions triggered by at-least-once event sources explicitly recommends this two-pronged approach - - use idempotent operations (conditional writes, upserts) wherever the downstream effect allows it; - fall back to storing processed-event IDs with a TTL as an explicit dedup gate for effects, like third-party API calls or notifications, that can't be made naturally idempotent.

  • Why is 'record a uniquely-identified transaction and derive the total' generally more robust than a dedup gate for something like a points balance?
    Because if the dedup record is ever lost, corrupted, or expires, a ledger-based system self-heals - replaying the same uniquely-identified transaction is still harmless, since the unique constraint or upsert on that transaction ID makes the replay a no-op. A pure dedup gate has no such backstop: if its 'already applied' record disappears, the underlying add-10 operation will genuinely re-apply and there's no way to detect or reverse that after the fact.
  • Can you always redesign a relative update into an absolute, keyed one? What if there's no natural unique ID for the triggering event?
    Not always - if the event genuinely lacks any stable identity, you either need the upstream system to add one, or you fall back to a dedup gate using whatever partial signal is available (a content hash plus a time bucket), accepting a small residual risk of either false-positive dedup (dropping a legitimately distinct event) or false-negative dedup (missing a true duplicate).
  • For the welcome-email example, why use (user_id, event_type, event_id) as the dedup key instead of just user_id?
    Keying only on user_id would prevent ever sending that user a second welcome email even for a legitimate reason, and it conflates unrelated events; keying on the specific triggering event lets the system distinguish 'redelivery of the same signup event' (dedup, skip) from 'a different, legitimate reason to email this user' (not a duplicate, should send).

It's like the difference between 'add one more log entry to this specific line' (which is safe to repeat, since re-writing the same line changes nothing) versus 'shout an announcement out the window' (which can't be undone once shouted, so the only real protection is a doorman keeping a checklist of announcements already made, and not letting anyone shout the same one twice).

saying these in an interview costs you the question

  • Proposes wrapping a relative update (+10) in a generic retry/idempotency-key mechanism without addressing that the operation itself is still non-idempotent
  • Treats 'send an email exactly once' as solvable purely by making the email API call idempotent, ignoring that the send itself is the irreversible step
  • Doesn't distinguish between operations that can be redesigned to be naturally idempotent versus ones that need a dedup gate
  • Uses a dedup key too coarse (e.g., just a user ID) to distinguish legitimate repeats from true duplicates

context