Some operations like sending an email or charging a card can't be made naturally idempotent. How do you make a consumer effectively exactly-once for those, conceptually?
answer
- stable key: business id / event UUID / TPO
- guard (check) then act then record
- effect + 'done' marker must commit together
- push idempotency into the provider (Idempotency-Key)
- key store needs TTL / bounded retention
basics
~20 sIf the side effect can't be a simple overwrite, you make it idempotent by giving each unit of work a stable unique id and recording 'already done' so a redelivered record is recognized and skipped. The action plus the 'done' record must commit together.
solid answer
~50 sNaturally-idempotent writes (upserts, PUTs) absorb duplicates for free. Inherently observable effects — emails, payments, outbound webhooks, counters — cannot. For those you derive a stable idempotency key from the record (a business id, or the topic-partition-offset, or an event UUID) and guard the action: before doing it, check whether that key was already processed; after doing it, record the key. The subtlety is atomicity — the side effect and the 'recorded' marker must commit together (or the side effect must accept the key, e.g. Stripe's Idempotency-Key header), otherwise a crash between them either repeats the action or wrongly marks it done. When the action is in a transactional DB, write its result and the processed-key in one transaction. When it's a remote API, push idempotency into that API. (Concrete store choices — unique DB key, Redis — belong to a sibling topic; here it's the conceptual pattern.)
go deeper
Know that non-idempotent effects need a unique key plus a 'check-then-skip' guard, and that you record what you've done.
Explain key sources, the guard/record steps, and that the effect and the marker must commit atomically.
Reason about the atomicity window, when to push idempotence into the provider vs a shared transaction, and key-store retention.
Define an org-wide convention for idempotency keys and atomic-commit patterns across heterogeneous sinks, including replay-safety and GC.
**Why some effects resist idempotence.** A naturally-idempotent operation reaches the same final state regardless of how many times it runs: `SET status='PAID'`, `PUT key=value`, `DELETE`, `UPSERT`. An *observable* or *accumulating* effect does not: sending an email produces a new email each time; `balance += amount` accumulates; charging a card moves money each time; appending to a log grows it. Redelivery (which at-least-once guarantees can happen) turns these into real, visible duplicates. **The dedup pattern (conceptual).** Convert the non-idempotent effect into an idempotent one by *remembering what you've already done*: 1. **Derive a stable idempotency key** for each unit of work. Good sources: a business identifier (orderId), an event UUID carried in the record, or the `topic-partition-offset` triple (unique per record, but tied to physical position — re-keyed topics break it). The key must be the same on the redelivery as on the first attempt. 2. **Guard:** before performing the effect, check whether this key is already recorded as done; if so, skip the effect (and still advance the offset). 3. **Record:** after performing the effect, persist the key so future duplicates are recognized. **The atomicity subtlety — this is the whole game.** Steps 2/3 only work if 'do the effect' and 'record the key' are atomic, or if the effect itself accepts the key. Three patterns: - **Shared transactional store:** the effect *is* a DB write, so write the business result and insert the idempotency key in the *same* DB transaction. A unique constraint on the key turns a duplicate into a constraint violation you can swallow. Now the side effect and its dedup record are one atomic commit. - **Push idempotence into the remote system:** for a payment or HTTP call, send the key to the provider (Stripe `Idempotency-Key`, or an idempotent PUT). The provider dedups; you don't need a local marker to be atomic with the call. - **Two-system, can't be atomic (e.g. send email + mark done in DB):** there's an unavoidable window. Choose ordering by which failure you tolerate: mark-then-act risks a lost email (at-most-once for that effect); act-then-mark risks a duplicate email (at-least-once). For low-stakes effects act-then-mark is typical; for money you push idempotence into the provider instead. **Garbage collection / retention.** A processed-key store grows unbounded; in practice you bound it (TTL, or only need to cover the redelivery window — roughly retention/rewind horizon). Too-aggressive expiry re-opens the duplicate window. **Ordering vs dedup.** Dedup answers 'did I already do this one?'; it does not by itself enforce ordering. With Kafka's per-partition ordering plus single-threaded-per-partition processing you usually get ordering for free, but parallel/async processing can reorder and you must reason about that separately. The takeaway: you can always *manufacture* idempotence around a non-idempotent effect by adding a remembered key — the engineering work is making 'effect' and 'remembered' commit together, or delegating dedup to the downstream system.
- Why is topic-partition-offset a risky choice of idempotency key?It's tied to the physical position in a specific topic. If records are re-published, the topic is recreated, or the data is replayed from another topic, the same logical event gets a different offset and dedup fails. A business id or event UUID is stable across replays.
- You write the result to a DB and then send an email; the email send isn't transactional. What ordering issue arises?The DB write and email can't be atomic. If you mark-done before sending and crash, the email is lost; if you send before marking and crash, a duplicate email goes out on retry. You pick the tolerable failure, or move dedup into the email provider.
saying these in an interview costs you the question
- Claiming you can make charging a card idempotent without any key or downstream support — you can't.
- Recording the 'processed' marker in a store unrelated to the side effect and assuming the two are atomic.
- Using topic-partition-offset as the key while planning replays or topic recreation.
- Letting the processed-key store grow forever with no TTL/bound.
- Conflating dedup (did I do this?) with ordering (in what order?).