skip to content

A payment-processing client automatically retries a 'charge card' API call after a timeout, without knowing whether the original request actually succeeded on the server before the timeout occurred. What can go wrong, and what design makes it safe to retry this kind of operation?

level: middleimportance: must knowfreq 80%

answer

  1. timeout = ambiguous outcome, not 'failed'
  2. idempotency key = one per logical operation, reused across retries
  3. server persists key -> dedups replay
  4. naturally idempotent: PUT absolute state; not: POST create, increment
  5. Stripe Idempotency-Key header

basics

~20 s

The first request might have actually succeeded even though the client never got a response, so retrying it blindly could charge the customer twice. Making the operation idempotent — using a unique request ID so the server can recognize and ignore a duplicate — prevents that.

solid answer

~50 s

A timeout tells the client nothing about whether the server processed the request — it could have been lost before reaching the server, processed and the response lost, or still in flight. Blindly retrying a non-idempotent operation like 'charge $50' risks applying the side effect twice. The standard fix is an idempotency key: the client generates a unique token per logical operation (not per retry attempt) and sends it with every attempt; the server persists a record of keys it has already processed and, on seeing a duplicate key, returns the original result instead of re-executing the side effect. This shifts the correctness burden from 'never retry' to 'retry safely,' and is a prerequisite for retry policies on any operation with a side effect — reads are naturally idempotent and need no such mechanism, but writes generally do unless they're naturally idempotent (e.g., a PUT that sets an absolute value).

go deeper

for a junior

Should recognize that retrying an operation like 'charge a card' twice could cause a double charge, and that some mechanism is needed to prevent it.

for a middle

Should be able to describe the idempotency-key mechanism at a high level: client generates one key per logical operation, server dedups on it.

for a senior

Should distinguish naturally idempotent operations (PUT absolute state) from ones needing engineered idempotency (POST create, increments), and know the key must be reused across retries, not regenerated.

for a principal

Should reason about the atomicity of the dedup check itself (check-then-act races), key retention windows relative to retry timeouts, and how this interacts independently with backoff/jitter policy design.

## Why a timeout is ambiguous When a client sends a request and the connection times out before a response arrives, the client is left in genuine ambiguity: it cannot distinguish between - the request never having reached the server, - the request being processed but the response getting lost on the way back, - or the request still executing on the server at that very moment. All three are indistinguishable from the client's point of view — a timeout is silence, not a 'no.' This ambiguity is exactly why retrying is attractive (the request might just have been a lost packet that never landed) and exactly why retrying is dangerous (the request might have landed and executed, and a naive retry would execute it again). For an operation with no side effect, like fetching an account balance, replaying it is harmless. For an operation with a side effect, like charging a credit card, placing an order, or incrementing a counter, replaying it can be actively harmful: the customer gets charged twice, two orders get placed, or a counter drifts upward every time a network blip happens to coincide with a retry. ## Naturally idempotent, or engineered to be **Idempotency** is the property that makes an operation safe to execute more than once with the same net effect as executing it once. - Some operations are **naturally idempotent** by their semantics: an HTTP `PUT` that sets a resource to an absolute state produces the same end state whether it runs once or five times, and a `SQL UPDATE ... SET status = 'shipped'` is idempotent for the same reason. - Others are **naturally non-idempotent**: an HTTP `POST` that creates a new resource, or an operation that increments a value, is not naturally idempotent, since running it N times produces N times the effect. ## The idempotency-key mechanism For this second category, idempotency has to be engineered, and the standard mechanism is an **idempotency key**. 1. The client generates a unique identifier once per logical operation — not once per HTTP attempt — and attaches it to every retry of that operation. 2. The server, on receiving a request with an idempotency key, checks a persisted record of keys it has already handled. 3. If the key is new, it executes the operation and stores the key alongside the result. 4. If the key is a duplicate, it skips re-executing the side effect and returns the previously recorded result instead. This converts 'at-least-once delivery is fine, because duplicates are detected and absorbed server-side' from 'retries risk duplication.' ## The trade-offs on both sides The trade-offs sit on both sides. - **On the client side**, consistently reusing one idempotency key across all attempts of a single logical operation (rather than minting a new key per retry, which would defeat the purpose entirely) adds bookkeeping the client must get right — the key has to survive across the retry loop, and if the client itself crashes and restarts mid-operation, it needs a durable way to remember or regenerate the same key for that same logical intent. - **On the server side**, tracking idempotency keys requires persistent storage with its own retention policy — keys can't be kept forever, so there's a window after which a 'duplicate' would no longer be recognized as one. There's also a subtler correctness requirement: the idempotency check and the side effect it's guarding must be atomic with respect to each other (typically via a database transaction or a unique constraint on the key), or a race between two near-simultaneous retries can both pass the 'have we seen this key' check before either records it, defeating the mechanism — a classic check-then-act race. ## Failure modes, and where it shows up Failure modes without proper idempotency handling are visible directly in financial and e-commerce systems: duplicate charges, duplicate order confirmation emails, or double-decremented inventory are textbook symptoms of retries applied to non-idempotent operations, and they tend to correlate suspiciously with periods of elevated network latency or partial outages. Stripe's Idempotent Requests API is a widely cited concrete real-world example: it requires clients to pass an `Idempotency-Key` header on state-changing requests like creating a charge, documenting that the key should be generated once per logical operation and reused across retries, and that Stripe stores the result keyed by that value for a bounded retention period so a retried request returns the original charge's result rather than creating a second charge. The general lesson generalizes beyond payments: any retry policy applied to a write operation needs an explicit answer to 'is this operation idempotent, and if not, what mechanism makes it safe to retry' before backoff and jitter parameters are even worth tuning.

  • If a client mints a brand-new idempotency key on every retry attempt instead of reusing one key per logical operation, does the mechanism still work?
    No — that defeats the purpose entirely, since the server would see each retry as a distinct, never-before-seen key and execute the side effect again each time. The key must be generated once when the logical operation begins and reused unchanged across every retry, which means the client has to hold onto it (in memory or durably) for the lifetime of the retry loop.
  • Why isn't a simple 'check if this key exists, then insert if not' server-side implementation sufficient on its own?
    Because that's a check-then-act sequence, and two near-simultaneous retries can both read 'key does not exist' before either one writes it, letting both proceed to execute the side effect. The check and the insert need to be made atomic, typically via a unique constraint on the key column combined with catching the constraint violation, or via a transaction with appropriate isolation.
  • Does making an operation idempotent eliminate the need for backoff and jitter on its retries?
    No, they solve different problems — idempotency makes it safe to retry without corrupting state, while backoff and jitter control when and how aggressively those safe retries are sent so they don't overload the dependency or synchronize across many clients. A fully idempotent operation retried in a tight, unjittered loop is still just as capable of causing a retry storm as a non-idempotent one.

Like mailing a request slip with a unique confirmation number stapled to it: if the postal service loses the reply and you resend the exact same numbered slip, the recipient's ledger recognizes the number and just re-confirms it was already handled, instead of paying out twice.

saying these in an interview costs you the question

  • Assumes a timeout means the request definitely failed and had no server-side effect
  • Treats all HTTP methods as safe to retry without distinguishing idempotent from non-idempotent ones
  • Suggests generating a new idempotency key on every retry attempt
  • Doesn't recognize the check-then-act race in naive idempotency-key storage
  • Thinks idempotency alone eliminates the need for backoff/jitter
  • Assumes idempotency keys can be retained indefinitely with no retention/expiry consideration

context