How do you make a GraphQL mutation safe for a client to retry?
answer
- Ask what the client cannot distinguish
- Retry unit versus atomic unit
- One stable key per logical action
- Store the key with the write
- Replay the recorded payload
basics
~20 sHave the client send a stable key it generates once per logical action and reuses on every retry. The server records that key with the outcome in the same transaction as the write, and on a repeat key returns the recorded result instead of writing again.
solid answer
~50 sThe dangerous case is the ambiguous failure: a timeout or a dropped connection leaves the client unable to tell whether the write happened, whether the response was merely lost, or whether a multi-field document half-applied. Blind retry then risks a second charge. The fix is an **idempotency key** — a value the client generates once for the action and resends unchanged on every attempt, carried as an argument on the mutation field. The server persists the key alongside the outcome inside the same transaction as the write; a second request with the same key skips the work and replays the stored payload, including any ids it generated the first time. Nothing in GraphQL specifies this: there is no reserved argument, no directive and no header for it, so you name and document the convention yourself. Retry ambiguous transport failures, never deterministic field errors, which will just fail again.
code
graphql · 7 linesmutation Capture($key: String!) {
capturePayment(
orderId: "ord_48213"
amountCents: 4183
idempotencyKey: $key
) { receiptId capturedAt }
}go deeper
Recall the core idea: a retry must carry the same key the first attempt carried, and the server uses that key to notice it has already done the work. Know that mutations are not safe to repeat by default.
Explain the ambiguous failure and the mechanics: key generated once per action, stored with the outcome in the same transaction, replayed rather than re-executed. Be able to contrast an absolute write with a relative one.
Demonstrate production judgement — concurrent attempts with one key, retention windows, replaying the original payload, and the discipline of retrying ambiguous failures only. Connect it back to sending one root mutation field per request.
Own the convention across the whole graph. Decide the key's name, scope and lifetime once, decide which mutations must honour it, and be able to argue why a per-team convention is worse than none at all.
## The ambiguous failure A client sends `capturePayment` for order `ord_48213` and gets nothing back: a socket reset, a proxy timeout, a phone that lost signal. Three worlds are consistent with that silence. 1. The request never reached the server. Nothing happened. 2. The request executed and the response was lost on the way back. The charge went through. 3. The document carried more than one root field and half of them applied before the connection died. The client cannot distinguish them, and the user is standing at the counter. If it retries blindly and the world was (2), the guest is charged twice. If it gives up and the world was (1), the restaurant is not paid. One team measured this the hard way after a network blip: 19 duplicate captures in a twelve-minute window, every one of them a client doing the sensible-looking thing. ## Why GraphQL sharpens the problem A retry re-sends the whole document. If the document had three root mutation fields and the failure landed on the third, the retry re-runs the first two — which succeeded. You get a second discount and a second void attempt on top of the capture you were actually trying to redo. The retry unit and the atomic unit are different sizes, and that mismatch is the bug. Which gives the first rule, before any machinery: **send one root mutation field per request**, so the thing the client retries is exactly the thing the server made atomic. ## The idempotency key The key is a value the client generates **once per logical action** — not once per attempt — and resends unchanged with every retry. A random identifier minted when the cashier taps "charge" is right; one minted inside the retry loop is useless, because each attempt then looks like a new action. ```graphql mutation Capture($key: String!) { capturePayment( orderId: "ord_48213" amountCents: 4183 idempotencyKey: $key ) { receiptId capturedAt } } ``` The server side has one property that carries the whole design: ```pseudocode resolve capturePayment(orderId, amountCents, idempotencyKey): begin transaction prior = idempotency.lockAndRead(callerId, "capturePayment", idempotencyKey) if prior.exists: commit return prior.payload # replay, do not write receipt = charge(orderId, amountCents) idempotency.write(callerId, "capturePayment", idempotencyKey, payload = { receiptId: receipt.id, capturedAt: now() }) commit # key and write commit together return payload ``` **The record and the write must commit together.** If the key is stored after the transaction, a crash in between leaves a charge with no key, and the retry charges again — which is the very failure you set out to prevent. Details that matter in production: * **Scope the key** to the caller and the field, not globally, so two tenants that happen to mint the same value do not collide. * **Store the payload**, not just a flag. The replay should return the same receipt id the first attempt returned, so the client's second read matches its first. * **Give it a retention window** comfortably longer than any client's retry budget — keeping keys 26 hours against a 90-second budget costs little and closes the gap left by a phone that reconnects much later. * **Handle concurrent attempts.** Two in-flight requests with one key is normal under aggressive retry. A unique index on the key plus a row lock makes the second wait for the first, or fail cleanly rather than charging. ## Shapes that need no key Some writes are idempotent by construction, and choosing those shapes is cheaper than the machinery: * **Absolute rather than relative.** `setPartySize(seats: 6)` run twice leaves six. `addSeats(delta: 2)` run twice leaves four. * **Client-supplied ids on creation.** If the client mints the order line's id, a duplicate create collides on the primary key and can be reported as "already exists" instead of inserting a second row. * **State transitions guarded by the current state.** "Mark ready if pending" is safe to repeat; "mark ready" with no guard may resurrect an order someone else has since cancelled. ## What not to retry Retry **ambiguous** failures — timeouts, resets, gateway errors — where you genuinely do not know the outcome. Do **not** retry a field error you received in a complete response. A permission denial, a validation failure or a business rejection is deterministic: the second attempt fails identically, you have burned a round trip, and an automatic retry loop can bury the real defect under noise. If you saw a response, you know what happened; retry is for when you do not. ## Specified or conventional? Conventional, entirely. GraphQL defines no reserved argument name, no directive and no metadata slot for idempotency, and the GraphQL over HTTP specification defines no header for it either. Pick a name, apply it consistently across every mutation that writes money or sends messages, and document it as part of your API contract — because a convention only half the clients follow protects nobody.
- Why must the idempotency record and the write commit in the same transaction?Because the gap between them is a failure window. If the charge commits and the process dies before the key is stored, the retry finds no key and charges again — exactly the duplicate you were preventing. Committing them together makes "the write happened" and "the key is recorded" a single fact, so the retry can trust what it reads.
- Where does the key belong — a field argument or a transport header?Either can work, and neither is specified. An argument keeps it inside the document, so it survives any transport and is visible to schema linting and to server-side validation. A header keeps it out of the schema but ties the convention to one transport and hides it from anything that only sees the operation. Most graphs put it in the schema for that visibility.
- A client received a complete response containing a permission error. Should it retry?No. That is a deterministic outcome, not an ambiguous one — the same request will be refused the same way. Retrying wastes a round trip and, in an automatic loop, hides the underlying defect behind traffic. Retry when you do not know what happened; surface the failure when you do.
- How long should keys be retained?Longer than the longest retry window any client could plausibly use, which is usually far longer than the one you designed for. A budget of about ninety seconds paired with a retention of a day or so leaves a wide margin at trivial storage cost. Retire keys with a sweep, and be aware that once a key expires a very late retry becomes a fresh action again.
saying these in an interview costs you the question
- Generates a new key inside the retry loop
- Stores the key after the write commits
- Retries deterministic field errors automatically
- Assumes mutations are idempotent by default
- Returns a bare flag instead of the stored payload
- Believes the specification defines an idempotency header