A GraphQL endpoint answers every read and write as a POST to one URL. How do you stop an automatic retry duplicating a write?
answer
- Every operation looks alike from outside
- The method cannot tell you what it is
- A timeout proves nothing about the commit
- Whoever wrote the operation owns the retry
- Repeat-safety belongs in the mutation input
basics
~20 sNothing outside a GraphQL request body distinguishes a read from a write — same URL, same method, same headers — so no intermediary can retry safely. Move retry decisions to the code that authored the operation, or make the write repeat-safe.
solid answer
~50 sThe GraphQL over HTTP shape puts the operation type inside the request: a mutation and a query are both `POST /graphql` with a JSON body, and only parsing that body reveals which is which. Every generic retry mechanism — a client library's default policy, a load balancer, a sidecar proxy — decides from the method and status alone, so it will happily replay a mutation whose response was merely lost. In a restaurant ordering graph that surfaces as two identical orders from one customer tap. Three fixes, in order of value. Disable blanket retries on the graph endpoint and retry only in the layer that knows what it sent. Make writes repeat-safe by taking a client-generated key as a mutation argument, recording it, and returning the first result on a replay. And send read-only operations over GET where they fit, so the method itself tells intermediaries the request is safe to repeat.
code
json · 11 lines{
"query": "mutation PlaceOrder($input: PlaceOrderInput!) {\n placeOrder(input: $input) {\n order { id state }\n }\n}",
"operationName": "PlaceOrder",
"variables": {
"input": {
"restaurantId": "r-4417",
"itemIds": ["item-2291", "item-2308"],
"requestKey": "ord-intent-83112"
}
}
}go deeper
Understand the basic fact first: every GraphQL operation goes to the same URL with the same method, so nothing outside the body says whether a request reads or writes. That is why a blind retry is risky.
Explain the mechanism end to end: where the operation type lives, which layers make retry decisions from the method alone, and how a lost response after a committed write produces a duplicate.
Show you have operated this. Name the layers that retry by default, describe how you would trace a duplicate write back to one of them, and design the repeat-safe mutation contract rather than just switching retries off.
Own the policy across services: who is allowed to retry a graph call, what every write-bearing mutation must accept as an intent key, and how that is enforced in schema review rather than rediscovered after an incident.
## Why the shape causes this Every request to a GraphQL endpoint looks the same from the outside. Same URL — one endpoint serves the whole schema. Same method — a POST, because the body is JSON. Same headers. The only thing that distinguishes `menuBoard` from `placeOrder` is a word inside a JSON string inside the body. HTTP's own machinery for deciding what may be repeated works entirely on the method. That machinery is everywhere: default retry policies in HTTP client libraries, connection-level retries in load balancers and sidecar proxies, user-driven refreshes. Point any of it at a GraphQL endpoint and it is working blind — it sees POST, applies whatever POST policy it was configured with, and cannot tell a repeat-safe read from a write. The failure has a specific shape. A client posts `placeOrder`. The server validates, charges, commits the row, and starts writing the response. The connection drops, or the client's deadline expires first. The client sees a timeout, which tells it exactly nothing about whether the server committed. If a retry policy sits underneath it, a second `placeOrder` executes and the customer has two orders. In an internal graph of eleven services, where the request crosses several hops each with its own defaults, the retry usually turns out to have come from a layer nobody remembered configuring. ## Ranking the fixes **Stop retrying from layers that cannot classify.** The most reliable change is also the least clever: turn off automatic retries for the graph endpoint everywhere below the code that composed the operation. That code knows the operation type, knows whether the mutation is repeat-safe, and can back off intelligently. A proxy layer cannot know any of it. This is a configuration and review problem more than an engineering one, and it is the change that actually prevents the incident. **Make writes repeat-safe in the schema.** Retries do not disappear; users press buttons twice and networks fail mid-flight. The durable answer is to let the mutation accept a client-generated key — a value the client creates once per user intent and reuses across every attempt — and have the server record the key with the outcome. A replay finds the recorded outcome and returns it instead of performing the work again. Two design points matter. The key belongs in the schema as an input field, not in a header: it is then typed, visible in the document, validated like any other input, and it survives being carried by any transport. And it must be generated at the point of *intent* — when the customer taps Place Order — not per HTTP attempt, or every retry invents a fresh key and nothing is deduplicated. **Let reads say they are reads.** Where a document is short enough to fit in a URL, sending query operations over GET restores the signal that the transport lost: the method itself now says this request is safe to repeat, and generic retry logic becomes correct rather than lucky. It does nothing for mutations, which is why it is the third fix rather than the first. ## Things that look like fixes and are not Appending the operation name to the URL as a parameter, or copying it into a header, makes traffic legible in logs and dashboards. That is worth doing for other reasons, but no generic retry policy is going to learn from it what is safe to repeat, and the body remains the authority on what actually executes. Giving mutations their own URL would restore the distinction, at the price of breaking the single-endpoint assumption every client and tool is built on. It is worth being able to discuss, and worth rejecting for a general graph. Relying on the server to notice "this looks like a duplicate" without a client-supplied key means guessing from field values and timing. Two genuinely separate identical orders exist — the same customer really can order the same coffee twice in a minute — so a heuristic either drops real writes or misses replays. ## What an interviewer is listening for First, that a timeout is not evidence of anything: the request may have fully succeeded and only the response was lost. Second, that the responsibility sits with whoever can read the operation, which means the client, not the transport. Third, that repeat-safety is a schema design decision — the key is part of the mutation's input contract, agreed with clients, not a piece of infrastructure bolted on afterwards.
- Would sending queries over GET solve the whole problem?Only half of it. Reads become legible to intermediaries, so generic retry logic stops being blind about them, and that is a real gain. Mutations still travel as POST bodies to the same URL, so nothing below the client can distinguish one write from another. GET also constrains document length, so only short read documents qualify.
- Where should a client-generated key for a repeat-safe mutation live?As an input field on the mutation, so it is typed, validated, visible in the document, and part of the contract clients agree to. A header is invisible to the schema and to anything that only sees the body, and it does not survive being carried by another transport. Generate the key once per user intent, not per HTTP attempt.
- The client saw a timeout. Is that evidence the mutation did not run?No. A timeout says the response did not arrive within the deadline; it says nothing about server-side execution. The mutation may have validated, committed and been half-way through serialising a response. Treating a timeout as failure is how duplicate writes get created, which is why the correct reaction is to re-drive the same intent key rather than to issue a fresh write.
It is a mailroom that forwards every envelope through the same slot: nothing on the outside says whether the letter inside is a question or an instruction to ship goods, so the mailroom can never safely post a second copy.
saying these in an interview costs you the question
- Assumes a timeout means the mutation never ran
- Says retrying a POST is harmless
- Expects a proxy to parse the GraphQL body
- Puts the deduplication key in a header only
- Generates a fresh key on every retry attempt
- Trusts the server to guess which writes are duplicates