A retried GraphQL mutation leaves a duplicate row in the client cache after an optimistic update. What went wrong?
answer
- A timeout is not a failure
- The prediction lives in a droppable layer
- Nothing rolled back because nothing failed
- Two commits, two ids, two references
- Idempotency key per user intent
basics
~20 sThe first attempt reached the server and committed; only its response was lost. The retry created a second object with a second id, so the two real responses wrote two entities and two list references. Rollback cannot help - nothing failed.
solid answer
~50 sAn optimistic update writes a predicted result into the store immediately and keeps it in a layer that is dropped when the real response arrives or the mutation errors. That machinery is not the fault here. A request that times out or drops its connection tells the client nothing about whether the server committed; if it did, a retry runs the write a second time, and a non-idempotent mutation happily produces a second object. The client then receives one genuine success, writes the entity, appends it to the list - and the earlier commit is discovered later, or arrives on another screen. The fix is at the write, not the cache: carry a client-generated idempotency key on the mutation input so the server returns the original result for a repeat, and never auto-retry a mutation the way you would retry a query.
code
graphql · 14 linesmutation HoldSeats($input: HoldSeatsInput!) {
holdSeats(input: $input) {
seatHold {
id
seatIds
expiresAt
}
event { id }
}
}
# variables:
# { "input": { "requestId": "8f3c1a94-...", "eventId": "ev-4471",
# "seatIds": ["B-14", "B-15"] } }go deeper
Know what an optimistic update is: the client writes the expected result before the server answers, and undoes that write if the mutation errors. Knowing it exists is enough at this level.
Explain the layer model - prediction stacked over confirmed data, dropped on success or error - and why an inverse edit or a snapshot restore is the wrong rollback when other results landed in the meantime.
Show the diagnosis: the cache faithfully mirrored a server that really did commit twice. Separate observed failure from unknown outcome, name the retry policy as the cause, and propose an idempotency key with a retention window.
Own the standing rule for the client platform - which operations may be retried at all, where intent keys are generated and persisted, and how long the server keeps them - so that no individual feature has to rediscover this incident.
## What an optimistic update actually is Before the request leaves, the client writes a *predicted* result into the store as if the server had already answered: the seat hold exists, the button is done, the row is on screen. The prediction is kept apart from confirmed data - conceptually a layer stacked over the canonical entities - so that reads see prediction-over-truth while the request is in flight, and the layer can be dropped whole when the real answer lands. On success the server's response is merged into the canonical data and the layer goes away; on error the layer goes away and the canonical data was never touched. Two properties of that design matter for the incident. - Rollback is **removal of a layer**, not an inverse edit and not a snapshot restore. Restoring a snapshot would also throw away every query result and every other mutation that landed while the request was in flight. Dropping the layer and re-deriving reads leaves those intact. - Rollback only fires on a **failure the client observed**. A write the client never learned about is not a failure to roll back; it is a success the client is unaware of. ## The incident A ticketing app, `holdSeats`, one fan taking two seats in a 1,240-seat room. The client writes an optimistic hold and shows the seats as held. The request crosses a flaky mobile link, the server commits the hold, and the response is lost on the way back. The client sees a timeout at 8 s, and the request layer retries once - as it was configured to do for any operation. The second attempt succeeds and returns `SeatHold:sh-5502`. The client merges it, appends the reference to the event's holds list, drops the optimistic layer, and shows a clean confirmed state. Meanwhile `SeatHold:sh-5471` from the first attempt exists on the server. The moment anything refetches - a navigation, a subscription push, another device - the list arrives with two holds and four seats are gone from a room that only ever agreed to two. Note what this is *not*. The cache never diverged from the server's truth; it faithfully reflected a server that really did have two holds. Chasing this in the store is chasing the symptom. ## Where it actually goes wrong, in order 1. **A timeout was read as a failure.** A client timeout, a dropped socket and a 502 from an intermediary all mean *unknown*, not *did not happen*. Only an explicit application error in the `errors` entry of a completed response tells you the write did not take effect - and even then only if the server's error semantics say so. 2. **A non-idempotent write was retried automatically.** Retrying a query is free; retrying a mutation is a second write unless the write is idempotent by construction or by key. Transport-level retry policies that do not distinguish operation type are the usual culprit. 3. **The mutation had no idempotency key.** With one, the second attempt is a lookup, not a write. ## The fix that holds Carry a key the client generates once per user intent and reuses across every retry of that intent: ```graphql mutation HoldSeats($input: HoldSeatsInput!) { holdSeats(input: $input) { # input.requestId is the idempotency key seatHold { id seatIds expiresAt } event { id } } } ``` The server records the outcome under that key and, on a repeat, returns the original result rather than committing again. The client is then free to retry, because a retry is now a read of a decision already made. Two constraints people miss: the key must survive an app restart if the retry can, so it belongs with the queued intent rather than in a component's memory; and the server must remember the key at least as long as the client's whole retry window, which is a storage decision, not a cache one. Alongside that: exclude mutations from any blanket retry policy, and let a failed write surface to the user with an explicit "try again" rather than reissuing silently. ## The other optimistic-update traps worth naming - **Predicting a server-assigned identity.** The prediction needs *some* key, and a temporary one is fine, but the real response arrives with the server's id - so the prediction is replaced, not merged, and any list reference must be swapped rather than duplicated. Predicting an id that looks real is how a phantom entity survives in the store forever. - **Predicting values the server owns.** Price after fees, queue position, remaining inventory: predict these and the UI flickers to a different number on confirmation, which is worse than a spinner. - **Concurrent optimistic writes.** Two in-flight predictions over the same entity must both be re-applied over the confirmed data as each resolves, in a defined order. This is exactly why layered rollback beats undo-by-inverse-edit. - **Refetching before the write settles.** A refetch issued in parallel with the mutation can return pre-write data and confirm it over the prediction, making the row flicker out and back. Sequence the refetch after the mutation's response. ## The senior signal Say plainly that the cache is not the bug. Show that you separate *observed failure* from *unknown outcome*, that you know the retry policy is where the duplicate was born, and that the durable fix is an idempotency key on the write with a defined retention window - not a cleverer rollback.
- Why is restoring a snapshot of the store the wrong way to roll back an optimistic update?Because the store moved on while the request was in flight. A snapshot restore discards every query result, subscription push and other mutation that landed in that window, so a failed write silently reverts unrelated, correct data. Keeping the prediction in its own layer over the canonical entities means rollback is dropping that layer and re-deriving reads, which leaves everything else exactly as the server left it.
- Where should the idempotency key come from, and how long must the server remember it?The client generates it once per user intent - one key for the whole retry sequence, not one per attempt - and it must survive whatever the retry survives, so it belongs with the queued intent rather than in component state. The server stores the outcome under that key for at least the client's full retry window plus a margin: long enough that a repeat is always answered from the record, short enough to bound the storage.
- When would you skip the optimistic update entirely for a write like this?When the server owns the values - assigned seats, queue position, price after fees - because a prediction that is later corrected is worse than a brief spinner. Also when failure is genuinely likely, as with contended inventory, since most predictions would be rolled back in front of the user. Optimistic updates pay off for high-frequency, near-certain writes whose result the client can compute exactly.
- The mutation succeeded and the response merged cleanly, yet another open screen still shows the old state. Is that the same bug?No. That is ordinary invalidation: the write landed on entities that screen does not read, or changed a list's membership rather than an entity's values, so nothing it references moved. The remedy is the usual one - repair the affected list or refetch its operations. Duplicate rows from a retry are a write-path problem; a stale second screen is a cache-coverage problem, and conflating them sends you fixing the wrong layer.
Posting a second cheque because the first one never cleared in your statement - the bank cashed both. The remedy is a reference number on the payment, not a better way of tearing up your own copy.
saying these in an interview costs you the question
- Treats a timeout as proof the server did not commit
- Retries mutations under the same policy as queries
- Rolls back by restoring a snapshot of the whole store
- Thinks a better rollback would have prevented the duplicate
- Relies on the server rejecting duplicates with no idempotency key
- Predicts a server-assigned id in the optimistic result