In a GraphQL API, how do you keep half-applied writes from becoming a standing source of corrupt data?
answer
- The field is the only atomic unit
- Enforce at the boundary, not in review
- One convention beats several good ones
- Measure responses with payloads and errors
- Reconcile what cannot be atomic
basics
~20 sMake the field the atomic unit and enforce that at the boundary rather than in review: reject multi-write documents, require one idempotency convention across the graph, measure the rate of partially-applied mutation responses, and reconcile what genuinely cannot be atomic.
solid answer
~50 sStart by accepting the physics: no document-level transaction exists, so the root field is the only atomic unit you can offer. Everything else follows. Move the seam server-side — publish coarse mutation fields named for business actions, so a single resolver owns the transaction or the orchestration and its compensation. Then **enforce** rather than advise: a server-side validation rule that rejects a mutation document carrying more than one root field is a gate, and a style guideline is not, because you do not ship every client. Standardise one idempotency convention for the whole graph rather than one per team. Instrument the thing you are trying to eliminate: the rate of mutation responses that carry both payloads and a non-empty `errors` list is a first-class signal, and it should be alarmed on, not merely available. For actions that genuinely span transactional boundaries, own the reconciliation yourself and publish how long inconsistency may last.
code
pseudocode · 5 linesvalidate(document):
for operation in document.operations:
if operation.type == MUTATION and count(operation.rootFields) > 1:
reject("A mutation document may carry one root field; "
"use a single field that owns the whole action.")go deeper
You are unlikely to be asked this, but take away the shape: the fix for half-applied writes lives in the schema's write surface and in server-side rules, not in a note telling client developers to be careful.
Be able to explain why one coarse mutation field is safer than three fine-grained ones, and what a validation rule limiting root fields would do. Knowing the tradeoff against reuse will already put you ahead.
Show that you would measure before enforcing. Describe the report-only rollout, the correctness metric you would emit, and the reconciler for whatever genuinely cannot be atomic.
Own the whole contract: where the atomic boundary goes, what is enforced at the boundary versus advised, one idempotency convention for the graph, and an explicit statement of how long inconsistency may last. Be equally clear about the writes not worth protecting.
## The problem is not a bug, it is a property Half-applied writes are not a defect somebody introduced; they are what happens when a protocol with no transactional scope meets clients that batch writes. You cannot patch them away in a resolver. The leadership question is where you put the boundary, what you enforce, and what you are prepared to clean up afterwards. ## 1. Concede the atomic unit and design to it The largest unit a single resolver controls is one root field, so that is the largest unit that can be atomic. Two consequences follow for the write surface: * **Publish fields named for business actions, not for steps.** `closeOutTable` rather than `applyDiscount` + `voidLine` + `capturePayment`. The action a product person can name is usually the right transactional boundary, and it is the one clients cannot accidentally split. * **Resist the pressure toward fine-grained writes.** Small mutation fields feel composable and reviewers like them. They are composable, and that is the hazard: composition is done by the client, in a document, with no transaction. Every fine-grained field you publish is an invitation to a multi-write document. The counter-argument is real and you should be able to state it: coarse fields have more arguments, more branching inside one resolver, and they generalise badly when the next screen needs three of the five steps. That is a genuine tradeoff. Resolve it by asking whether the steps share an invariant. If they do, they are one field. If they merely happen to occur together on one screen, they are separate fields and the screen may half-apply, which is fine when the invariant is nothing more than tidiness. ## 2. Enforce at the boundary, not in review A convention that lives in a style guide is honoured by the teams that read it. You do not ship every client — there is a mobile release from six weeks ago in the field, and a partner integration nobody on your team wrote. So put the rule where it cannot be skipped: a custom validation rule that rejects any mutation document carrying more than one root field, applied on the server before execution. Be honest about what that is: **a rule you added**, not something the specification provides. There is no built-in limit on root fields. It is exactly the same class of control as a depth or cost cap — a custom validation rule the server runs. Ship it behind a flag, measure how many operations it would reject before turning it on, and give the teams it breaks a coarse field to migrate to. If you also run a registry of approved documents, that is the cheaper place to enforce it, because a document that never gets approved never reaches production. ## 3. One idempotency convention, graph-wide Retry safety cannot be a per-team decision, because retries cross team boundaries. Decide once: the argument's name, whether it is required on every write that touches money or messaging, its scope and its retention. Then make the schema linter enforce it for new mutation fields. Two teams with two conventions is materially worse than one convention with gaps, because a caller cannot tell which regime a field is under. ## 4. Measure the thing you are trying to eliminate You cannot manage this without a number, and the number is available: **the rate of mutation responses that carry both a filled `data` map and a non-empty `errors` list**. Emit it per operation name. It is not a latency metric or an availability metric; it is a correctness metric, and it belongs on the same board. Alert on the rate, not on individual occurrences, and keep a correlation identifier on each so an on-call engineer can reconstruct which fields applied. A related number is worth keeping: how many mutation documents in production traffic carry more than one root field. That one should be trending to zero, and if it is not, your enforcement is not where you think it is. ## 5. Own the reconciliation you cannot design away Some actions genuinely span transactional boundaries — a payments provider on one side of the wire and your own store on the other. No schema shape fixes that. You have three positions and should pick deliberately: * **Orchestrate server-side and compensate.** One field, one owner, explicit compensating writes, and a documented statement of what a failed action leaves behind. Expensive to build, but the cost lands once. * **Make the action asynchronous.** The mutation records an intent and returns a handle; consistency arrives later and the client observes it. This is honest — it stops pretending the action was synchronous — but it changes the client's model, so it is a product decision, not just a technical one. * **Publish the steps and let each client sequence them.** The cheapest thing to build and the most expensive thing to own. Every client reimplements compensation, each one differently, and every bug becomes a data bug that outlives the release that caused it. Choose this only when the steps are genuinely independent. Whatever you pick, run a reconciler for the residue — charges without orders, orders without tickets — and publish a number for how long inconsistency may last, say a four-minute upper bound, so that downstream teams can design against something concrete rather than against hope. ## 6. Know when to stop spending Not every half-applied action deserves this. A favourite flag that did not save alongside an analytics ping that did is not corruption; it is a nuisance. Reserve orchestration, compensation and reconciliation for invariants somebody would actually be paged about — money, inventory, entitlements, anything that leaves the building. Spending saga machinery on a preferences screen is how a platform team runs out of credibility before it runs out of work.
- What is the strongest argument against collapsing every multi-step write into one coarse mutation field?Reuse. A coarse field encodes one screen's sequence, and the next screen usually wants three of the five steps, so you end up with several near-duplicate fields and a lot of branching in one resolver. The test is whether the steps share an invariant. If they do, one field is right despite the awkwardness. If they only co-occur on a screen, keep them separate and accept partial application.
- How do you roll out a rule rejecting multi-root-field mutation documents without breaking existing clients?Measure first — run the rule in report-only mode and count which operations and which callers it would reject. Publish coarse replacements for the sequences you find, migrate the clients you control, and give partners a deadline with the new field ready. Turn on enforcement per caller rather than globally, so a partner you cannot reach does not fail on the same day as your own app.
- Which single metric would you put on the board for this?The proportion of mutation responses carrying both a filled `data` map and a non-empty `errors` list, broken down by operation name. It measures the failure itself rather than a proxy, it is cheap to emit at the point the response is serialised, and a rise in it is always worth a look — unlike latency or error-rate, which move for a dozen unrelated reasons.
- When is accepting a half-applied write the right call?When no invariant crosses the steps and nobody would be paged for the inconsistency — a preference saved without its analytics event, a label applied without its audit note. Building orchestration and reconciliation for those consumes the budget you need for the writes that move money or inventory, and a platform team that spends it in the wrong place loses the argument for the right place.
saying these in an interview costs you the question
- Relies on a style guide to stop multi-write documents
- Claims the specification limits root fields per mutation
- Lets each team invent its own idempotency convention
- Tracks only latency and availability for write correctness
- Pushes orchestration onto clients it does not ship
- Builds compensation machinery for writes nobody is paged about