How do you set a team-wide policy for keeping a GraphQL client cache correct after mutations?
answer
- Defaults, not per-feature judgement
- Correct by default, fast by exception
- The schema floor: return what you changed
- Refetch for membership, patch only with evidence
- Every screen survives a cold load
basics
~20 sPick a default that is correct without thought - mutations return the objects they changed, and membership changes refetch the affected operations - then allow hand-written store patches only on measured hot paths, and measure the refetch traffic the policy costs.
solid answer
~50 sTreat it as a defaults problem, not a per-feature one. Set a floor every mutation must meet: return the changed objects with their identity and every field the write can move, so value staleness never needs a decision. Make **refetching the affected operations** the default for membership changes, because it is correct by construction and needs no client-side knowledge of ordering or filtering. Allow direct store patching only where a measurement justifies it, behind reviewed helpers rather than at arbitrary call sites, and only for lists whose shape the client fully controls. Then instrument the cost - requests per mutation, time to settle, share of refetched bytes that were unchanged - and tighten where it hurts. Finally, keep the escape hatch honest: every screen must survive a cold load, so the cache is never the only thing making the UI correct.
code
pseudocode · 10 lines# colocated with the module that owns the lists, not with the call site
invalidates(mutation = "joinWaitlist") = [
operation("EventWaitlistPanel"), # ordered by position, server-filtered
operation("VenueTonightSummary") # derived counts, server-computed
]
# tier 2 exception, reviewed helper, client-controlled ordering only
patches(mutation = "appendBoxOfficeNote") = [
listPatch(entity = "Event", field = "boxOfficeNotes", strategy = APPEND)
]go deeper
You will follow the team's rule rather than set it. Know what your codebase's default is after a write - returning the changed object, refetching a list, or patching the store - and why the default exists.
Be ready to justify the mechanism you chose for a specific write and to say what it costs in round trips or maintenance. Recognise when your list's ordering or filtering makes a hand-written patch the wrong call.
Argue for defaults over case-by-case judgement, and bring evidence: the requests-per-mutation and unchanged-bytes numbers that decide whether a hot path earns an exception. Know how invalidation knowledge rots and how a test catches it.
Own the whole policy - the schema-side floor, the default for membership, the exception process, the retry and prediction rules, the staleness budget per surface, and the invariant that every screen is correct from a cold store.
## Why this needs a policy at all Each individual case has a defensible answer, which is exactly the problem: 40 engineers making 40 defensible local calls produce a codebase where nobody can say what a given mutation invalidates. The failure modes at scale are two, and they are opposite. - **"Refetch everything to be safe."** Every write triggers a broad reload. Correct, and the backend absorbs a multiple of the traffic the product actually generates; on a ticketing platform, a checkout flow with five writes can turn into dozens of extra documents at exactly the moment the system is busiest. - **"Never refetch, patch by hand."** Fast and cheap - and the invalidation logic is now scattered across hundreds of call sites, each encoding a list's ordering and filtering. It rots silently: someone adds a status filter to a list argument and eleven patch sites become subtly wrong with no test failing. A policy exists to make the cheap correct thing the path of least resistance, and to make the expensive or brittle thing a deliberate, reviewed exception. ## The three-tier default **Tier 0 - the schema floor.** Every mutation returns the objects it changed, carrying identity plus every field the write can move. This costs nothing extra at runtime and removes an entire class of bug from the client's hands. It is a contract the client platform asks of the schema, and it belongs in schema review, not in a feature ticket. On a wide type - a 37-field `Event` accumulated over four years - "every field the write can move" is a real analysis, and the honest answer is often to return the whole object rather than guess. **Tier 1 - the default for membership.** Creates, deletes and reorders refetch the operations that read the affected list. It is correct by construction: the server re-applies its own ordering and filtering, and the client needs no model of either. It costs round trips and settles after the mutation response, so the UI contract has to define what the user sees in that window - usually the optimistic prediction, or a short pending state. **Tier 2 - the measured exception.** A direct store patch is allowed when a hot path shows the refetch cost, and only for lists whose ordering and filtering the client can reproduce exactly. The patch lives in a small reviewed set of helpers, not at the call site, so there is one place to fix when a list gains an argument. That ordering is deliberate: correctness by default, performance by exception with evidence. ## What the policy needs beyond a default **A mutation-to-affected-operations map.** The hardest part of any refetch rule is knowing what to refetch. Left implicit, it is folklore. Make it an artefact: a declaration next to each mutation naming the list-owning areas it disturbs, colocated with the module that owns those lists so it moves when they do. It will still rot; the mitigation is a test per critical flow that runs the mutation and then asserts the screen's data against a cold read. **Rules on retries and predictions.** Which operations may be retried automatically, where an intent key is generated and persisted, and which writes are allowed an optimistic prediction at all. These decide whether the cache can be corrupted by the transport, and no single feature team should be answering them. **A staleness budget by surface.** Not every screen deserves the same rigour. Seat inventory, money and anything the user is about to act on are strict. A venue description or a marketing banner can be stale until the next natural navigation. Writing the budget down is what stops the strict answer from being applied uniformly and blowing up the traffic bill. **Measurement.** Track requests issued per mutation, the p95 time from write to settled UI, and the fraction of refetched payload bytes that turned out unchanged. That last one is the signal that a broad refetch is being used where a narrower repair belongs. A regression here should be as visible as a latency regression. ## Alternatives worth considering, and their limits If several clients care about the same entities, pushing the canonical object over a subscription after a write converges everyone with no refetch. That is a good fit for genuinely shared, high-churn state such as remaining inventory, and a poor one everywhere else: it adds a long-lived connection, a fan-out problem and an ordering problem to solve a staleness issue that one refetch would have handled. Shortening a store's retention, or evicting aggressively after writes, is superficially attractive and mostly moves the cost: reads that would have been served locally become requests, at a time you do not choose. ## The invariant to defend Whatever the policy, the cache is an accelerator, not a source of truth. Every screen must be correct after a cold load with an empty store, and the invalidation strategy must never be the only thing making it correct. That single rule is what keeps a wrong invalidation decision a performance bug and a brief inconsistency rather than a data-integrity incident - and it is what lets you tune the rest of the policy without fear.
- How do you keep the map from mutations to affected operations from rotting?Two mechanisms, because neither is enough alone. Colocate the declaration with the module that owns the list, so it is in the diff when the list changes. Then defend it with a test per critical flow: run the mutation, apply the policy, and assert the resulting screen data matches a cold read from an empty store. That test fails when someone adds a filter argument and forgets the declaration, which is exactly the silent case.
- What tells you the policy has become too expensive?Requests issued per mutation trending up, p95 time from write to settled UI stretching, and above all the share of refetched payload bytes that came back unchanged. A high unchanged fraction says broad refetching is doing work nobody needed. Treat these as first-class metrics with the same visibility as latency, and use them to promote specific hot paths to a reviewed store patch rather than loosening the default globally.
- When do you deliberately accept a stale client cache after a write?When the surface is non-authoritative and the blast radius is small - descriptive content, secondary summaries, anything the user is not about to act on - and the next natural navigation will refetch anyway. The decision belongs in a written staleness budget per surface, so that the strict treatment reserved for inventory and money is not applied uniformly. What is never acceptable is staleness on a value the user is about to base a commitment on.
- Why insist that every screen work from an empty store when the cache is nearly always warm?Because it demotes every invalidation mistake from an integrity incident to a performance bug. If a screen is correct only when the store already holds patched data, a missed invalidation shows the user wrong information with no recovery path. If it is correct from cold, the worst case is a stale render until the next read. It also keeps the cache replaceable, which is what lets you change the policy at all.
saying these in an interview costs you the question
- Makes refetch everything the standing rule to be safe
- Bans refetching and hand-patches the store at every call site
- Treats cache correctness as purely a client concern
- Assumes optimistic updates remove the need for a policy
- Has no way to say which operations a given mutation affects
- Relies on the cache being warm for a screen to be correct