A GraphQL response stopped being shared-cacheable after one per-viewer field was added. Why, and what would you change?
answer
- One field changed the whole answer's audience
- Scope folds to the most restrictive value
- Public only if every selected field is public
- Split the operation along the viewer seam
- Never relabel a private field public
basics
~20 sScope folds to the most restrictive value, so one field hinted private makes the entire response private and no cache shared between users may store it. The fix is to move the per-viewer field out of that document into its own operation.
solid answer
~50 sThe scope half of the fold is contagious. A response is public only if **every** selected field is public, so adding one per-viewer field - a permission flag, a personalised count, a viewer's own notes - flips the whole response to private, and every cache shared between users must stop storing it. The lifetime usually collapses too, because per-viewer fields are typically hinted with a zero max age. The remedy is at the **document** level, not the hint level: split the operation so the large, stable, viewer-independent part stays public and cacheable, and fetch the per-viewer field in a second, small, uncached operation. Never fix it by marking the field public - that trades a cache miss for one viewer's data being served to another. And be honest that keeping shared caching alive costs a second round trip.
code
graphql · 14 linesquery ClaimDashboard($id: ID!) {
claim(id: $id) {
id
policySummary { policyType coverageStart }
timeline(first: 25) { id occurredAt kind }
}
}
query ClaimViewerFlags($id: ID!) {
claim(id: $id) {
id
viewerCanReopen
}
}go deeper
Know the headline rule: a response counts as public only if every field it selected is public, so one per-viewer field takes shared caching away from the whole answer.
Explain why the response, not the field, is the unit that gets stored, and why marking the field public or extending its lifetime are both wrong fixes for a scope problem.
Show the diagnosis and the operation-level remedy: split the document along the public/private seam, weigh the extra round trip, and instrument the computed policy per operation so the next collapse is caught on its deploy.
Own the design consequence: if operations are the unit of cacheability, then who may add a field to a hot document is an organizational control, and you need a story for what the API team promises about hit rate as consumers multiply.
## The incident A four-person platform team runs an insurance claims graph. Their busiest document is the adjuster dashboard: one claim, its policy summary, and its `timeline` - a list of claim events that grew without a bound, because claims opened in 2019 have accumulated over four thousand events each and nothing ever trimmed them. Serving that timeline from the origin is expensive; served from a shared cache in front of the service it was nearly free. Their measured hit rate on that operation sat at 78%, and p95 for the dashboard at 41 ms. One sprint they added a small field so the UI could grey out a button: ```graphql type Claim @cacheControl(maxAge: 300) { id: ID! policySummary: PolicySummary! timeline(first: Int!): [ClaimEvent!]! viewerCanReopen: Boolean! @cacheControl(maxAge: 0, scope: PRIVATE) } ``` The field is correct: whether *this* adjuster may reopen *this* claim genuinely differs per viewer, so private with a zero lifetime is the honest hint. The client team added it to the existing dashboard document, because that is where the button lives. Within a day the hit rate on that operation was 3%, p95 was 890 ms, and the unbounded timeline was being rebuilt at the origin on every single dashboard open. ## Why one field did that The fold that produces a response's policy is conservative in both dimensions. Lifetime takes the minimum; **scope takes the most restrictive value seen**. A response is public only if every selected field is public. One private field makes the response private, and a cache shared between users must not store a private response at all - not the whole body, and not the public parts of it, because what is being stored is one opaque body and the cache has no notion of GraphQL fields inside it. So the 300-second, viewer-independent bulk of the dashboard - the expensive part - lost its shared cache because of a boolean. The zero max age made it worse, but the scope alone would have been enough. This rule is not something to work around. A shared cache handing one adjuster's `viewerCanReopen` to a different adjuster is a data-leak incident, and the whole value of the private scope is that the conservative fold makes that impossible by construction rather than by everybody remembering. ## The fix that works The leverage is in the **document**, not in the hint. Cacheability is a property of what a document selects, so split the selection along the public/private seam: - One large operation for the viewer-independent material - policy summary, timeline page - which stays public with its 300-second lifetime and goes on being served from the shared cache. - One small operation for the per-viewer material - `viewerCanReopen`, and any other flag of the same kind - which is private, uncached, and cheap to compute because it touches a permission check rather than an unbounded list. The cost is an extra round trip, and a UI that must render before the second answer arrives. In this case that is a good trade: the button starts disabled and enables when the flag lands, while the expensive part of the screen paints from cache. Two second-order options are worth naming. Sometimes the per-viewer field can be **derived on the client** from data it already has - if the document already returns the claim's status and the viewer's role, a reopen flag may not need to be a server field at all. And sometimes the right answer is to accept the origin cost, if the operation is rare - the fix is only worth its complexity for genuinely hot documents. ## What not to do **Do not mark the field public** to get the hit rate back. That is trading a correctness property for a performance number, and the failure it invites is silent and cross-viewer. **Do not raise the private field's max age.** Scope, not lifetime, is what excluded the shared cache; a longer lifetime on a private field only extends how long one viewer's own private cache keeps it. **Do not assume the browser's cache is affected.** A private cache belonging to one viewer is precisely the cache still permitted to store a private response, so a per-user cache may keep working while the shared tier stops. Keying a shared cache per viewer is a different lever with different costs - it belongs to the design of the cache key, and it does not change the fold at all. ## Making it visible next time A four-person team cannot review every field addition against every document, so the defence is instrumentation rather than vigilance: - Record the **computed lifetime and scope per operation name** and treat a collapse - public to private, or minutes to zero - as an alertable regression on the deploy that caused it, not a slow discovery from origin traffic. - Track shared-cache hit rate per operation, and pair it with origin cost for the operations that select unbounded lists, since those are where the loss of a cache hurts most. - Lint the schema change itself: adding a private or unhinted field to a type that appears in a heavily cached operation is a review-worthy event, and it is mechanically detectable if you know which operations select which types. The interview point behind all of this is that response caching in GraphQL is a **whole-document property assembled from field-level declarations**, so the unit of design is the operation. Teams that add fields freely and expect caching to survive have the granularity wrong.
- If a shared cache must not store the response, can the requesting browser's own cache still keep it?Yes. Private scope excludes caches shared between users, not caches belonging to one viewer, so a per-user cache may still store the response for whatever lifetime the fold produced. That is worth remembering for two reasons: it means the private label is about audience rather than about storage in general, and it means a shared device or a browser profile handed to someone else is a real, if narrow, exposure worth thinking about for sensitive fields.
- How would you catch this regression on the deploy that introduced it rather than weeks later?Emit the computed policy - lifetime and scope - as a dimension on your per-operation metrics, and alert when an operation's scope flips to private or its lifetime drops to zero. Pair that with shared-cache hit rate per operation. Both signals move on the deploy itself, whereas origin cost and latency drift up slowly enough that a small team attributes them to growth.
- The product team insists the flag must be in the same round trip. What do you do?Then the document is private, and you say so plainly rather than mislabelling the field. From there the options are to make the expensive part cheap enough to serve from the origin - bounding that unbounded timeline is overdue anyway - to cache the expensive fragment inside the service where per-viewer separation is possible, or to buy the round trip back by having the client fetch the flags once for a set of claims rather than per claim.
saying these in an interview costs you the question
- Marks the per-viewer field public to restore hit rate
- Thinks a shared cache stores only the public fields
- Believes raising the private field's max age fixes it
- Says private scope also bars the requesting browser's cache
- Blames the lifetime when scope was the cause
- Treats cacheability as a schema property, not a document one