skip to content

questions

4

Why can a CDN store a GraphQL response that reports a failure, and how do you stop it?

level: middleimportance: must knowfreq 58%

answer

  1. The cache reads the status, not the body
  2. Execution can succeed and still be partial
  3. One bad minute, one frozen minute
  4. Only the origin knows the result failed
  5. Branch on errors before writing the lifetime

basics

~20 s

A shared cache decides from the HTTP status line, but GraphQL reports field failures in the body — a partly-null response carrying an errors entry still arrives as 200. The origin must mark those responses uncacheable before they leave.

solid answer

~50 s

HTTP caching is status-driven: 200 is one of the statuses a shared cache may store by default, so once a query rides a GET URL the edge keeps whatever came back under it. GraphQL is body-driven: when a resolver fails, the specification requires the failure to appear as an entry in the response's `errors` list with the field nulled in `data`, and the transport itself succeeded, so the status is still 200. The result is a poisoned entry — one downstream blip frozen at the edge and replayed to every viewer asking for that operation until the lifetime expires. The fix belongs at the origin, not the edge: the code that writes the response has both the GraphQL result and the HTTP response in hand, so it should branch on `errors` being present and emit no storable lifetime. Adopting the GraphQL over HTTP working draft's response media type maps pre-execution failures to 4xx, but field errors are legitimately 200 and always will be.

code

json · 17 lines
json
{
  "errors": [
    {
      "message": "credits unavailable",
      "path": ["release", "tracks", 0, "credits"]
    }
  ],
  "data": {
    "release": {
      "id": "rel_4471",
      "title": "Nightjar Sessions",
      "tracks": [
        { "id": "trk_9083", "title": "Chalk Line", "credits": null }
      ]
    }
  }
}

go deeper

for a junior

Recall that a GraphQL failure usually arrives as HTTP 200 with an errors list in the body, and that a cache in front decides from the status code alone. Those two facts together are the whole answer at this level.

for a middle

Be ready to explain the mechanics: which failures produce a present-but-holey data map, why 200 is storable by default, and where in the request path the check has to live so that every cache benefits from it.

for a senior

An interviewer expects the incident shape — a transient dependency failure amplified into minutes of degraded reads for every caller — plus how you would detect it, how long the tail runs, and why the fix is a branch at the response writer rather than an edge rule.

for a principal

Own the policy: cacheability is decided once, at the origin, by code that sees the execution result, and no cache configuration anywhere is allowed to encode GraphQL semantics. Be able to argue that partial-success caching is not worth the reasoning cost.

## Two different senses of "it worked" A shared cache — a CDN node, a reverse proxy, any store sitting between the caller and the origin — decides what it may keep from the request method, the URL, the response status and the freshness the origin declared. HTTP treats 200 as cacheable by default. Nothing in that decision reads the response body, and nothing in it knows what GraphQL is. GraphQL puts the outcome in the body. The specification defines a response map of `data`, `errors` and optionally `extensions`, and requires that a field which failed during execution be reported as an entry in `errors`, carrying a `path` to the exact position of the failure, while `data` is still present and still carries every field that succeeded, with `null` where the failure landed (bubbling up to the nearest nullable ancestor if the failed field was non-null). Execution completing with holes is a normal, specified outcome, not a transport failure. So the transport says "200, fine" while the body says "a third of this is missing". Put a shared cache between the two and they drift apart in the worst possible direction: the layer deciding what to keep is precisely the layer that cannot see the failure. ## The poisoning, concretely A music catalogue serves a public `ReleaseDetail` read over GET, keyed on a persisted operation identifier plus the release id, with a lifetime of 137 seconds. Its `credits` field is backed by a separate credits service. That service was unhealthy for 96 seconds one morning; the resolver raised, and the origin returned a perfectly well-formed 200 with `data.release.tracks[*].credits` set to `null` and one entry in `errors` per failed field. The edge stored it, and every node that stored a copy then served the hole to every caller for the full 137 seconds — including well after the credits service recovered at 09:14, because a stored entry does not learn that the world got better. 8,214 requests were answered from that poisoned entry. The origin's error rate had already returned to zero by then, which is why the incident looked, from the dashboards, like it ended twenty minutes before the complaints did. That is the whole hazard in one line: a cache turns a transient failure into a persistent one, and turns a failure that hit one caller into a failure that hits everyone. ## Request errors, field errors, and what the over-HTTP draft changes Two failure classes behave differently and interviewers like the distinction. **Request errors** happen before execution: the document did not parse, it failed validation, a variable could not be coerced. The specification says `data` must not be present in that response at all. **Field errors** happen during execution: a resolver raised, or returned `null` for a non-null field. `data` is present, holes and all, and `errors` describes them. Under the long-standing `application/json` media type, both classes come back as 200 and both are therefore storable. The GraphQL over HTTP specification — still a working draft, not a ratified spec — changes this for its own media type `application/graphql-response+json`: a request error yields a 4xx rather than a 200, while a well-formed request whose execution produced field errors still yields 200. Adopting it therefore fixes the class where nothing ran, and cannot fix the class where something ran and partly failed. The first class bites in a way worth naming. After the catalogue removed a deprecated `Track.previewUrl` field, an old client pinned to a document that still selected it kept sending that operation. Under the legacy media type each of those requests was a 200 whose body was nothing but a validation error — storable, and duly stored, so the failure kept being served from the edge for the lifetime even after an emergency schema rollback had restored the field at the origin. ## Fixing it at the origin The rule is one branch: the response writer is the only place holding both the GraphQL result and the HTTP response, so cacheability is decided there. If `errors` is present, the response goes out with no storable lifetime; if it is absent, it gets whatever lifetime the operation earned. Keeping both decisions together means no configuration anywhere else has to know about GraphQL's failure convention. Resist being clever about partial success. "The failing field was only a sidecar, cache it anyway" is defensible exactly once and then becomes a rule nobody can reason about. ## What not to do **Do not push the check to the edge.** Some edge platforms can inspect a JSON body and refuse to store on a condition, which looks like a tidy central fix. It is not: you are reimplementing knowledge only the origin has, in a system deployed on a different cadence, and it protects one cache — a browser cache, a corporate proxy or a second CDN in the path will do none of it. **Do not rely on errors being rare.** Rarity is the wrong axis: a cache multiplies each bad response by however many callers ask for that key during the lifetime, so a low origin error rate and a large blast radius are perfectly compatible. ## Specified, or convention? Specified by GraphQL: the response map's shape, that field errors are reported in `errors` with a `path`, and that `data` is absent when the request itself failed. A working draft, not final: the status-code mapping under `application/graphql-response+json`. And how a response declares whether a cache may store it is plain HTTP, covered by the HTTP caching topic — what matters here is that GraphQL hands you a body-level signal no cache will ever look at, so the origin has to translate it into something a cache does look at.

  • Does returning a 5xx whenever any field errors solve this?
    It stops the storage, but it throws away partial results the caller could have used and misreports a working server as broken — every client's retry and circuit-breaker logic now fires on a single missing sidecar field. Keep the 200 and the partial data, and make the response non-storable instead. The one place a non-200 is right is the request-error class, where nothing executed and there is no partial result to preserve.
  • How long can a poisoned entry actually be served after the origin recovers?
    Up to a full lifetime past the last time a cache stored a failing copy, independently at every node that stored one. Recovery at the origin does not reach into a stored entry, so the tail is bounded by the lifetime you chose, not by how long the incident lasted. That is the argument for short lifetimes on operations you cannot purge quickly.
  • A response has data fully populated and an empty errors array. Is that storable?
    The specification says `errors` must not be present in the response at all when no errors occurred, so an empty array is a server bug worth fixing. Treat presence of the key as the signal, and you will refuse to store responses that a stricter reading would allow — which is the safe direction to be wrong in, and it disappears once the server stops emitting the key.

saying these in an interview costs you the question

  • Assumes a 200 means the whole response is good
  • Thinks caches inspect the JSON body
  • Says GraphQL returns 500 when a resolver fails
  • Puts the errors check in the edge configuration
  • Treats a rare error rate as a small blast radius
  • Believes origin recovery clears stored entries

context

open as a page

What must an edge cache key contain for a GraphQL query sent over GET?

level: seniorimportance: should knowfreq 47%

basics

~20 s

An identity for the operation that will run, the variables in a canonical form, and — whenever a selected field depends on who is asking — a dimension of the caller's identity. Miss the third and one viewer receives another's data.

open as a page

How do you decide which GraphQL operations belong behind an edge cache and which stay at the origin?

level: principalimportance: should knowfreq 44%

basics

~20 s

Per operation, never as a global switch. Promote reads whose origin cost times request rate is large, whose variables have low cardinality, that carry no viewer-scoped field, and whose staleness budget you can state in seconds.

open as a page

Why is a GraphQL request's operationName unsafe as a shared cache key?

level: seniorimportance: nice to knowfreq 19%

basics

~20 s

Because operationName only picks which operation to run out of a document that defines several. It is a client-chosen label, not an identity: two unrelated documents may reuse it, and it says nothing about the variables.

open as a page