skip to content

Why can a GraphQL response carry HTTP 200 OK and still report that the operation failed?

level: juniorimportance: must knowfreq 68%

answer

  1. Two layers, one status line
  2. The courier delivered; the cargo is mixed
  3. Partial results need a success status
  4. Only well-formed requests get the 200 rule
  5. The legacy media type froze the behaviour

basics

~20 s

The HTTP status describes the transport, not the GraphQL result. A request the server parsed and executed answers 200 even when execution produced errors, and under the legacy application/json response type the GraphQL over HTTP draft requires 200 for every well-formed request.

solid answer

~50 s

HTTP and GraphQL are two stacked layers, and the status code belongs to the lower one: 200 means the server understood the request and is returning the response document it promised. Whether every selected field resolved is reported inside the body instead. That matters most for a failure **during** execution — one resolver in a large selection throws while the rest succeed, so the server has a genuine partial result to hand back and a 500 would be a lie. Under the older `application/json` response media type the rule is stronger still: the GraphQL over HTTP working draft says a well-formed GraphQL request gets 200 whatever the outcome, even one rejected before execution began. Non-2xx statuses are reserved for requests that never became well-formed GraphQL requests at all — an unparseable body, a rejected method, an unsatisfiable Accept, an authentication failure at the edge.

code

graphql · 10 lines
graphql
query DepotList($depotId: ID!) {
  depot(id: $depotId) {
    vehicles {
      id
      registration
      odometerKm
      lastServiceRecord { completedAt technician }
    }
  }
}

go deeper

for a junior

Be ready to say plainly that the status code reports the HTTP exchange while the GraphQL outcome lives in the body, and that a 200 does not mean every field resolved. Interviewers ask this to check you have actually read a GraphQL response.

for a middle

Explain the mechanics: a resolver that throws after execution began yields a partial result at 200, while a request rejected before execution is 200 only because the legacy application/json media type freezes it there for compatibility.

for a senior

Show the operational consequence. Talk about error rate that has to come from bodies rather than statuses, about retry and circuit-breaker rules that never fire, and about what a partial result should and should not trigger for the caller.

for a principal

Own the observability contract: what a platform team standardises so every service reports GraphQL failures comparably, and whether you are willing to move a client fleet onto status-carrying responses for the sake of infrastructure that keys on the status line.

## Two layers, one status line An operation travels over HTTP, but HTTP is only the courier. The status code answers a transport question — did the server receive something it could act on, and is it returning the document that was asked for? — and it has no vocabulary for the GraphQL question, which is whether every field in the selection produced a value. So a `200 OK` on a GraphQL endpoint should be read as *"here is your response document"*, not as *"everything worked"*. Nearly every surprise in this area comes from importing the REST habit where the status line is the result. ## The failure that most obviously must be 200 Take a fleet telematics graph. `Vehicle` is a broad type — 37 fields, everything from `odometerKm` to warranty and maintenance data pulled from three different backends. A depot dashboard selects 12 of them for a list of trucks. One of those fields, `lastServiceRecord`, is backed by a maintenance system that times out after 2.4 seconds. Eleven of the twelve fields resolved. The dashboard can render the whole page except one column. If the server answered `500`, it would be telling the client to throw away a response that is overwhelmingly usable, and most HTTP client stacks would do exactly that — take the non-2xx branch and never look at the body. So a failure raised inside a resolver, once execution has begun, answers 200 and reports the failure inside the response body. This is true under **both** response media types the GraphQL over HTTP working draft defines; nothing about it is legacy behaviour. ## The failure where the status is a choice, not a necessity The other failure moment is *before* execution: the document did not parse, it failed validation against the schema, or a variable could not be coerced to its declared type. Nothing ran. There is no partial result, only a rejection. Here the status genuinely could be a 4xx, and the working draft's newer response media type `application/graphql-response+json` says it must be. But under the older `application/json` — which is what the overwhelming majority of deployed servers still answer with — the draft directs the server to use 200 for any **well-formed GraphQL over HTTP request** regardless of the outcome. That is the compatibility rule: clients were written against servers that always answered 200, and they read failures out of the body. Changing the status on those clients breaks them, so the legacy media type freezes the behaviour. ## What "well-formed" is doing in that sentence The 200 rule is not a promise that a GraphQL endpoint never returns 4xx. It is scoped to requests that actually arrived as GraphQL requests. Outside that scope, ordinary HTTP applies: * a request body that is not valid JSON, or that carries no operation to run, was never a well-formed GraphQL request — 400; * a method the endpoint does not accept for that operation — a normal method-level rejection; * an `Accept` header listing no media type the server can produce — 406; * authentication or authorization refused by a layer in front of GraphQL — 401 or 403, decided before any document is looked at. So the honest summary is: once a request has become a GraphQL operation, the status stops carrying result information under `application/json`; before that point it behaves like any other HTTP endpoint. ## What this costs you in production The 200-with-errors convention has real operational consequences, and interviewers often push here after the definition. * **Monitoring lies by default.** A dashboard counting non-2xx responses shows a perfectly healthy endpoint while every operation on it is failing. Error rate for a GraphQL service has to be derived from response bodies or from server-side instrumentation, not from the status line. * **Status-keyed infrastructure never engages.** Retry policies, circuit breakers and alerting rules that trigger on 5xx sit idle through an outage in a downstream system. * **Caching sees success.** A response that is entirely errors looks, to anything keying on the status, exactly like a good one. * **Client code cannot take shortcuts.** "If the request succeeded, use the body" is wrong. Every response has to be inspected for failures even at 200. ## Where the specification actually stands One precision worth carrying into an interview: GraphQL itself specifies an execution algorithm and a response format, not a transport. The rules above come from the **GraphQL over HTTP specification, which is a working draft** — widely followed and the basis for what servers do, but not a ratified edition of the GraphQL specification. The 200-with-errors behaviour predates the draft; the draft wrote down the deployed convention and then added a second media type under which the status becomes meaningful again for pre-execution failures. ## The trap in the question Asked "should a GraphQL error be a 500?", the weak answer is a flat yes or no. The strong answer separates the two moments: a resolver that failed after execution began is a 200 with a partial result under either media type, and a request rejected before execution is a 200 under `application/json` and a 4xx under `application/graphql-response+json`.

  • When does a GraphQL over HTTP endpoint answering application/json legitimately return a 4xx?
    When the request never became a well-formed GraphQL request. An unparseable JSON body, a body with no operation to run, a method the endpoint refuses for that operation, an `Accept` header naming nothing the server can produce, or an authentication check that rejected the caller before any document was read. The 200 rule is scoped to well-formed GraphQL requests; everything short of that is ordinary HTTP.
  • If failures come back at 200, how should the service's error rate be measured?
    From the response body or from server-side instrumentation, not from the status line. Count operations whose response reported failures, and separate the two classes — rejected before execution versus a resolver that threw — because they have different owners. Keep the transport-level non-2xx count too, since it still catches the requests that never reached execution at all.
  • Does the newer application/graphql-response+json media type make resolver failures return 4xx?
    No. That media type only changes the status for failures where execution never began and the response therefore has no data entry. Once execution started, the response carries a data entry — possibly null — and the status stays 200 under both media types, because a partial or nulled result is still a result the client is meant to read.

The status code is the delivery receipt, not a quality inspection: the parcel arrived intact, which says nothing about whether one of the items inside is broken.

saying these in an interview costs you the question

  • Says a GraphQL failure should always be a 500
  • Treats 200 as proof every field resolved
  • Counts non-2xx responses as the endpoint's error rate
  • Thinks a GraphQL endpoint can never return 4xx
  • Skips reading the body whenever the status is 200

context