A document that succeeds in a GraphQL explorer fails from the application. How do you diagnose it?
answer
- Two requests, not two documents
- Capture both completely before theorising
- The response shape names the layer
- Who was the tab authenticated as?
- You tested twelve; the app asked for 250
basics
~20 sStop comparing documents and compare requests. Capture both calls in full — body, headers, credentials, target environment, variable magnitudes — then replay the application's exact request from the explorer. The difference is almost always identity or scale, not the document.
solid answer
~50 sThe explorer proves one thing only: this document is valid against the schema that endpoint introspected. It proves nothing about the request the application sends. So capture both requests completely — the `query`, `variables` and `operationName` in the body, every header, whatever cookies the browser attached invisibly, and which environment the URL points at — and diff them. Then read the failing response properly, because its shape names the layer: a body with no `data` key means the operation was rejected before execution, by parsing, validation, or a limiter; a body with `data` present and an errors entry carrying a `path` means execution ran and one field failed, which points at authorization or a backend. The usual culprits, in order, are a privileged identity in the explorer tab, variable magnitudes far larger than the ones you tested, and production-only policy such as a registered-document allowlist.
code
json · 8 lines{
"errors": [
{
"message": "Operation exceeds the maximum permitted cost.",
"extensions": { "code": "COST_LIMIT_EXCEEDED", "cost": 5312, "maximum": 2000 }
}
]
}go deeper
Know that a green run in the explorer is not proof the application will succeed, and that the two calls can differ in credentials, in variable values and in which environment they hit. Check the endpoint URL before anything else.
Explain what the response shape tells you: no data key means rejected before execution; data present with an error carrying a path means one field failed during execution. Be able to compare two requests field by field, headers included.
Demonstrate the method: capture both requests fully, diff them, replay the application's exact request from the explorer, and reason from the error path. Name identity, variable magnitude and production-only policy as the usual causes without guessing.
Own the systemic answer. Decide how a shipping document is proven before release rather than by hand in an explorer, what identity and data scale that proof runs at, and how much production-only policy the team is willing to have that lower environments never exercise.
## Reframe the question first The instinct is to compare the two documents, find them identical, and conclude the server is behaving inconsistently. It almost never is. What differs is not the document but everything wrapped around it, and the useful reframing is: **these are two different HTTP requests that happen to contain the same document text.** So the first move is capture, not theory. For each of the two calls, write down the request body (`query`, `variables`, `operationName`), every header, whatever the browser attached that no pane shows, the endpoint URL and therefore the environment, and the exact response — status, body, and the full `errors` array including each entry's `path` and `extensions`. Two complete captures side by side usually end the investigation before any hypothesis is needed. ## Read the failure shape before guessing The failing response tells you which layer to look at, and it is worth being precise here because a great deal of GraphQL failure arrives with a 200 status. - **No `data` key at all, errors present.** The operation was rejected before execution began: a parse or validation failure, an unregistered document under an allowlist, or a limiter refusing it on depth or estimated cost. The document never ran. - **`data` present with a null in it and an errors entry carrying a `path`.** Execution happened and one field failed. The `path` names the exact field, which is the single most useful string in the whole investigation. - **A non-2xx with no GraphQL body at all.** You did not reach the executor — a proxy, an authentication gate, or a rejected preflight. Nothing about the document is implicated. ## The five differences that actually cause this **1. Identity.** By a wide margin the most common. The explorer tab is authenticated as *you*, often by a session cookie the browser attaches silently, or by a broadly-scoped token someone pasted into the headers pane months ago. The application authenticates as a service or as an end user with a narrower scope. Field authorization then returns null for the fields your identity could see, with an error whose `path` names them — and the whole response still arrives as a 200 with partial data. **2. Scale.** You tested with the numbers a human types. A music catalogue graph makes this vivid: in the explorer, `ReleaseCredits` on release `rel_9f3c14` with `first: 12` returns in 240 ms. The application requests the credits page for a box set of 1,412 tracks with `first: 250`, and the operation is rejected before execution by a cost limiter, or executes and times out at the gateway. Same document text; different variable magnitude; different outcome. **3. Production-only policy.** Environments are not identical, and the ones that differ most are exactly the ones you cannot poke. A registered-document allowlist accepts only documents shipped through the build, so an ad-hoc document typed into an explorer is rejected in production and accepted everywhere else. Rate limits keyed on a client identifier behave differently for a browser tab making one call than for a service making thousands. **4. Environment.** The explorer is pointed at a staging deployment one schema version ahead. The field exists there and not in production, and the failure is an unknown-field error that looks absurd until you check the URL. Always confirm the endpoint, not the tab title. **5. The request the application makes is not the one you think.** A client library may add its own headers, may deduplicate or batch calls, and may serve the value from its own cache so the network call you are hunting for never happens at all — which presents as a field that returns stale data rather than as an error, and sends people looking at the server for a problem that lives in the caller. ## Close the loop by replaying The decisive step is to reproduce the *application's* request from the explorer: paste the application's exact document, its exact variables, and its headers into the panes, and clear whatever privileged credential the tab was carrying. If it now fails in the explorer, the difference was in the request and the capture diff will name it. If it still succeeds, the difference is outside the request — the environment, the network path, or the client library's own behaviour before it ever issues the call. One habit prevents most recurrences: never let the explorer be the only place a shipping document has run. The documents an application sends should be executed against the real schema automatically, before release, with the identity and the variable magnitudes production will actually use. An explorer is an excellent instrument for exploring a graph and a poor one for proving that anything works.
- The failing response is a 200 whose body has a data key and one error with a path. What does that rule out?It rules out everything that happens before execution. The document parsed, validated against the schema, and was accepted by any limiter or allowlist in front of the executor — otherwise there would be no `data` key at all. Execution then ran and one field failed. The `path` names it exactly, so the investigation narrows to that field's resolver and its inputs: authorization for the identity on the request, or a backend it depends on.
- Why is the explorer's authentication so often the difference, even when nobody set it deliberately?Because a browser-hosted explorer is frequently authenticated by a session cookie the browser attaches on its own, with nothing in the headers pane to show for it. The tab is therefore signed in as a developer — often with broader scopes — while the application authenticates as a service or an end user. Field authorization then diverges, and the response arrives as a 200 with nulls and error paths rather than as an obvious authentication failure.
- What practice stops this class of surprise from recurring?Do not let the explorer be the only place a shipping document has ever run. Execute the application's own documents against the real schema automatically before release, using an identity with production's scopes and variable magnitudes matching real usage rather than the small numbers a human types. The explorer stays what it is good at — exploring a graph — instead of standing in as evidence that a document works.
- How do you tell a limiter rejection from a plain validation failure when both arrive without a data key?Read the error entries. A validation failure names the offending part of the document — an unknown field, a wrong argument type — and typically carries a source location. A limiter rejection describes a budget rather than the document's shape, and servers conventionally put a machine-readable code and often the computed score into `extensions`. If neither is clear, shrink the variables: a request that succeeds at a smaller page size was refused on cost, not on validity.
It is the classic "works on my machine", with the machine being your browser tab: the tab is signed in as you, asks for small pages, and talks to whichever environment you last typed in.
saying these in an interview costs you the question
- Assumes an explorer run proves the document works in production
- Compares document text only, never headers or variable sizes
- Believes a failure must arrive as a non-200 status
- Forgets the browser attaches cookies the headers pane never shows
- Never checks which environment the explorer points at
- Re-runs it in the explorer repeatedly instead of replaying the app's request