skip to content

A shipped client's GraphQL requests all fail with a syntax error — how do you localise the fault?

level: seniorimportance: should knowfreq 30%

answer

  1. Ask what the failure proves never happened
  2. The schema is not a suspect here
  3. Read the position against the sent text
  4. Look at how the client builds the string
  5. Constant document, values as variables

basics

~20 s

A syntax error proves the schema, resolvers and data stores never ran, so the fault is in the bytes the client sent. Capture that exact document text, read the error's line and column against it, and inspect how the client assembles it.

solid answer

~50 s

A syntax error is the most eliminating failure GraphQL produces: the server could not turn the text into a syntax tree, so it never consulted the schema, never validated, never called a resolver. That excludes a schema rename, a bad deploy and a database outage in one stroke — a rename surfaces as an unknown-field **validation** error, never a parse one. So the investigation is entirely on the client side: reproduce the exact document text the request carried, and read the error's `locations` line and column against *that* text, not against the source file the query was written in, because clients minify and assemble documents. The usual causes are values interpolated into the document string instead of being sent as variables, a build or minification step mangling or truncating the assembled text, and a proxy cutting the body short. The durable fix is a document text that is constant per build, with values travelling as variables.

code

pseudocode · 7 lines
pseudocode
# broken: the value becomes part of the document's grammar
document = 'query { listings(q: "' + userInput + '") { id } }'
# userInput = 'the "Willow" conversion'  ->  string literal ends early

# safe: the document text is constant, the value travels apart
document  = 'query Search($q: String!) { listings(q: $q) { id } }'
variables = { "q": userInput }

go deeper

for a junior

Know that a syntax error means the server could not read the document at all, so nothing about the schema, the resolvers or the data is implicated yet.

for a middle

Be able to reproduce it: capture the exact document text the client sent, and read the error's line and column against that text rather than against the file the query was written in.

for a senior

Lead with the elimination argument — schema, resolvers and data stores never ran — then name the client-side causes: values interpolated into the document instead of sent as variables, and build steps that mangle or truncate the assembled text.

for a principal

Own the prevention: make the document text a build artefact rather than something computed at runtime, and check it against the schema before release, so a failure class that hits every user of a release simultaneously cannot ship.

## Start with what the failure excludes The most valuable property of a syntax error is everything it proves *did not happen*. Parsing is the first phase of the pipeline and it consults nothing but the document text. If parsing failed, then: the schema was never consulted, no validation rule ran, no operation was selected, no variable was coerced, no resolver was called, no authorisation check fired, nothing touched a data store. That single observation collapses the search space. On a real-estate listings graph where a mobile release started failing on the same afternoon as a schema change, the instinct is to blame the rename — and the instinct is wrong, because a rename cannot produce a parse error. A client asking for a field that no longer exists produces an **unknown field** error from validation, with a different message and a different shape. If the error says *syntax*, the schema is not a suspect. Neither is the deploy, the cache, the database, or the identity provider. The fault is in the bytes of the `query` string the client sent. That is the whole territory. ## Get the exact text The first concrete step is to obtain the document as the server received it, not as a developer wrote it. Three routes: log the raw document on a parse failure (truncated, and note that if user data is inside the document text, that fact is itself the bug you are hunting); capture a request from the shipped build with the client's own network tooling; or reconstruct it by running the client's build pipeline over the source. Then read the reported `locations` against that text. The line and column index the document source the server parsed. A client that minifies its operations onto one line will report line 1 and a large column number; a document assembled by concatenating fragment strings has line numbers that correspond to no file anyone can open. Candidates who compare the reported line against the `.graphql` file in the repository and conclude the server is confused have skipped this step. ## The usual causes **Interpolated values.** Far and away the most common. A client builds the document by string concatenation and drops a user-supplied value into it. As long as the value is tame, everything works and the code passes review; the first search for a listing whose blurb contains a double quote terminates the string literal early and the rest of the document becomes nonsense. Note the exact mechanism: GraphQL string literals are delimited by double quotes only — there is no single-quoted string form — so an apostrophe in a street name is harmless while a double quote is fatal. The fix is not better escaping. It is to stop putting values in the document at all: the document text becomes a constant and the value travels in the request's separate variables map, where no character in it can change the document's grammar. **Build-step mangling.** The document lived in a template literal or a tagged string that a minifier, transpiler or bundler rewrote. Or the document was assembled from several fragment strings and one of them was tree-shaken away, so the concatenation spliced the text `undefined` into the middle of a selection set. This class is nasty precisely because it appears only in the shipped build, never in local development — which is exactly the symptom in question. **Truncation.** A proxy or gateway with a body-size limit cut the request short, and the parser reports an unexpected end of input. The tell is that the reported position is at or near the end of the document and the message mentions end-of-file. Check the length of what arrived against the length of what was sent. **Encoding oddities.** Smart quotes from a document pasted through a word processor, a byte-order mark somewhere other than the start of the source, or a body read with the wrong charset. Rarer, but they produce baffling messages about an unexpected character. ## What not to chase Do not go looking for a performance cause. Parsing a couple of kilobytes of document is microseconds; against a request budget in the hundreds of milliseconds, the parser is invisible. A syntax error is a correctness and shape problem, never a latency one, and a stopwatch will teach you nothing about it. Also be careful about how the failure surfaces in your monitoring. A parse failure is a request error rejected before execution, and depending on how a client's dashboards are built — counting only errors that appear inside otherwise-successful responses, say — a wave of them can be invisible on the very panel a team watches. Confirm the failure rate against a signal that counts rejected requests. ## Prevention, which is the part a senior is really being asked about The class of bug is "the document text is computed at runtime", and the class of fix is "the document text is a build artefact". Concretely: write operations as static documents; pass every value as a variable; extract the operation text at build time so the exact string that will be sent is an artefact you can test; and check those extracted documents against the schema in continuous integration, which catches both syntax errors and validation errors before a release rather than after it. Once the document is constant per build, a parse failure in production becomes almost impossible — and if one appears anyway, the truncation and encoding branches above are the only two left standing.

  • Why is a syntax error strong evidence that a schema rename is not the cause?
    Because a rename can only be noticed once the document is compared with the schema, and that comparison happens in validation — one phase after parsing. A client asking for a field that no longer exists gets an unknown-field validation error. If the server reports a syntax error it never reached validation at all, so no schema change, however breaking, can explain it.
  • The error reports line 3, column 12, but the developer's file has nothing wrong there. What is happening?
    The position indexes the document text the server received, which is rarely the file as written. Clients minify operations onto a single line, assemble them from several fragment strings, or inject them from a bundle. Reconstruct the sent text first — from a raw log or a captured request — and read the position against that; against the source file the numbers are meaningless.
  • Escaping user input before interpolating it into the document would fix this. Would you accept that?
    No. Escaping puts the correctness of every request on a hand-written encoder that has to be right for every input forever, and it leaves the document text varying per request, so it can never be extracted, tested or registered ahead of time. Sending values in the variables map removes the failure mode structurally instead of defending against it.

saying these in an interview costs you the question

  • Blames a resolver or the database for a syntax error
  • Assumes a schema change or deploy caused it
  • Escapes user input into the query string instead of using variables
  • Reads line and column against the source file, not the sent text
  • Thinks single quotes delimit a GraphQL string
  • Hunts for a latency cause inside the parser

context