skip to content

Which parts of a GraphQL document's text does the parser ignore?

level: juniorimportance: should knowfreq 44%

answer

  1. The lexer throws some characters away
  2. Punctuation that separates nothing
  3. Not JSON: no trailing-comma error
  4. Hash to end of line, no block form

basics

~20 s

Whitespace, line breaks, a leading byte-order mark, comments running from a hash to the end of the line, and commas. Commas are ignored tokens rather than separators, so a document parses identically with all of them, some, or none.

solid answer

~50 s

GraphQL's lexer splits the source text into tokens and throws away a category the specification calls **ignored tokens**: a leading byte-order mark, whitespace, line terminators, comments and — the one that surprises people — **commas**. A comma is exactly as meaningful as a space, so `{ streetAddress, priceCents }` and `{ streetAddress priceCents }` are the same document, and a doubled or trailing comma is legal rather than an error. That is the opposite of JSON, where commas are structural. A comment starts with `#` and runs to the end of the line; there is no block-comment form, and a comment is discarded at lex time, so it never reaches a response and is not introspectable. Everything else carries meaning: punctuators such as `{ } ( ) : $ @ ! [ ] ...`, names, and int, float, string and enum literals. Names themselves are strict — letters, digits and underscores, never starting with a digit.

code

graphql · 6 lines
graphql
query ListingCard {          # a comment runs to end of line
  listing(id: "LST-4417") {
    streetAddress,
    priceCents,
  }
}

go deeper

for a junior

Be ready to say plainly that commas and line breaks in a GraphQL document carry no meaning, and that a comment starts with a hash and runs to the end of the line. Interviewers use this to check you have actually read a document rather than only copied one.

for a middle

Name the whole ignored-token category — byte-order mark, whitespace, line terminators, comments and commas — and contrast it with JSON, where a comma is structural and both a missing one and a trailing one are parse errors.

for a senior

Explain why insignificant commas make generated and machine-edited documents safe to assemble, and why comments never survive lexing, so anything an API consumer must read belongs in a schema description rather than a hash line.

for a principal

Own the tooling consequence: because formatting carries no meaning at all, document formatting and linting can be standardised across every client team with zero behavioural risk, and generated documents produce stable diffs.

## Two layers: lexing, then parsing GraphQL's grammar is defined in two stages. A **lexical** stage turns the raw source text of a document into a stream of tokens, and a **syntactic** stage assembles those tokens into definitions, selection sets, arguments and so on. The lexical stage is where the question of "what does the parser ignore" is settled, and the specification answers it with an explicit list. The **ignored tokens** are: a Unicode byte-order mark at the start of the source, whitespace (space and horizontal tab), line terminators, comments, and commas. Ignored tokens may appear between any two real tokens and are discarded — they never become part of the parsed document. Everything that is left is meaningful: punctuators (`{` `}` `(` `)` `[` `]` `:` `$` `@` `!` `|` `&` `=` and the spread `...`), names, and the literal forms for integers, floats, strings, booleans, null and enum values. ## The comma surprise Because the comma sits on that list, a comma in a GraphQL document means precisely what a space means: nothing. All of these are the same operation over a real-estate listings graph: ```graphql query { listing(id: "LST-4417") { streetAddress priceCents } } ``` ```graphql query { listing(id: "LST-4417") { streetAddress, priceCents } } ``` ```graphql query { listing(id: "LST-4417") { streetAddress,, priceCents, } } ``` A doubled comma is legal. A trailing comma is legal. A comma between an argument list and the selection set that follows it is legal. None of them changes the token stream that reaches the parser. This is the reverse of JSON, where a missing comma and a trailing comma are both syntax errors, and it catches out anyone who assumes the two formats share a grammar because they both use braces. The design is deliberate and it pays off in machine-generated text. A tool that concatenates selected fields, or a person editing a long selection set by hand, does not have to track which element is last. Reordering two lines produces a one-line diff instead of a three-line one. Formatters can therefore normalise commas however a team likes, and the behaviour of the request is untouched. ## Comments A comment begins with `#` and runs to the next line terminator. There is no `/* ... */` form; writing one is a syntax error. Because a comment is an ignored token, it may appear wherever whitespace may — inside an argument list, between a field and its selection set, above an operation. What matters operationally is that a comment does not survive. It is discarded when the document is lexed, so it never appears in a response, never reaches a resolver, and is not visible through introspection. Commentary that has to be readable by consumers of the API is written as a **description** — a string attached to a definition in the schema's SDL, which is retained and exposed — not as a `#` line. Confusing the two is the usual beginner mistake: `#` is for whoever opens the file, a description is for whoever reads the schema. ## Layout is free, inside string literals it is not Since whitespace and line terminators are ignored, a document may be minified onto one line — which is exactly what most clients send over the wire — or spread across forty indented lines in a repository. The two parse identically. That freedom has one boundary: **inside a string literal, characters are content, not layout.** A space inside `"221B Baker Street"` is part of the value. The other boundary is that two adjacent names still need something between them: `streetAddress priceCents` is two fields, `streetAddresspriceCents` is one name. One practical consequence of free layout is worth carrying into the next topic: when a server reports the line and column of a syntax error, those numbers index the text the server actually received, which for a minified document is usually line 1. ## Names are not free The tokens that are *not* ignored come with their own strictness. A GraphQL **Name** — used for operations, fields, arguments, types, fragments, directives and enum values — matches `[_A-Za-z][_0-9A-Za-z]*`. Letters, digits and underscores only; never a leading digit; no hyphens, dots or spaces; and names are case-sensitive, so `priceCents` and `pricecents` are different. A legacy column called `list-price` or `2024_total` cannot become a field name verbatim; it has to be renamed in the schema. Names beginning with two underscores are reserved by the specification for the introspection meta-fields, so a schema must not define its own. ## How this comes up in an interview Usually as a fast filter: "are commas required?", "how do you comment a query?", or a snippet with a doubled comma and the question of whether it compiles. The point of the question is not the trivia — it is whether you have read the grammar rather than pattern-matched GraphQL onto JSON. Answer with the category, not the anecdote: commas are ignored tokens, alongside whitespace, line terminators, a leading byte-order mark and `#` comments.

  • What characters may a GraphQL name contain?
    A name matches `[_A-Za-z][_0-9A-Za-z]*`: letters, digits and underscores, never starting with a digit, and no hyphens, dots or spaces. Names are case-sensitive. Names beginning with two underscores are reserved by the specification for introspection meta-fields, so a schema must not define its own field or type with that prefix.
  • Does GraphQL have a block-comment form?
    No. The only comment form is `#` to the end of the line; a `/* ... */` block is a syntax error. Multi-line commentary is several `#` lines. Documentation that must reach API consumers is written instead as a string description on a schema definition, which is retained and readable through introspection, whereas a comment is discarded when the document is lexed.
  • If commas mean nothing, why do people still write them?
    Habit from JSON and from most programming languages, plus readability when a selection set or an argument list is collapsed onto one line. Because they are ignored tokens, the choice is purely a formatting convention: a team can have a formatter add them, strip them or leave them alone with no effect whatsoever on the parsed document or the response.

Ignored tokens are the punctuation of a shopping list: the commas and line breaks help you read it, but the shop assistant fills exactly the same basket whether you wrote the items on one line, on six, or with commas everywhere.

saying these in an interview costs you the question

  • Says commas separate fields and are required
  • Writes /* ... */ block comments in a document
  • Expects a trailing comma to be a syntax error
  • Thinks a # comment reaches the response or introspection
  • Assumes a field name may contain a hyphen
  • Believes a minified one-line document parses differently

context