Why can a shallow GraphQL document still be expensive enough to need a cost limit?
answer
- Shape is not volume
- Count objects, not levels
- Every list slice multiplies what is below
- 213 × 184 × 466 in four levels
basics
~20 sCost follows how many objects a document asks the server to produce, not how deeply it nests. Three levels of list fields, each requesting a few hundred items, multiply into millions of objects while staying shallow.
solid answer
~50 sNesting depth bounds the **shape** of a document; it says nothing about the **volume** the document asks for. Every list field carries a slice argument — `first: 213` by widespread convention, not by any rule in the GraphQL specification — and each nested level multiplies the level above it. A four-level selection over a solar-array telemetry graph asking for 213 sites, 184 inverters per site and 466 readings per inverter is 18,263,472 reading objects, from a document a reader would call small. Complexity scoring prices that: give every field a configured weight, multiply a list field's subtree by the slice it requests, sum the document, and compare the total to a budget before anything executes. The score is an upper bound on the work the document authorises — exactly the number a public endpoint needs in order to refuse a request cheaply instead of discovering the cost halfway through execution.
code
graphql · 12 linesquery SiteTelemetry {
sites(first: 213) {
name
inverters(first: 184) {
serial
readings(first: 466) {
wattage
capturedAt
}
}
}
}go deeper
Be ready to say, in one sentence, that a GraphQL document's cost comes from the number of objects it asks for, and that nested list slices multiply. Have a concrete example ready with real numbers.
Explain the mechanics: where the multipliers come from, why a nesting cap and a cost score measure different things, and why the score has to be computed before any resolver runs rather than after.
Show you know the score is an upper bound rather than a latency estimate, and be able to say what that costs you in practice — documents refused that would have been cheap, and weights that need calibrating against measured work.
Own the framing that a single GraphQL endpoint has no natural unit of work, so requests-per-minute is not a usable currency. Be ready to argue what unit you would bill in and why the organisation should pay for that machinery.
## Shape is not volume A GraphQL document is a description of the response a client wants, and two properties of that description are easy to confuse. Its **shape** is how many levels of selection it nests and how many fields it names. That is a property of the text, and it is what a nesting cap or a field-count cap measures. Its **volume** is how many objects the server must produce in order to answer it. That is a property of the *data behind* the text, and nothing in the text's size tells you about it. A control that reads only the shape can happily admit a document that asks for the entire database, because the expensive thing about a GraphQL document is almost never how far down it goes. It is multiplication. ## Where the multiplication comes from A field that returns a list is normally given a slice argument — `first`, `last`, `limit` — saying how many items the client wants. Those names are convention: `first`/`last` come from the Relay server specification's connection convention, and the GraphQL specification itself defines no pagination argument and no cost model at all. Whatever the client selects *underneath* a list field is produced once per item in that slice. Nest a second list field inside the first, and its slice multiplies. Nest a third, and you have a product of three factors. ## A worked case Take a solar-array telemetry graph: sites own inverters, and inverters emit power readings. ```graphql query { sites(first: 213) { inverters(first: 184) { readings(first: 466) { wattage capturedAt } } } } ``` Four levels of selection, five named fields, forty-odd characters per line. A nesting cap of six waves it through without comment. The server, meanwhile, is being asked for 213 sites, 213 × 184 = 39,192 inverters, and 39,192 × 466 = 18,263,472 readings, each carrying two scalars — roughly 36.5 million leaf values. Change one digit, `466` to `999`, and the document is the same length while the work more than doubles. This is why "the document looked small" is never a defence. The client controls the multipliers, and a text-length or shape check cannot see them. ## Why the layers below do not save you Three tempting answers fail here. *The response will be huge, so something downstream will kill it.* By the time the response is huge the work is already done: rows read, objects allocated, the request thread or event loop occupied. An abuse control that triggers after the cost has been paid is not a control. *The database will refuse an unreasonable query.* The database is not asked one unreasonable query. It is asked 39,192 reasonable ones — each individually fast, collectively an outage. Nothing at that layer sees the aggregate. *Rate limiting by requests per minute will bound it.* A limit of sixty requests per minute is meaningless when a single request can be worth millions of object fetches. Requests are not a unit of work on a GraphQL endpoint; that is precisely the property that makes one endpoint with one HTTP verb harder to police than a hundred REST routes with fixed response shapes. ## What a cost score adds Complexity scoring replaces "how big is the text" with "how much work does this text authorise". The mechanism is deliberately crude and entirely static: 1. Every field in the schema is given a **weight** — usually 1 by default, higher for fields backed by an expensive call. 2. A list field's subtree cost is **multiplied** by the slice its arguments request. 3. The document's total is compared to a **budget**. Over budget, the server answers with an error and executes nothing. For a public endpoint, the value of this is that the rejection is cheap. Parsing and scoring a document costs microseconds; executing it costs an outage. And because scoring happens before any resolver runs, there is no half-served response, no orphaned backend work, and no partially charged downstream. ## What the score is and is not The number is an **upper bound**, computed from the document and its variables alone. It says: *no matter what the data looks like, this document cannot ask for more than this.* It is not a latency prediction and not a measurement — a document scored at 400,000 may return in 20 ms because a cache was warm, and a document scored at 900 may be slow because one field is a slow external call whose weight was set too low. That gap between authorised work and actual work is the source of both the control's strength (it is decidable before execution) and its main operational annoyance (it can refuse documents that would have been cheap). Two other things it is not: it is not authorization — a cheap document can still ask for data the caller may not see — and it is not part of the GraphQL specification. Every weight, every multiplier convention and every budget in this area is server configuration and local policy.
- Does the GraphQL specification define the `first` argument or a cost model for it?No. The specification defines the type system, document validation and the execution algorithm; it defines no cost model, no weights and no pagination arguments. `first` and `last` come from the Relay server specification's connection convention, and plenty of schemas use `limit` or `take` instead. Any weight table, multiplier rule or budget you see is server configuration, so two servers can score the same document differently and both be correct.
- Why not let the backend refuse the work instead of scoring the document up front?Because the backend never sees one outrageous request — it sees tens of thousands of individually reasonable ones, and refusing one of them mid-flight leaves you with a half-executed operation, wasted work and a partial response. Scoring rejects before any resolver runs, at parse-and-validate cost, and returns a single clear error. It also gives you a number you can log and budget against, which a scattering of downstream failures does not.
A shipping manifest can be one page long and still order forty containers. Counting the pages tells you nothing about the freight bill.
saying these in an interview costs you the question
- Claims a nesting cap already bounds the work
- Equates cost with how deep the document nests
- Assumes a short document means a small response
- Says the database will refuse anything unreasonable
- Counts requests per minute as a unit of work
- Treats page size as purely a client-side concern