A GraphQL document validates cleanly and still returns 388,514 rows — why didn't validation stop it?
answer
- Ask what inputs the check needs
- Document and schema, nothing else
- An optional argument bounds nothing
- Put the bound in the type system
- Clamp where the data is fetched
basics
~20 sValidation compares a document only to the schema. A list field whose paging argument is optional is a legal selection however many rows come back, so nothing static can object. The bound must come from the schema and a resolver clamp.
solid answer
~50 sNothing was violated: the field is declared, its `first` argument is nullable with no default, and no literal is mistyped — so the pass had nothing to object to. Result size is not among its inputs. Three things are structurally invisible to it: the values behind variables, since the specification's rules are defined over document and schema alone and a value-dependent rule could not be reused across requests; the *meaning* of a type-correct value, such as a page size being unreasonable for this domain; and how much work a field will do, since no static analysis knows a list's length. The fix is layered. Put the bound in the type system where a static pass can see it — declare `first: Int!`, accepting that this breaks deployed clients and needs a rollout. Then clamp at execution, because only the resolver sees the coerced number. Then measure returned collection sizes, not just latency.
code
graphql · 4 linestype PayRun {
reference: String!
payslips(first: Int!, after: String): [Payslip!]!
}go deeper
Recall that a static pass only compares the document to the schema. If an argument is optional, omitting it is legal no matter how much data comes back, so nothing before execution can object.
Explain why the values behind variables are outside the pass's inputs, and how declaring a paging argument as Non-Null converts a production surprise into a failure the client's build catches.
Show the layered response — bound it in the schema, clamp it at execution, measure returned sizes — and be candid that making an argument required breaks deployed clients and therefore needs a rollout, not just a commit.
Own the general rule: a check that needs only document and schema belongs in the static pass; a check that needs a value, a row or an identity cannot live there at all. Then decide organizationally how much of that safety a shared schema is required to encode up front.
## The incident A payroll and benefits graph serves 4,182 employees. A reporting screen sends this operation: ```graphql query PayRunReport($runId: ID!) { payRun(id: $runId) { reference payslips { netPayCents employee { fullName } } } } ``` `PayRun.payslips` was declared as `payslips(first: Int, after: String): [Payslip!]!` — paging arguments were added, but as nullable ones so that nothing already deployed would break. The screen omits `first`. During the year-end run the field resolved every payslip attached to that run since 2019: 388,514 rows, a list that grew without a bound, one enormous response and a service under pressure. The document was valid, and validation was right to say so. `payslips` is declared on `PayRun`. `first` is nullable with no default, so omitting it breaks no rule. No literal is mistyped, no fragment is unused. Every rule the pass owns was satisfied. ## The three things the pass structurally cannot know **Values that arrive with the request.** The specification's validation rules are defined over a document and a schema. A rule that wanted to reject `first: 5000` could see it only when the client wrote it as a literal — and the moment the client sends `first: $pageSize` instead, the number is in the variables map, not in the document. A check that read variable values would also be unsafe to reuse: servers keep validated documents around precisely because validation's answer depends only on document and schema, and a value-dependent rule would make yesterday's answer wrong today. **What a type-correct value means.** Validation can prove `first` is an `Int`. It cannot know that anything above 200 is unreasonable here, that `E-4182` is not an employee in this company, or that a date range spans six years. Meaning lives in the domain, and the type system only carries as much of it as you chose to encode. **How much work a field will do.** Nothing static knows a list's length. A pre-execution cost check has to *assume* a page size for each list field, and a field with no required bound defeats that assumption too — the assumed number and the real one part company exactly where it matters. ## Where each check actually belongs **In the schema, first.** The strongest fix is to move the bound into the type system, because that is the only place a static pass can see it. Declare the argument `first: Int!` and every document that omits it becomes invalid at validation — the failure moves from production to the client's build. Be honest about the cost: making an argument required is a breaking change for every deployed client, so it lands as a new field, or behind a rollout that fixes callers first. **In a static analysis alongside validation, second.** Servers commonly run additional pre-execution rules of their own over the same document. They are not part of the specification, and the important thing here is that they inherit the same blind spot: they reason about shape, so they can only bound what the schema lets them count. **At execution, third and non-negotiably.** Whatever the schema says, the resolver clamps. A maximum page size applied where the data is fetched is the only check that sees the actual number after coercion, including the value that came from a variable, and it is the only one that holds when a client is wrong. ## The generalisation worth stating in an interview Ask what a proposed check needs as input. If it needs only the document and the schema, it can be a static rule and should be one, because a static rule fails fast, fails identically for everyone, and can run in an editor. If it needs a value, a row, a clock or an identity, it cannot be a validation rule at all, and pushing it there produces either a rule that is wrong or one that quietly cannot be reused. The last piece is operational. If nothing in the schema bounds a list, add the measurement that would have caught this: record returned collection sizes per field rather than watching only latency, because a list that grew without a bound looks fine on a latency chart right up until the day it does not.
- Why not simply add a server rule that rejects a page size above 200 before execution?Because it only works when the client writes the number as a literal in the document. Sent as a variable, the value is not in the document at all, so a document-and-schema rule cannot see it. A rule that reached into the variables map would also stop being a property of the document, which is what makes validated documents reusable. The check has to sit where the coerced value exists — at execution.
- What does making the paging argument Non-Null cost you?Every deployed document that omitted it becomes invalid, and they fail totally rather than degrading. Old mobile builds can stay pinned for months, so the change lands either as a new field alongside the old one, or behind a rollout that fixes and ships callers first, with the argument's usage measured until the tail reaches zero.
- If the schema cannot bound it, what would have told you sooner?Instrumentation of returned collection sizes per field, not only latency and error rate. A list that grows steadily produces a rising size distribution long before it produces a visible latency problem, and a per-field size histogram with an alert on its upper percentile turns a cliff into a slope you can see coming.
saying these in an interview costs you the question
- Blames validation for missing a data-dependent problem
- Proposes a validation rule that reads variable values
- Thinks a static pass can estimate result size exactly
- Adds a required argument with no client rollout plan
- Relies on the schema bound with no resolver clamp
- Watches only latency for an unbounded list field