How do you catch, before release, that a client's GraphQL documents fail against the deployed schema?
answer
- Run the server's pre-execution check early
- Collect the operations the client can send
- Which schema you compare against decides everything
- Two artefacts deploy on separate clocks
- Additive schema first, then the client
basics
~20 sExtract every executable document from the client's source, validate each against the schema currently deployed to the target environment, and fail the build on any error. Validating against your own branch's schema proves only self-consistency.
solid answer
~50 sThe check itself is simple: collect the client's operations — from `.graphql` files and from document strings in source — and run each through validation against a schema, failing the build on the first error. The judgement is in which schema you validate against. A client and a server deploy independently, so the only schema that matters is the one that will be live when the client ships: fetch it from the target environment or from a registry that publishes the deployed version. Validating against the schema built from the same pull request catches typos and nothing else, because both sides move together in the repository and neither reflects production. That constrains release order: an additive schema change deploys first, then the client that uses it. Ship them the other way and every request is rejected at validation before a resolver runs.
code
pseudocode · 11 linesschema = fetchDeployedSchema(env: "production") // hard-fail if unavailable
documents = extractDocuments(sourceDir: "client/src")
failures = []
for doc in documents:
errors = validate(schema, doc)
if errors is not empty:
failures.add(doc.operationName, doc.file, errors)
if failures is not empty:
fail("%d client documents invalid against the deployed schema" % failures.size)go deeper
Know that a client's operations can be checked against a schema in the build, and that an unknown field is rejected before anything executes rather than coming back as a null value.
Explain the two halves — extracting documents from client source with fragments resolved, then validating each against a schema — and what validation proves about a document.
The answer hinges on which schema you compare against and on deploy order. Show you validate against the environment's live schema, fail hard when the fetch fails, and ship additive schema changes ahead of the clients that use them.
Own the release contract across independently deployed clients and servers: how long old client versions stay in the field, what that implies for the oldest schema you must validate against, and who is accountable when the two clocks diverge.
## The check A GraphQL server validates every incoming document against its schema before executing it. That check is deterministic and needs nothing but the document and the schema — which means you can run it in a build, days before a client reaches a user. The job has two halves. **Extraction**: gather every executable document the client can send — standalone `.graphql` files, and operation and fragment strings embedded in source, usually recognised by a tagged template or a well-known helper call. **Validation**: run each collected document against a schema, and fail the build with the operation name and location on any error. Fragments must be resolved across files first, since a document that references a fragment defined elsewhere is incomplete on its own. A useful side effect: documents assembled at runtime by string concatenation cannot be extracted, so this gate quietly pushes a client toward static documents — which is also the precondition for registering documents ahead of time. ## Which schema, and why it is the whole question This is where candidates separate. Validating the client's documents against the schema built from the same pull request feels rigorous and proves almost nothing, because both artefacts came from the same commit. It catches a misspelled field. It cannot catch the failure that actually happens in production, which is a **version skew** between two independently deployed things. The schema to validate against is the one that will be serving when this client release is live: - fetch the SDL from the target environment's endpoint, or - pull the schema version a registry records as deployed there. If you cannot force clients to upgrade — a mobile app, a long-lived browser session, an embedded terminal in a warehouse — a document must additionally be valid against every schema version still in the field, which is an argument for validating against the *oldest* supported deployed schema as well as the newest. ## The ordering assumption that breaks A warehouse inventory graph added `Pallet.lastCycleCountAt`, and the handheld scanner client selected it in the same sprint. CI was green: the client's documents validated against the branch's schema, which of course contained the new field. The client build was promoted on Tuesday; the server deploy was scheduled behind a migration and landed on Thursday. At the shift-change peak of roughly 1,200 requests per minute, every scanner request came back as a request error. Not a nulled field, not partial data: an unknown field is caught during validation, so execution never begins, `data` is absent entirely, and there is no degraded mode to fall back to. Roughly 41 minutes of scanning was lost before the schema change was rushed out. The rule that falls out is unglamorous and absolute. **Deploy the additive schema change first, then the client that uses it.** Removals run in the opposite order: stop clients selecting the field, wait out the deployed client versions, then remove. The CI gate enforces the first half automatically — a client validating against the live schema simply cannot pass until the server change is out. ```json { "errors": [ { "message": "Cannot query field 'lastCycleCountAt' on type 'Pallet'.", "locations": [{ "line": 6, "column": 5 }] } ] } ``` Note the absence of a `data` key and of a `path`. That shape is diagnostic: it says the request never executed. ## What this gate cannot see Validation is a structural check of a document against a set of types. Passing it means every selected field exists on the type it is selected on, every variable is declared and used compatibly, every fragment can apply, and leaf selections are shaped correctly. It says nothing about: - **authorization** — the field exists; whether this client's token may select it is application logic at execution time; - **cost and depth limits** — an unspecified control most production servers add; a perfectly valid document can be rejected for being too expensive; - **runtime nulls and field errors** — a nullable field that is always null in practice validates cleanly; - **semantics** — a field that now returns quantities in cases rather than units is invisible to validation; - **deprecation** — selecting a deprecated field is valid, and stays valid until it is removed. Surfacing that is a linting or usage question, and the removal policy itself belongs to schema evolution work. ## Making it operational Run the gate on every client build, not nightly, so the failure lands on the change that caused it. Report failures by operation name and file position; a wall of errors from one missing fragment is unactionable. Cache the fetched schema per pipeline run so a hundred jobs do not hammer the endpoint, but never cache it across days — a stale cached schema recreates exactly the skew the gate exists to prevent. And keep the fetch a hard failure: if the environment's schema cannot be retrieved, fail the build rather than silently falling back to the branch's copy, which turns the gate back into the self-consistency check you started with.
- Your pipeline validates client documents against the schema built from the same pull request. What does that miss?Every skew failure, which is the only kind that reaches users. Both artefacts come from one commit, so the check is self-consistent by construction and passes for a field that exists only on the branch. It catches typos. Validate against the schema fetched from the target environment — and against the oldest schema still serving, if old client versions cannot be forced to upgrade.
- A client document passes this CI gate and still fails in production. Name two causes.Authorization and cost. Validation proves the selected fields exist and are used compatibly; whether this caller's token may read a field is application logic evaluated during execution, and a valid document can still exceed a depth or cost limit and be rejected before execution. Runtime field errors and changed field semantics also pass validation cleanly.
- How does this gate change the order in which you deploy a schema change and a client change?It enforces additive-first. A client selecting a new field cannot pass validation against the live schema until the server change is deployed, so the sequence becomes schema, then client. Removals invert it: stop clients selecting the field, wait until no deployed client version still sends it, then remove it from the schema.
- Some documents in the client are built by concatenating strings at runtime. What do you do?Treat it as a defect in the client rather than a limit of the gate. A concatenated document cannot be extracted, so it is invisible to the check and would also be ineligible for registering documents ahead of time. Move the variation into variables and directives such as `@skip` and `@include`, so the document text stays static and analysable.
saying these in an interview costs you the question
- Validates client documents against the branch's own schema
- Assumes an unknown field returns null with an error
- Thinks passing validation proves the caller is authorized
- Deploys the client before the additive schema change
- Silently falls back to a local schema when the fetch fails
- Caches the fetched schema for days across pipeline runs