With a build-time persisted document registry, what does a GraphQL request carry instead of the document text?
answer
- The text stops travelling
- Written down one release early
- Build extracts, server stores, client points
- A key into a map the server holds
- Hash of the normalized document text
basics
~20 sAn identifier for a document the server already holds, plus the variables. A build step extracts every operation from the client source into a manifest and publishes it ahead of the release; the shipped client carries identifiers, never query text.
solid answer
~50 sOnly an identifier and the variables. A build step walks the client's source, extracts every GraphQL document it finds, normalizes each one, and writes an identifier — usually a hash of the normalized text — into a **manifest** next to the document it stands for. That manifest is published to the server; the shipped client contains identifiers, not text. At runtime a request carries the identifier, the variable values, and an operation name when the stored document defines more than one operation. The server looks the identifier up and executes the stored document if it holds it. None of this is in the GraphQL specification — the manifest format and the parameter that carries the identifier are a convention between one client build and one server, so implementations differ on the field name. You buy a small request and a document the server has seen before; you give up asking anything new at runtime.
code
graphql · 9 linesquery TrackConsignment($waybill: String!) {
consignment(waybill: $waybill) {
status
currentLeg {
carrier
arrivedAt
}
}
}go deeper
Be able to say exactly what leaves the client: an identifier and the variable values, never the query text. Remember the mapping from identifier to document was published to the server before that client shipped.
Explain the build step in order — extract, normalize, hash, publish the manifest — and why the identifier is opaque to the server, which needs it only as a lookup key rather than as something it can decode.
Be ready to separate convention from specification here: GraphQL specifies none of it, so the manifest format, the parameter name and the not-found behaviour are agreements between your client build and your server rather than portable guarantees.
Own the consequence of fixing the document set at build time: nothing new can be asked at runtime, so ad-hoc tooling, one-off scripts and any third-party consumer need a deliberately designed separate path or none at all.
## The document leaves the client at build time, not at request time A GraphQL client normally carries its operations as text and sends that text on every request. A persisted-document registry moves the text out of the request entirely, one release ahead of time. The move happens during the build. A tool walks the client source, collects every GraphQL document it can find — string constants, tagged template literals, `.graphql` files — and for each one produces a pair: an identifier, and the exact document text that identifier stands for. Those pairs are written into a **manifest**, which is nothing more than a map from identifier to document. The manifest is published to the server. The client artefact that ships alongside it no longer contains query text; where the text used to sit, the build has substituted the identifier. The identifier is usually a cryptographic hash of the document text, because a hash needs no coordination — two builds of the same operation agree on it without talking to each other, and any change to the operation produces a different one. Nothing requires a hash: sequential numbers or `release/operationName` work equally well, because the identifier is opaque to the server, which needs only a key it can look up. What does matter is that the text is **normalized before it is hashed** — canonical whitespace, comments stripped, fragments emitted in a stable order — so that a reformat with no semantic change does not mint a second entry for the same operation. ## What travels Three things: the identifier, the variable values, and an operation name where the stored document defines more than one operation. The variables are unchanged from an ordinary request — they are per-call data, and folding them into the identifier would defeat the whole scheme by minting an entry per value. For the same reason, interpolating a value into the document text instead of passing it as a variable produces a distinct document, and therefore a distinct registry entry, for every value. There is no specified field name for the identifier. The GraphQL specification defines the language and the execution algorithm and says nothing about how a request travels; the GraphQL over HTTP specification governs the transport. Persisted documents sit on top as a convention, so the request parameter, the manifest format and the behaviour on a miss are agreements between one client build and one server rather than portable guarantees. Say that plainly in an interview — it separates people who have run this from people who have read about it. ## The server side is a lookup, not a negotiation On receiving a request the server takes the identifier, looks it up in whatever store holds the registry, and — if it finds an entry — proceeds exactly as if the client had sent the text: parse it (or reuse an already-parsed form), validate it against the schema, select the operation, execute with the supplied variables. If it finds no entry there is nothing to execute, and it answers with an error rather than a result. The absence of a fallback is the defining property. A related convention derives the same kind of identifier at runtime and lets a client teach the server the text when the server has not seen it; a build-time registry deliberately does not, because the text is not in the shipped artefact to send. That single design choice is where every deploy-ordering and retention consequence comes from. ## What the identifier does and does not identify It identifies **one document**, byte-for-byte after normalization. It does not identify a response, a user, a schema version or a result set. Interviewers probe this because the identifier looks so much like a cache key that people reach for it as one, and a shared cache keyed on the identifier alone ignores both the variables and the viewer. On a freight-tracking graph, one `TrackConsignment` identifier sent by two accounts is one key and two entirely different answers — which is how a cache serves one shipper the other's consignment row. Nor is the manifest discoverable. The server does not learn it by introspection or by watching traffic; something in the release process put it there. ## What you buy and what you give up The gains are concrete: requests shrink from kilobytes of document text to a short key; the server only ever executes documents that existed in a build; and a short key is small enough to ride in a URL, which reopens transport options that a body-carried document closes off. The loss is equally concrete: nothing new can be asked at runtime. Any tool, script or ad-hoc exploration that composes a query on the fly is either excluded or needs a separate path that accepts document text — an authenticated route, a non-production endpoint, or a second server. Deciding which of those you offer, and to whom, is a real design question rather than an afterthought.
- If the identifier is a hash of the document text, what does reformatting an operation cost you?A new identifier and a second manifest entry, even though nothing semantic changed. That is harmless when the manifest ships first — the old entry stays for older clients — but it is why extraction tools normalize before hashing: canonical whitespace, comments stripped, fragments in a stable order. Without normalization the registry accumulates several entries for one operation and per-operation usage metrics split across them.
- Is the operation identifier usable on its own as a cache key for the response?No. It names the document and nothing else. A shared cache keyed on the identifier alone ignores both the variables and the viewer — on a freight-tracking graph that is exactly how one shipper gets served another account's consignment row. A response key needs the identifier, the variable values, and whatever authorization scope the response depended on.
- How does the server know which operation to run when a registered document defines several?The same way it would with document text: the request carries an operation name and the server selects that operation from the stored document. Most extraction tools sidestep the question by emitting one operation per entry, which also makes usage metrics and retention decisions per-operation rather than per-file.
It is a will-call ticket rather than a parcel: the goods were delivered to the counter yesterday, and the ticket only names which crate to hand over.
saying these in an interview costs you the question
- Thinks the server can execute an identifier it does not hold
- Says the client computes the identifier at request time
- Treats the operation identifier as a response cache key
- Believes variable values are hashed into the identifier
- Assumes the server discovers the manifest by introspection
- Claims the GraphQL specification defines the request parameter