Why does a cacheable GraphQL GET carry an operation identifier rather than the document text?
answer
- The query string is not a hint
- Bytes are compared, meaning is not
- Formatting churn should not cost a miss
- Short and constant beats long and variable
- Canonical variables or the key fragments
basics
~20 sSize and key stability. A percent-encoded document is long enough to approach URL limits, and because the whole URL is the cache key, any difference in formatting or parameter order between clients mints a second entry for the same operation.
solid answer
~50 sFor a GET, a shared cache keys on the method plus the complete URI, so whatever sits in the query string *is* the cache key. Document text makes a bad one on two counts. It is long — percent-encoding costs three characters per escaped brace, quote or newline, and URL and header limits along the path are commonly a few kilobytes. And it is unstable: pretty-printed versus minified, an injected `__typename`, a reordered field or a different parameter order all produce a new key for an identical result. A short identifier — from a hash handshake or a build-time manifest — is constant in size, identical across clients shipping the same operation, and changes only when the operation genuinely changes. Persisted identifiers are a widespread convention, not part of the GraphQL specification. Variables still ride the URL and still belong in the key, so they must be serialized canonically or they fragment it too.
code
graphql · 7 linesquery SegmentMatches($sourceHash: String!, $targetLocale: String!, $minScore: Int!) {
segmentMatches(sourceHash: $sourceHash, targetLocale: $targetLocale, minScore: $minScore) {
score
target
unit { id updatedAt }
}
}go deeper
Know that on a GET the whole URL is what a shared cache stores the response under, and that sending a short identifier instead of the full operation text keeps that URL small and the same across clients.
Explain both halves precisely: percent-encoding inflation against real URL and header limits, and byte-level key instability from formatting, field order, parameter order and encoding choices. Say clearly that persisted identifiers are convention, not specification.
Demonstrate the discipline that makes it pay: canonical variable serialization across every client, a shared derivation of identifiers so two applications share one entry, and measurement of hit rate against distinct-key count rather than assuming the swap worked.
Own it as a cross-team contract — one identifier scheme, one variable-serialization rule, and a deploy ordering that guarantees the server knows an identifier before any client can send it — and weigh the readability and tooling cost against the cache economics.
## The URL is the whole key A shared cache stores a response under the request that produced it, and for a GET that means the method plus the complete URI — path and query string, byte for byte, before any application ever parses it. That is what moving a GraphQL query onto a GET buys: the operation stops hiding in a body no intermediary reads, and becomes part of an addressable identity. It also means the *content* of that identity matters enormously. Whatever you put in the query string is not a hint about the key; it **is** the key. So the choice between carrying the document text and carrying a short identifier for it is a choice about the quality of your cache key, and secondarily about whether the request fits in a URL at all. ## Why the document text is a poor key **It is long.** Percent-encoding is not free: every escaped character costs three (`{` becomes `%7B`, a space `%20`, a newline `%0A`), and GraphQL documents are dense with braces, quotes, parentheses, `$` and newlines. On a translation-memory graph, a `SegmentMatches` operation with two named fragments runs to about 2,900 characters of source and lands near 4,400 once escaped. URL length has no specified ceiling, but browsers, reverse proxies, CDNs and server request-line and header buffers all impose one — configurable defaults in the region of 4 to 8 kilobytes are typical — so a document of that size is already sitting on the edge of the cliff before variables are added. **It is unstable.** The cache compares bytes, not meaning. Every one of these produces a different key for the same operation: - one client ships the document pretty-printed, another minified; - a codegen step starts injecting `__typename` into selection sets; - a developer reorders two fields or renames a fragment; - the client emits `query` before `variables` in the query string, and another emits them the other way round; - one encoder writes spaces as `%20`, another as `+`. None of those changes what the server returns. All of them halve or fragment your hit rate. An invented but entirely typical outcome: the same `SegmentMatches` read shipped by a web client and a desktop client differed only in indentation, and the shared cache held two entries for it, with the measured hit rate stuck near 51% for a read that should have been near-perfectly shared. ## Why an identifier is a good key An identifier — a content hash produced by a hash handshake at runtime, or a name plus version pulled from a manifest built at compile time — fixes both problems at once. It is **short and constant**: sixty-odd characters regardless of how large the operation is, so the URL length stops tracking document complexity and starts tracking only the variables. It is **stable across clients**: two applications that ship the same operation derive the same identifier, so they share one cache entry instead of minting one each. And when it is derived from the document, it is **stable for exactly as long as the operation is unchanged**. A cache miss then coincides with a real change to what the client is asking for, which is precisely when a miss is correct. Formatting churn no longer costs you anything, because formatting no longer reaches the URL. There is a side benefit worth one sentence in an interview and no more: an endpoint that only accepts known identifiers stops being an arbitrary-query surface, because a URL naming an unknown identifier is simply rejected. Say plainly which parts of this are specified. GraphQL over HTTP describes how a request's parameters ride a URL. Persisted operation identifiers — the hash handshake and the build-time manifest alike — are a **widespread convention**, not part of the GraphQL specification, and even the parameter name that carries the identifier varies between servers. ## Variables are still in the key, and they can still fragment it Swapping the document for an identifier does not make the URL constant, and should not: two callers asking for different target locales must get different responses, so the variables belong in the key. But they arrive as a JSON object serialized to text, and text has more freedom than the object does. `{"targetLocale":"de-DE","minScore":72}` and `{"minScore":72,"targetLocale":"de-DE"}` are the same variable map and two different cache entries. So the discipline that makes an identifier pay off is canonical serialization on the client: stable key order, no incidental whitespace, one agreed encoding. Teams that skip this get the short URL and keep the fragmentation. ## What it costs An identifier is not readable. A URL you could once paste into an explorer becomes opaque, and diagnosis now needs the manifest or the hash-to-document store to interpret it. There is a lookup step on the server, a deploy-ordering obligation so the server knows an identifier before a client sends it, and — for the runtime handshake variant — a cold round trip the first time an identifier is unknown. And none of this makes the response cacheable on its own. A short, stable URL only makes the response *addressable*; whether a shared cache may store it, and for how long, and whether it is viewer-specific, are separate decisions the server still has to signal.
- If the identifier is a hash of the document, what happens to cached responses when the operation changes?The hash changes, so every request for the new version carries a new URL and misses. That is correct behaviour rather than a cost: a miss now coincides exactly with a real change to what the client asked for. The old entries are never served to the new operation and simply age out, so there is no invalidation step to run at deploy time for the operation change itself.
- The URL is now short and stable. Does that make the response cacheable?No. It makes the response *addressable* — a shared cache can now form a key for it at all. Whether that cache may store it, for how long, and whether the content is viewer-specific are separate decisions the server still has to signal, and a response that varies by viewer must not be stored under a key that does not distinguish viewers.
- What debuggability do you lose by putting an identifier in the URL instead of the document?The URL stops being readable: you can no longer see what was asked from an access log or a pasted link, and diagnosis needs the manifest or the hash-to-document store to interpret it. Teams usually compensate by logging the resolved operation name server-side after lookup, rather than by adding a redundant name parameter that could itself vary and fragment the key.
saying these in an interview costs you the question
- Thinks the cache normalizes GraphQL syntax before keying
- Assumes reformatting a document is cache-neutral
- Says persisted identifiers are part of the GraphQL specification
- Leaves variables out of the cache key
- Serializes variables with unstable key order
- Believes a short URL alone makes a response cacheable