skip to content

Which Automatic Persisted Queries requests pay two round trips, and why does that keep recurring?

level: seniorimportance: should knowfreq 46%

answer

  1. Ask how many stores there are
  2. Warm-up is a rate, not an event
  3. New process, new ignorance
  4. One shape decays, the other is flat
  5. Hashed string must equal uploaded string

basics

~20 s

One request pays two trips per document per document-store — and the store is usually per server process, so the cost repeats on every replica, every deploy, every restart and every scale-out. Misses are a steady rate, not a one-off warm-up.

solid answer

~50 s

The double round trip is paid once per unique document per **store of known documents**, and the trap is assuming there is one store. A typical server keeps a bounded in-memory map per process, so the real unit is once per document per replica, and it resets on every restart, deploy, scale-out or eviction. Misses are therefore a continuous low rate with a spike at each release, not a warm-up that ends. A separate, permanent miss loop appears when the string a client hashes is not the string it later uploads: the server files or verifies against different text, so every request costs two trips forever. Diagnose by shape — cold-start misses decay after a deploy and correlate with instance age, while a mismatch holds a flat ratio all day. Remedies: back the store with a shared one, prime on boot, and make hash and upload share one serialization.

code

pseudocode · 8 lines
pseudocode
response = post(endpoint, { operationName, variables, extensions: persistedQuery(hash) })

if response.errors contains message "PersistedQueryNotFound":
    # second round trip: same request plus the document text
    response = post(endpoint, { operationName, variables, query: documentText,
                                extensions: persistedQuery(hash) })

return response

go deeper

for a junior

Know that the extra round trip happens the first time a server sees a hash, and that a server which has just restarted has seen nothing. That alone explains why misses reappear after every release.

for a middle

Explain the unit of cost — once per document per store — and enumerate what resets a store: deploys, restarts, scale-outs and eviction. Be ready to say why variables never multiply that count.

for a senior

Demonstrate diagnosis: read not-found responses as a share of traffic across a release boundary, separate a decaying cold-start curve from a flat mismatch ratio, and pick the matching remedy rather than guessing between them.

for a principal

Own the tradeoff. A shared document store removes per-replica relearning at the price of a request-path dependency; a warm-up pass shifts the cost off users; and on a small fleet the honest answer is to measure the trickle and leave it alone.

## Where the cost actually lands The handshake costs one extra round trip for a hash the server does not recognise. The interesting question is how often "does not recognise" happens in a running system, and the answer turns entirely on a word people gloss over: the server's **store of known documents**. Most implementations default that store to a bounded in-memory map inside each server process. Nothing is shared. So the unit of cost is not "once per document", it is *once per document per process*, and it resets on every event that gives you a fresh process: * a deploy — every replica starts empty; * a restart, a crash, a rescheduled instance; * a scale-out, where a brand-new replica must learn every document a client sends it; * an eviction, when the map is bounded and the number of distinct documents in flight exceeds it. That converts what sounds like a one-time warm-up into a **rate**. A healthy system shows a small continuous trickle of not-found errors, with a visible step at each release and a smaller bump at each scale-out. A candidate who says "you pay it once and then it is free" has not run one of these. ## An incident that makes the shape concrete A hospital appointment graph is served through a gateway fronting a 62-subgraph supergraph, with 41 gateway replicas behind a load balancer. The team adopted the hash-first scheme along with a shared document store, and the rollout plan had two steps: switch the shared store on, then ship the client bundle that starts sending hashes. The steps went out in the other order — an ordering assumption that broke, because the client bundle was on a faster release train and nobody had written the dependency down. For a few minutes each of the 41 replicas learned each of the 23 documents in the bundle independently: 738 not-found responses in the first 90 seconds, and two distinct latency modes on the appointment-list view. Nothing broke — every user got a correct answer, one round trip later — which is exactly why the miss went unnoticed until someone looked at the error counter. What happened *after* the shared store was enabled is the lesson: the trickle went to nearly zero, but the deploy-time step did not vanish — the fleet now learns each document once instead of 41 times, so the spike shrinks by roughly the replica count rather than disappearing. ## The other failure: a miss ratio that never decays There is a second, more embarrassing pattern. Suppose a client computes its hash over the document as written in source, but the code that actually issues the request serialises a parsed form of the document — reordered whitespace, normalised indentation, a stripped trailing newline. Now the digest the client claims and the digest of the text it uploads are different values. The server, which verifies before storing, either refuses to store or files the text under the hash it computed rather than the one claimed. Either way the client's next hash-first request misses again. That is a permanent two-trip system: every request doubled, forever, for every user on that client build. It is worth remembering that a single trailing newline is enough to do it — the document in the previous section hashes to `fc7cd86c…` without one and `28ce4466…` with one. **Telling the two apart is the diagnostic skill.** Cold-start misses have a characteristic shape: they spike at a deploy, decay over minutes, correlate with instance age, and fall as a share of traffic as the fleet warms. A hash mismatch has no shape at all — a flat ratio near one miss per request, constant across the day, confined to one client build or platform, and completely indifferent to deploys of the server. Plot not-found responses as a fraction of requests over a release boundary and the two are unmistakable. ## What to do about it * **Share the store.** Backing the document store with a shared one turns "once per document per replica" into "once per document per fleet", which is the single biggest reduction available. The cost is a lookup on the request path and a new dependency to fail; most deployments keep a small local cache in front of the shared one, so a warm document costs nothing extra. * **Size it honestly.** The store's population is the number of *distinct document texts your clients ship* — normally dozens. If that number is unbounded, something upstream is generating documents rather than reusing them, and no cache size helps. * **Prime on boot.** A new instance can preload from the shared store, or the deploy can drive a short warm-up pass, so the first real users are not the ones paying. * **Make hash and upload share one serialization.** The single string that is hashed must be the single string that is sent. This is a client-build property, and it is worth a test. * **Accept it when it is cheap.** With a handful of replicas and a dozen documents the trickle is negligible, and a shared store buys a dependency in exchange for latency you cannot measure. Knowing when *not* to fix this is part of the answer. ## Traps The common errors are believing there is one store behind the whole fleet; forgetting that a deploy empties an in-memory one; assuming a sticky load balancer removes the cost (it reduces per-client repetition, but every replica still has to learn every document, and stickiness collapses at a deploy anyway); and blaming the network for what is a constant, structural miss ratio.

  • Why doesn't a sticky load balancer solve the cold-hash cost?
    Stickiness keeps one client on one replica, so that client stops re-teaching the same document repeatedly — but every replica still has to learn every document from whichever client reaches it first, so the fleet-wide cost is unchanged. It also collapses exactly when it would help most: a deploy replaces every instance at once, so all the affinity is discarded and every document is unknown again.
  • How do you distinguish cold-start misses from a permanent hash mismatch in your metrics?
    By shape. Plot not-found responses as a share of requests across a release boundary. Cold starts spike at the deploy, decay over minutes, and correlate with instance age. A mismatch is flat — a near-constant ratio approaching one miss per request, unaffected by server deploys, and usually confined to a single client build or platform. The second needs a client fix; the first needs a shared store or a warm-up.
  • What does moving to a shared document store actually cost you?
    A lookup on the request path and a dependency that can fail or slow down, which is why most deployments keep a small in-process cache in front of it so warm documents never touch the network. You also inherit its consistency and eviction behaviour. In exchange the fleet learns each document once rather than once per replica, which is the largest single reduction available.
  • When is it right to leave the miss rate alone?
    When the arithmetic is small. A handful of replicas shipping a dozen or two distinct documents produces a trickle of extra round trips that is invisible next to normal request latency, and the users affected still get correct answers one trip later. Adding a shared store to fix that buys a new dependency in exchange for a saving you cannot measure — a bad trade you should be willing to name.

saying these in an interview costs you the question

  • Says the cost is one extra round trip, once, ever
  • Assumes every replica shares one document store
  • Forgets a deploy empties an in-memory store
  • Thinks a sticky load balancer removes the cold cost
  • Blames the network for a flat, constant miss ratio
  • Ignores eviction from a bounded document store

context