In a batched GraphQL request, why do per-operation caching and tracing stop working?
answer
- Downstream layers key on what they can see
- The composite key varies if any member varies
- Cacheable only when every member is
- One request, one log line, one span
- Rebuild the signal inside the executor
basics
~20 sBecause the unit the infrastructure sees is the batch, not the operation. One composite POST body means one cache key, one status, one access-log line and one span, so hit rates, error rates and latency can no longer be attributed to individual operations.
solid answer
~50 sEvery layer between client and server keys on what it can see, and in a batch that is the whole array. A shared cache or CDN sees one POST whose body is a composite of several operations and their variables, so it is uncacheable by default; even a body-aware cache would key the entire array, so one different variable in one member misses the whole batch and the hit rate collapses. The per-operation route to network caching - a GET whose URL carries a document identifier and variables, as Automatic Persisted Queries does with a hash - is foreclosed the moment operations are fused into one POST body. Observability degrades the same way: one HTTP request is one access-log line, one status and one server span, so per-operation latency and error rates exist only if the batch executor opens a child span and emits metrics per member. A client's normalized cache is unaffected.
code
pseudocode · 16 lineshandleBatch(members):
batchSpan = startSpan("graphql.batch", { memberCount: length(members) })
results = awaitAll(members.map(m ->
withSpan("graphql.operation",
{ operationName: m.operationName,
index: indexOf(m),
parent: batchSpan },
() -> execute(m))))
for (m, r) in zip(members, results):
emitMetric("graphql.operation.duration", r.durationMillis,
{ operationName: m.operationName, failed: hasErrors(r) })
batchSpan.end()
return resultsgo deeper
Take away the core idea: caches, logs and traces identify work by what the request shows them, and a batch shows them one composite request instead of several named operations.
Explain the mechanics both ways - a POST body that varies whenever any member varies cannot be usefully keyed, and one HTTP request yields one status, one log line and one span unless the server adds more.
Show what you would rebuild: child spans and metrics per member inside the executor, per-member logging, and an honest account of why a personalized batched response should not sit behind a shared cache at all.
Own the ledger. Batching trades header overhead for lost cacheability and lost attribution, and the loss is silent because the dashboards keep rendering, so make the decision explicitly and set the instrumentation as a platform requirement if you allow it.
## Everything downstream keys on what it can see A cache, a log pipeline and a tracing system all identify work by whatever the request exposes. For an unbatched GraphQL request that is one operation: one name, one document, one variables map, one outcome. Batching replaces that unit with a composite, and every layer that was reasoning about operations silently starts reasoning about batches instead. Nothing errors; the signal just gets coarser. ## Why network caching goes away A batched request is a POST, and a shared cache keys on method and URL, so it is uncacheable without a bespoke rule. Suppose you write that rule and key on the body as well. Now consider a supporter dashboard on a charity donations graph sending nine members, one of which carries `donorId`. Every distinct donor produces a distinct body, so every batch is a unique key even though eight of its nine members are byte-identical across all users. Cacheability composes multiplicatively in the wrong direction: the batch is cacheable only when *every* member is, and its key varies whenever *any* member's variables vary. The campaign banner that would have had a 96% hit rate on its own has a hit rate of roughly zero once it is fused to a per-donor member. The route that per-operation caching actually takes is the opposite of a composite POST. Send a query as a GET whose URL identifies the document - Automatic Persisted Queries does this with a hash standing in for the document text, with variables alongside - and the request becomes a normal cacheable URL that a shared cache can key, vary on identity, and serve. Batching closes that door by construction: you cannot express nine documents and nine variable sets as one cacheable URL. It is worth separating this from a client's **normalized cache**, which is unaffected. That cache lives in the client, keyed by object identity from the results after they arrive, and it does not care how the results travelled. Losing network-level caching while keeping client-level caching is exactly the sort of distinction an interviewer is listening for. ## The cache-key hazard, concretely Composite keys are also easy to get dangerously wrong, and this is the failure that turns an efficiency discussion into an incident review. Imagine a shared cache placed in front of that donations graph, configured to key batched POSTs on the request body so that dashboards would hit. Two supporters open their dashboards. Their batch bodies are byte-identical, because the per-donor member takes its identity from the session rather than from a variable - and the cache key does not include anything derived from the session. The second supporter is served the first supporter's row: their name and giving total, out of cache, with a 200. The cause is not exotic. Any cache key that omits the identity dimension of a personalized response leaks across viewers, and a composite body makes the mistake much easier to make, because the body looks like a complete description of the request when in fact the authorizing credential lives in a header. It is also why caching batched, personalized responses at a shared layer is usually the wrong idea outright rather than a keying problem to solve. ## Why tracing and metrics go coarse One HTTP request produces one access-log entry, one status, one duration and, by default, one server span. If nine operations rode inside it, then: - **Latency** is attributed to the batch. A per-operation p99 cannot be computed, and the 840 ms donor-profile member is invisible inside a request recorded as 861 ms. - **Errors** are attributed to nothing. A member failing validation rides inside a 200, so error-rate panels built on status stay flat. - **Operation-name dimensions collapse.** Dashboards and alerts sliced by operation name have no single name to use, and a server that records one gets whichever member it happened to pick. The fix is to instrument inside the batch executor rather than at the HTTP boundary: open a child span per member carrying its operation name and its index, and emit duration and outcome metrics per member. Log per member too, so an access log stops being the record of what happened. ```pseudocode handleBatch(members): batchSpan = startSpan("graphql.batch", { memberCount: length(members) }) results = awaitAll(members.map(m -> withSpan("graphql.operation", { operationName: m.operationName, index: indexOf(m), parent: batchSpan }, () -> execute(m)))) batchSpan.end() return results ``` That is a small amount of code, but it has to be written deliberately, and it is the standing cost of the convention: every per-operation signal that came for free on an unbatched endpoint must now be rebuilt inside the server. When a team weighs batching, this is the half of the ledger that tends to be forgotten, because the loss is silent - the dashboards keep rendering, they just stop meaning what they used to.
- Does batching also defeat a client's normalized cache?No, and the distinction matters. A normalized client cache stores objects by identity from results that have already arrived, so it is indifferent to how they travelled. What batching defeats is caching in the network - shared caches, a CDN, a server response cache - because those key on a request that has become a composite POST rather than an identifiable operation.
- Why does a body-aware cache key not rescue hit rates for a batch?Because cacheability composes the wrong way. The batch is cacheable only if every member is, and its key changes whenever any member's variables change. Fusing one per-user member to eight shared ones makes the shared eight uncacheable too, so a banner that would have hit on almost every request now misses on almost every request.
- What is the minimum instrumentation that makes a batched endpoint observable again?A child span per member carrying the operation name and index, and per-member duration and outcome metrics emitted from inside the executor. Per-member logging on top of that replaces the access log as the record of what happened. Without those, latency and error rate are only known per batch, and status-based alerting is blind to member failures.
It is a group order under one restaurant bill: the total is recorded faithfully, but nobody can tell afterwards which dish was slow, which one was wrong, or which one could have been reused.
saying these in an interview costs you the question
- Says a shared cache can key a batched POST as it would a GET
- Claims batching improves cache hit rates by combining requests
- Confuses a client's normalized cache with network caching
- Assumes tracing gives per-operation spans automatically
- Keys a cached personalized response on the request body alone
- Trusts status-based error rates on a batched endpoint