If POSTed GraphQL responses are never stored by HTTP caches, why can a field still return stale data?
answer
- The pressure moves, it does not disappear
- Below the transport nothing has a header
- Ask which tier each field read
- One response is not one snapshot
- No URL left to purge
basics
~20 sLosing the HTTP cache tier relocates caching rather than removing it. Client stores, server-side response caches, batch loaders and the data layer behind resolvers all still cache — below the boundary, where no header, metric or URL purge reaches them.
solid answer
~50 s"Uncacheable over HTTP" is a statement about one tier, not about the system. When the URL stops being a usable key, the caching does not go away — it moves inward, to places HTTP tooling cannot observe: the client's own store, a server-side response or parsed-document cache, per-request batch loaders, and above all the data layer each resolver reads from, which may be a shared key-value cache, a precomputed rollup, or a read replica running behind its primary. A stale value therefore has an origin **inside** the service. The diagnostic consequence is sharp: there is no `Age` header to read, no cache-hit counter on a reverse proxy, and no purge-by-URL. Two fields in one response can even carry data of different ages, because each was resolved independently against a different tier. You have to instrument staleness yourself, per field, or you cannot answer where it came from.
code
graphql · 8 linesquery SiteDashboard($site: ID!) {
site(id: $site) {
name
totalOutputKw
lastReadingAt
inverters { id status }
}
}go deeper
Remember that a value can be old even when nothing in the network stored it, because caches also live inside the server and inside the client library you are using.
Be able to list the tiers in order — client store, server response cache, per-request loader, shared data cache, replica or rollup — and say which of them a network trace can and cannot see.
Show the diagnosis: bisect tier by tier, explain why a purge changed nothing, and articulate that one response is not one snapshot because each field chose its own source.
Own the consequence for the organisation: caching, invalidation and cache observability stop being infrastructure your platform team runs and become product code your team writes, tests and staffs.
## The tier did not vanish; it moved The single-endpoint POST shape removes the HTTP cache tier. It removes nothing else. Caching is not a feature of HTTP that a system either has or lacks — it is a pressure that appears anywhere the same expensive answer is computed twice, and closing one outlet just routes it to another. After the move, a GraphQL response can carry stale bytes from at least five places: * **The client's own store.** A client library that keeps results locally will happily render a value it fetched earlier without asking again. This is invisible to the network entirely — no request is even made. * **A server-side response cache.** Some servers keep whole results keyed by operation plus variables plus a viewer identity, which is the same idea as an HTTP cache reimplemented above the transport, where the operation is finally readable. * **A per-request batch loader.** Within a single execution, a loader that has already fetched an object by id will return the same instance for a later field. It is short-lived by design, but inside one long-running operation it does mean two selections of the same entity are the same snapshot. * **A shared data cache.** The key-value cache your repository layer consults before touching the store of record. This is usually the biggest one and usually the one holding the stale value. * **The store of record itself.** A read replica lagging its primary, or a precomputed rollup refreshed on a schedule, serves data that was true a while ago. It is not called a cache, but it behaves as one for every purpose that matters here. ## A worked incident A solar-array telemetry graph has a 340 ms p99 budget on the site dashboard operation. To hold it, the team put the expensive aggregate behind a rollup refreshed every 60 seconds, while identity and status fields kept reading the primary directly: ```graphql query SiteDashboard($site: ID!) { site(id: $site) { name totalOutputKw # served from a 60s rollup lastReadingAt # read live from the primary inverters { id status } } } ``` Operators report that a site shows `lastReadingAt` seconds old while `totalOutputKw` is visibly wrong — by their measurement, up to 47 seconds behind. The first instinct is to blame a cache in the network path, so someone purges the CDN and nothing changes, which is unsurprising: the CDN never held anything, because the request is a POST to one URL. The actual cause is that **one response is not one snapshot**. Each field is resolved separately, and each resolver chose its own tier. HTTP caching, whatever its limits, gave you the opposite guarantee: a stored response was a single coherent capture of one moment with one `Age`. Field-level resolution against heterogeneous tiers gives up that coherence, and the resulting artefact — internally inconsistent data inside a single successful 200 — is one this shape produces and REST largely does not. ## What you lose operationally, and what to build back Three capabilities go away with the HTTP tier, and each has to be rebuilt inside the service if you want it: **Observability.** There is no `Age`, no `X-Cache: HIT`, no per-URL hit rate. The infrastructure dashboard your ops team already runs reports one hot endpoint at 100% origin traffic and tells you nothing else. Recovering it means emitting your own signals: a per-resolver cache-hit metric tagged with the field, and, when it matters to consumers, a provenance value returned in the response's extensions entry or as an explicit field such as `computedAt` on the aggregate. **Purging.** You cannot invalidate by URL, because the URL is the whole API. Invalidation has to happen at whatever internal key the data cache actually uses — an entity identity, a tenant, a tag — and that is application code someone must write and test. **A single freshness contract.** Under URL caching, one lifetime governed one response. Now every tier has its own lifetime and they compose in ways nobody wrote down. The practical discipline is to make freshness an explicit part of schema design: decide per field what age is acceptable, name the aggregate field in a way that admits it is an aggregate, and prefer returning an explicit timestamp over letting a client assume everything in the payload is simultaneous. ## How to answer this in an interview Do not stop at "caching moved inside". The sharp version is: the HTTP tier's real contribution was not only speed but a **coherent, observable, purgeable unit** — one response, one age, one key. Losing it costs those three properties independently, and a field returning stale data is the symptom of the first two failing at once. Then name where you would actually look, in order: the client store, the server response cache, the shared data cache, the replica or rollup — and say how you would tell them apart, which is by whether a request left the client at all, and whether the value changes when you bypass each tier in turn.
- How would you actually localise which tier served the stale value?Bisect by tier from the outside in. First check whether a request left the client at all — if not, it is the client store. If it did, re-run the same operation with the server's response cache bypassed; if the value corrects, you have it. Otherwise instrument the resolver: log the data-layer key and whether it hit the shared cache, and compare a direct read of the primary against what the resolver returned. Each step rules out exactly one tier.
- Is field-level staleness inside a single response ever acceptable, or should you force one snapshot?It is often acceptable and sometimes the only way to hold a latency budget — an aggregate refreshed on a schedule next to live identity fields is a reasonable trade. What is not acceptable is leaving it implicit. Make the age visible: name the field so it reads as an aggregate, return the time it was computed, and document the acceptable age. Forcing one coherent snapshot means resolving everything against one consistent read, which costs you the cheap tier you introduced in the first place.
- Your infrastructure dashboard shows the endpoint at 100% origin traffic. What does that number actually tell you?Only that no HTTP cache is storing anything, which was already known from the transport shape. It says nothing about the internal hit rates that determine real cost, so it is a metric that looks alarming and carries no information. The replacement is application-level: hit rate per internal cache, tagged by field or by data-layer key, plus per-operation latency, because the endpoint's aggregate p99 mixes a trivial lookup with an expensive report.
Damming a river does not stop the water; it finds new channels. Closing the HTTP cache does not stop caching — it just moves it somewhere with no gauge on it.
saying these in an interview costs you the question
- Assumes an uncacheable transport means nothing cached
- Purges a CDN that never stored the response
- Expects an Age header from an internal cache
- Treats one response as one coherent snapshot
- Tries to invalidate internal caches by URL
- Reads endpoint-level p99 as one workload's latency