skip to content

Normalized Client Caches

A client shreds each response into entities keyed by type and id, so one fetched object updates every view showing it. Interviewers ask what happens when a read wants a field the store never kept.

part ofGraphQLoverview, primer and where to startread it →
on this pageshow

questions

3

What does a normalized GraphQL client cache store, and how does it differ from a document cache?

level: juniorimportance: must knowfreq 68%

answer

  1. Two caches, two different keys
  2. One object shown on many screens
  3. Responses are shredded, not stored whole
  4. Type name plus identifier as key
  5. References replace nested objects

basics

~20 s

A normalized client cache shreds each response into individual objects stored under an identity key, usually the object's __typename plus its id, with nested objects replaced by references. A document cache instead stores a whole response under the operation plus its variables.

solid answer

~60 s

A document cache is a lookup table: the key is the operation the client sent plus the variables it sent, the value is the response body. Repeat that exact request and you get the stored body back; change one variable and it is a miss. A normalized cache does more work on the way in. It walks the response, and for every object it can identify it computes a storage key — conventionally the object's `__typename` combined with its `id` — and writes that object's fields into a flat entity map, replacing the nested object in its parent with a reference to that key. A later read is executed against the store rather than looked up: the client walks the requested selection set from a root entry, following references. The payoff is consistency — one fetch of an object updates every screen showing it. None of this is in the GraphQL specification; it is a client-side convention, and the server does nothing to enable it beyond returning `__typename` and a stable identifier when asked.

code

json · 11 lines
json
{
  "data": {
    "pallet": {
      "__typename": "Pallet",
      "id": "PLT-8842",
      "binCode": "A-14-3",
      "quantity": 236,
      "sku": { "__typename": "Sku", "id": "SKU-7719", "name": "Anchor bolt M12" }
    }
  }
}

go deeper

for a junior

Be ready to say plainly what the two caches key on: operation plus variables for a document cache, type name plus id for a normalized one. Knowing that nested objects become references is enough at this level.

for a middle

Explain the write path and the read path separately — shredding on write, executing the selection set against the store on read — and why that lets an operation never sent before be answered from cache.

for a senior

Expect to justify the cost. Talk about the memory and CPU of normalizing every response, the debugging burden of a store update changing an unrelated screen, and when a plain operation cache is the better call.

for a principal

Own the contract this convention imposes on the schema: stable, never-recycled identifiers on anything fetched from two places, and a shared position on which types get identity across an 11-service graph so every client keys them the same way.

## Two very different meanings of "the client cached it" When someone says a GraphQL client cached a result, they mean one of two mechanisms, and interviewers ask this question to find out whether you can tell them apart. **A document cache** (also called a response or operation cache) is the simpler one and is the same shape as an HTTP cache. The key is the identity of the request: the operation document (or an identifier standing in for it) plus the exact variables. The value is the response body, stored whole. Sending the same operation with the same variables is a hit; anything else is a miss. It is cheap, it is easy to reason about, and it has one structural weakness — the same object appearing in five different responses is stored five times, and refreshing one of those responses leaves the other four stale. **A normalized cache** attacks exactly that weakness. Instead of storing responses, it stores *objects*. ## The shredding step On every write — that is, every time a response arrives — the client walks the response tree. For each object node it asks: can I identify this? The near-universal convention is to combine the object's `__typename` with its `id` field, producing a key such as `Pallet:PLT-8842`. `__typename` is a specified meta-field available on every object, interface and union type, which is why clients lean on it; the `id` half is pure convention. The object's scalar fields are written into a flat map under that key, merging with whatever was already there. Wherever that object sat inside its parent, the client leaves a *reference* — a pointer to the key rather than the data. The result is not a tree of responses. It is a graph: a flat map from entity key to a bag of fields, plus a root entry holding the pointers that the top-level fields of each operation returned. That mirrors the shape of the schema itself, which is why the technique fits GraphQL so naturally. Consider a warehouse inventory graph. A dashboard lists 34 pallets in a bin, a detail screen shows one pallet, and an audit screen shows the same pallet inside a stock-count record. In a document cache those are three independent stored bodies holding three copies of pallet `PLT-8842`. In a normalized store there is exactly one `Pallet:PLT-8842` entry and three references to it. ## Reads are executed, not looked up Because the store holds objects rather than answers, a read cannot be a hash lookup. The client takes the selection set of the operation being read and *executes* it against the store: start at the root entry, resolve each selected field, and when a value is a reference, jump to that entity and keep resolving. If every selected field is found, the client assembles a result shaped exactly like a server response and returns it without a network call. Two consequences fall straight out of this. First, an operation that was never sent can still be served entirely from cache, provided some earlier response happened to store every object and field it selects. Second, and this is the flip side, the read fails as soon as one selected field is absent from the store — which is why partial misses are the classic follow-up to this question. ## Consistency is the whole point The reason teams pay the normalization cost is write propagation. When any response — a query result, a mutation payload, a subscription message — writes new fields for `Pallet:PLT-8842`, every view whose read touched that entity is affected at once. Client libraries track which active reads depend on which entity keys, so an update to one object refreshes the list, the detail screen and the audit screen together, with no explicit invalidation logic per screen. ## What is specified and what is not Nothing here is in the GraphQL specification. The specification defines the type system, the executable document and the execution algorithm; it says nothing about client storage. `__typename` is specified. The convention that a field named `id` carries a stable, globally unique identifier is a server-side convention that clients rely on, and clients let you override the key derivation per type precisely because it is a convention rather than a rule. That matters in practice. A server team does not "turn on" normalization; it enables it by giving objects stable identifiers and by not reusing an identifier for two different things. If a type has no identifier, the client cannot key it, and it falls back to storing that object inline under its parent — losing the sharing and reintroducing the duplicate copies a document cache would have had. ## When the simpler cache is the right one Normalization is not free: it costs CPU on every write, memory for the entity map, and a debugging burden when a store update changes a screen nobody was looking at. A client that mostly issues read-once, non-overlapping operations — reporting exports, one-shot search — gets little from it, and a plain operation-plus-variables cache is easier to reason about. The judgement call is whether the same objects genuinely appear across many concurrent views.

  • Which part of the storage key is defined by the GraphQL specification and which is convention?
    `__typename` is specified: it is a meta-field available on every object, interface and union type, and always resolves to the concrete object type's name. Combining it with a field named `id` to form the key is purely a client convention, and so is the expectation that `id` is stable and unique across the whole schema. Clients expose per-type key configuration precisely because a schema may identify objects some other way, or not at all.
  • What does a normalized cache do with an object that carries no identifier?
    It cannot give the object its own entry, so it stores the object inline under the field of its parent entity. That is safe but loses the benefit: the same logical object fetched through two different parents becomes two independent copies, and updating one leaves the other stale. It is one reason schema authors are pushed to give identity to any type a client will fetch from more than one place.
  • Does a normalized client cache remove the need for HTTP or CDN caching?
    No — they sit at different layers and solve different problems. The client store only helps a single client process that has already fetched the data; it is per-user, in-memory and empty on a cold start. Shared caching in front of the server is what spares the origin repeated work across users and sessions. A team typically wants both, and the client store never sees a request the network layer answered.

A document cache photocopies each report and files the copies; a normalized cache files each fact once in a card index and rebuilds any report from the cards.

saying these in an interview costs you the question

  • Says the client keys cached data by URL like an HTTP cache
  • Claims normalization is required by the GraphQL specification
  • Thinks the server builds or controls the client's store
  • Assumes every object can be normalized without an identifier
  • Confuses the client's store with a server-side response cache
  • Believes a normalized read is a single hash lookup

context

open as a page

Why does a normalized client cache miss when a read selects a field it never stored?

level: middleimportance: should knowfreq 51%

basics

~20 s

A read runs the whole selection set against the entity store and must find every selected field. One field absent from an entity makes the read incomplete, so the client reports a miss and goes to the network rather than returning a partly filled object.

open as a page

In a normalized client cache, what happens when two operations write the same entity with different values?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

The writes merge field by field into one shared entry, and for any field both carry, the write that lands last wins. Nothing compares freshness, so a slow response can overwrite a newer value and visibly change screens that never issued it.

open as a page