skip to content

Normalized Cache

InMemoryCache normalizes results by __typename and id, so an object is stored once and shared by every query that references it. Knowing readQuery, writeQuery, modify, keyFields, and the fetch policies is what separates people who have debugged an Apollo cache from people who have not.

on this pageshow

explore

questions

5

In Apollo Client 4, how does InMemoryCache store an author shown on fifty feed posts, and what breaks if the query omits the author's id?

level: juniorimportance: must knowfreq 58%

answer

  1. one record, many pointers
  2. type name plus identifier
  3. references instead of copies
  4. no id means an embedded copy

basics

~20 s

InMemoryCache stores each object with __typename and id once, under a cache ID such as Author:42, and every post references it. Without id in the feed query, each post embeds its own author copy and renames stop propagating.

solid answer

~50 s

When a result is written, `InMemoryCache` computes a cache ID for every object: by default `__typename` plus `id` (falling back to `_id`), so the author becomes `Author:42`. That record is stored once in a flat map, and each `Post` keeps a reference `{ __ref: "Author:42" }` in its `author` field, so fifty posts read one record and any later write to `Author:42` shows on all of them. Apollo Client 4 adds `__typename` to nested selections automatically, so the usual gap is a missing `id`. Without it the author cannot be identified and is stored inline inside each post, as that post's private copy. A rename written to `Author:42` from another query then never reaches the feed, which keeps the old name until it is refetched. `cache.identify(object)` tells you which ID, if any, the cache would use.

code

graphql · 11 lines
graphql
query Feed {
  feed(first: 20) {
    id
    body
    author {
      id
      name
      avatarUrl
    }
  }
}

go deeper

for a junior

Recall the default cache ID, __typename plus id with _id as a fallback, and that entities are stored once and referenced everywhere else.

for a middle

Explain what happens to an object without an id: it is stored inline in its parent, so writes to the entity elsewhere cannot reach it. Show it with cache.extract or cache.identify.

for a senior

Treat a missing id in a shared fragment or query as a correctness bug, not a size issue, and know how to spot embedded copies when stale data appears in one view but not another.

for a principal

Frame id selection as a team convention: every entity type exposes an id or declared key, and review or codegen checks catch selections that drop it before stale-data bugs reach production.

## What InMemoryCache does with a result `InMemoryCache` is the cache that ships with Apollo Client. It does not keep query responses as opaque blobs. When a response arrives, the cache walks the result tree and splits it into **entities**: objects it can identify. Each entity is stored once, in a flat map keyed by a **cache ID**, and wherever the entity appeared in the tree the cache stores a **reference** to it instead: a small object of the form `{ __ref: "Author:42" }`. This process is called **normalization**. Its payoff is that one real-world object has one cached record, no matter how many queries or list positions mention it. ## How the default cache ID is computed Unless you configure otherwise, the cache uses a built-in function, `defaultDataIdFromObject`, which reads two fields: - **`__typename`**: the GraphQL type name. Apollo Client 4 adds `__typename` to every nested selection set of the outgoing document automatically, and since 4.0 there is no `addTypename` option to turn that off. - **`id`**, or **`_id`** when `id` is absent. A number or string is used as is; any other value is JSON-encoded. | Object in the response | Default cache ID | Stored as | |---|---|---| | `{ __typename: "Author", id: 42 }` | `Author:42` | its own entity | | `{ __typename: "Author", _id: "a9" }` | `Author:a9` | its own entity | | `{ __typename: "Author", name: "Ada" }` | none | inline, inside its parent | | `{ __typename: "Post", id: 42 }` | `Post:42` | its own entity, distinct from `Author:42` | The type-name prefix is why a post and an author that share the number 42 never collide. ## Walking through the feed Suppose a feed query selects fifty posts, and every post selects `author { id name avatarUrl }`. Many of those posts were written by the same author, id 42. 1. The response contains fifty nested author objects, many of them identical copies. 2. The cache computes `Author:42` for each copy and merges their fields into one record. 3. Each `Post:N` entity stores `author: { __ref: "Author:42" }`. 4. The root query's `feed` field stores a list of references to the post entities. 5. When the feed is read back, the cache follows the references and rebuilds the tree. Now the author renames themselves on a profile screen, and the result that comes back contains `{ __typename: "Author", id: 42, name: "Ada L." }`. That write lands on the single `Author:42` record, every post in the feed resolves its author through that record, and the feed shows the new name without another request. ## What goes wrong when the id is missing If the feed query selects `author { name avatarUrl }` without `id`, the cache cannot compute an ID for the author. It does not guess and it does not fail; it stores the author **inline**, as a plain nested object inside each post's `author` field. - Fifty posts now hold fifty private author copies. - A write to `Author:42` from the profile screen updates that entity, but the feed's embedded copies are not linked to it, so the feed keeps the old name. - The stale copies change only when the feed query itself is refetched and overwrites them. - Anything that needs the entity by ID, such as `cache.identify`, `cache.modify` on `Author:42` or `cache.evict`, cannot reach the embedded copies. The fix is almost always to select `id` wherever the type appears. When the type has no `id` field at all, you declare its key in a type policy instead, which is a separate configuration step. ## Checking what the cache holds Two calls make normalization visible while debugging: - `cache.identify(object)` returns the cache ID the cache would use for an object, or `undefined` when it cannot compute one. - `cache.extract()` returns the whole normalized map, where you can see `ROOT_QUERY`, the `Post:N` entities and whether each post's `author` is a `__ref` or an embedded object. ```ts const id = cache.identify({ __typename: "Author", id: 42 }); // "Author:42" const missing = cache.identify({ __typename: "Author", name: "Ada" }); // undefined: this object would be stored inline ``` ## What changed in Apollo Client 4 Normalization itself works as it did in Apollo Client 3. The visible change is that `__typename` injection is no longer optional: the `addTypename` option was removed from `InMemoryCache` in 4.0, so every document sent through the cache asks for `__typename` in its nested selection sets. Test fixtures that previously switched typename injection off now need `__typename` in their mocked responses for normalization to behave as it does in production.

  • Can you stop Apollo Client 4 from adding __typename to your queries, and would you want to?
    No. Apollo Client 4 removed the `addTypename` option from `InMemoryCache`, so `__typename` is always added to nested selection sets. You would not want to disable it anyway: the default cache ID, type policies and matching fragments on interfaces all depend on knowing each object's type. Mocked responses in tests must include `__typename` for normalization to match production.
  • The backend's Author type exposes `_id` rather than `id`. Does InMemoryCache need any configuration?
    Not for that case. The default ID function falls back to `_id` when `id` is absent, so the author becomes `Author:<_id value>` automatically. Configuration is needed only when the identifying field has another name, such as `handle`, or when several fields together identify the object; then you declare `keyFields` in a type policy for that type.
  • What does `cache.identify` return for an object that has __typename but no identifying field?
    It returns `undefined`, because no cache ID can be computed. That is a quick diagnostic: if `cache.identify` on an object from your result gives `undefined`, the cache is storing that object inline in its parent, so writes to the entity from elsewhere will not reach it.

A normalized cache works like a shared address book. Each post's author line says see entry Author:42, so updating that one entry changes every page that points to it. A post that copied the author's details onto its own page, because it never recorded the entry number, keeps the old details.

saying these in an interview costs you the question

  • Apollo stores each query's response whole, so every query keeps its own author copy.
  • The cache ID is just the id value, so Post 42 and Author 42 collide.
  • In Apollo Client 4 you must set addTypename to true for normalization to work.
  • Leaving id out only costs memory; a rename still reaches every post.
  • Apollo adds id to every selection automatically, just as it adds __typename.
open as a page

In Apollo Client 4, after a user deletes a feed post, when do you remove it with cache.evict and cache.gc, cache.modify, or readQuery and writeQuery?

level: middleimportance: must knowfreq 50%

basics

~10 s

cache.evict removes the Post entity everywhere, and cache.gc then drops what became unreachable. cache.modify edits specific fields in place, bypassing merge functions. readQuery and writeQuery rewrite the result of one exact query and variables.

open as a page

In Apollo Client 4, the Author type has no id and is identified by a unique handle; how do you make InMemoryCache normalize it?

level: middleimportance: should knowfreq 42%

basics

~10 s

Declare the key in a type policy: typePolicies: { Author: { keyFields: ["handle"] } }. Authors are then stored under IDs like Author:{"handle":"ada"}, and every query that selects an Author must also select handle.

open as a page

In Apollo Client 4, the console warns 'Cache data may be lost when replacing the stats field of a Post object'; why does it happen, and what is the safe fix?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Post.stats is an object with no cache ID, so it is stored inline, and each write replaces it, dropping subfields another query selected. Give it an identity with keyFields, or declare merge: true when stats always belongs to its post.

open as a page

In Apollo Client 4, fetchMore with the next cursor returns page two of a feed, yet the screen still shows only page one; what is wrong in the cache configuration?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Without a field policy, Query.feed is stored once per argument set, so page two lands under its own after-cursor key and the original query never sees it. Set keyArgs to exclude the cursor and add a merge function, or use relayStylePagination.

open as a page