skip to content

In Apollo Client 4, how does InMemoryCache store an author shown on fifty feed posts, and what breaks if the query omits the author's id?

level: juniorimportance: must knowfreq 58%

answer

  1. one record, many pointers
  2. type name plus identifier
  3. references instead of copies
  4. no id means an embedded copy

basics

~20 s

InMemoryCache stores each object with __typename and id once, under a cache ID such as Author:42, and every post references it. Without id in the feed query, each post embeds its own author copy and renames stop propagating.

solid answer

~50 s

When a result is written, `InMemoryCache` computes a cache ID for every object: by default `__typename` plus `id` (falling back to `_id`), so the author becomes `Author:42`. That record is stored once in a flat map, and each `Post` keeps a reference `{ __ref: "Author:42" }` in its `author` field, so fifty posts read one record and any later write to `Author:42` shows on all of them. Apollo Client 4 adds `__typename` to nested selections automatically, so the usual gap is a missing `id`. Without it the author cannot be identified and is stored inline inside each post, as that post's private copy. A rename written to `Author:42` from another query then never reaches the feed, which keeps the old name until it is refetched. `cache.identify(object)` tells you which ID, if any, the cache would use.

code

graphql · 11 lines
graphql
query Feed {
  feed(first: 20) {
    id
    body
    author {
      id
      name
      avatarUrl
    }
  }
}

go deeper

for a junior

Recall the default cache ID, __typename plus id with _id as a fallback, and that entities are stored once and referenced everywhere else.

for a middle

Explain what happens to an object without an id: it is stored inline in its parent, so writes to the entity elsewhere cannot reach it. Show it with cache.extract or cache.identify.

for a senior

Treat a missing id in a shared fragment or query as a correctness bug, not a size issue, and know how to spot embedded copies when stale data appears in one view but not another.

for a principal

Frame id selection as a team convention: every entity type exposes an id or declared key, and review or codegen checks catch selections that drop it before stale-data bugs reach production.

## What InMemoryCache does with a result `InMemoryCache` is the cache that ships with Apollo Client. It does not keep query responses as opaque blobs. When a response arrives, the cache walks the result tree and splits it into **entities**: objects it can identify. Each entity is stored once, in a flat map keyed by a **cache ID**, and wherever the entity appeared in the tree the cache stores a **reference** to it instead: a small object of the form `{ __ref: "Author:42" }`. This process is called **normalization**. Its payoff is that one real-world object has one cached record, no matter how many queries or list positions mention it. ## How the default cache ID is computed Unless you configure otherwise, the cache uses a built-in function, `defaultDataIdFromObject`, which reads two fields: - **`__typename`**: the GraphQL type name. Apollo Client 4 adds `__typename` to every nested selection set of the outgoing document automatically, and since 4.0 there is no `addTypename` option to turn that off. - **`id`**, or **`_id`** when `id` is absent. A number or string is used as is; any other value is JSON-encoded. | Object in the response | Default cache ID | Stored as | |---|---|---| | `{ __typename: "Author", id: 42 }` | `Author:42` | its own entity | | `{ __typename: "Author", _id: "a9" }` | `Author:a9` | its own entity | | `{ __typename: "Author", name: "Ada" }` | none | inline, inside its parent | | `{ __typename: "Post", id: 42 }` | `Post:42` | its own entity, distinct from `Author:42` | The type-name prefix is why a post and an author that share the number 42 never collide. ## Walking through the feed Suppose a feed query selects fifty posts, and every post selects `author { id name avatarUrl }`. Many of those posts were written by the same author, id 42. 1. The response contains fifty nested author objects, many of them identical copies. 2. The cache computes `Author:42` for each copy and merges their fields into one record. 3. Each `Post:N` entity stores `author: { __ref: "Author:42" }`. 4. The root query's `feed` field stores a list of references to the post entities. 5. When the feed is read back, the cache follows the references and rebuilds the tree. Now the author renames themselves on a profile screen, and the result that comes back contains `{ __typename: "Author", id: 42, name: "Ada L." }`. That write lands on the single `Author:42` record, every post in the feed resolves its author through that record, and the feed shows the new name without another request. ## What goes wrong when the id is missing If the feed query selects `author { name avatarUrl }` without `id`, the cache cannot compute an ID for the author. It does not guess and it does not fail; it stores the author **inline**, as a plain nested object inside each post's `author` field. - Fifty posts now hold fifty private author copies. - A write to `Author:42` from the profile screen updates that entity, but the feed's embedded copies are not linked to it, so the feed keeps the old name. - The stale copies change only when the feed query itself is refetched and overwrites them. - Anything that needs the entity by ID, such as `cache.identify`, `cache.modify` on `Author:42` or `cache.evict`, cannot reach the embedded copies. The fix is almost always to select `id` wherever the type appears. When the type has no `id` field at all, you declare its key in a type policy instead, which is a separate configuration step. ## Checking what the cache holds Two calls make normalization visible while debugging: - `cache.identify(object)` returns the cache ID the cache would use for an object, or `undefined` when it cannot compute one. - `cache.extract()` returns the whole normalized map, where you can see `ROOT_QUERY`, the `Post:N` entities and whether each post's `author` is a `__ref` or an embedded object. ```ts const id = cache.identify({ __typename: "Author", id: 42 }); // "Author:42" const missing = cache.identify({ __typename: "Author", name: "Ada" }); // undefined: this object would be stored inline ``` ## What changed in Apollo Client 4 Normalization itself works as it did in Apollo Client 3. The visible change is that `__typename` injection is no longer optional: the `addTypename` option was removed from `InMemoryCache` in 4.0, so every document sent through the cache asks for `__typename` in its nested selection sets. Test fixtures that previously switched typename injection off now need `__typename` in their mocked responses for normalization to behave as it does in production.

  • Can you stop Apollo Client 4 from adding __typename to your queries, and would you want to?
    No. Apollo Client 4 removed the `addTypename` option from `InMemoryCache`, so `__typename` is always added to nested selection sets. You would not want to disable it anyway: the default cache ID, type policies and matching fragments on interfaces all depend on knowing each object's type. Mocked responses in tests must include `__typename` for normalization to match production.
  • The backend's Author type exposes `_id` rather than `id`. Does InMemoryCache need any configuration?
    Not for that case. The default ID function falls back to `_id` when `id` is absent, so the author becomes `Author:<_id value>` automatically. Configuration is needed only when the identifying field has another name, such as `handle`, or when several fields together identify the object; then you declare `keyFields` in a type policy for that type.
  • What does `cache.identify` return for an object that has __typename but no identifying field?
    It returns `undefined`, because no cache ID can be computed. That is a quick diagnostic: if `cache.identify` on an object from your result gives `undefined`, the cache is storing that object inline in its parent, so writes to the entity from elsewhere will not reach it.

A normalized cache works like a shared address book. Each post's author line says see entry Author:42, so updating that one entry changes every page that points to it. A post that copied the author's details onto its own page, because it never recorded the entry number, keeps the old details.

saying these in an interview costs you the question

  • Apollo stores each query's response whole, so every query keeps its own author copy.
  • The cache ID is just the id value, so Post 42 and Author 42 collide.
  • In Apollo Client 4 you must set addTypename to true for normalization to work.
  • Leaving id out only costs memory; a rename still reaches every post.
  • Apollo adds id to every selection automatically, just as it adds __typename.