skip to content

How should a normalized cache merge an incoming page into an already-cached list field?

level: middleimportance: should knowfreq 46%

answer

  1. Last write wins is not what you want
  2. A function of existing plus incoming
  3. The request's arguments say where it goes
  4. The same window can arrive twice
  5. Fresh metadata, not the stored copy

basics

~20 s

A merge rule receives what is already stored under the field's key plus the incoming page, and returns the value to store. It must place the incoming items by the request's arguments rather than blindly concatenating them.

solid answer

~50 s

Once the window arguments are out of the storage key, every page writes to the same entry, and the cache's default of last-write-wins would replace the list. A merge rule is a function of the existing value, the incoming value and the arguments of this request, returning what gets stored. The naive body — `existing.edges + incoming.edges` — is wrong in several ordinary situations: the same window arriving twice, from a retry or a remount, duplicates rows; a backward page belongs at the front, not the end; a fetch carrying no cursor is a reset and must replace rather than extend; and the page metadata must be taken from the incoming page, because the stored end cursor describes a stale window. A robust rule writes items at the position the arguments imply and de-duplicates by item identity. This is client-cache convention — the specification has no merge concept.

code

pseudocode · 8 lines
pseudocode
merge(existing, incoming, args):
    if args.after is null and args.before is null:
        return incoming
    base  = existing or { edges: [] }
    known = setOf(e.node.id for e in base.edges)
    fresh = [e for e in incoming.edges if e.node.id not in known]
    edges = args.before ? fresh + base.edges : base.edges + fresh
    return { edges: edges, pageInfo: incoming.pageInfo }

go deeper

for a junior

Recall that something has to decide what happens when two pages hit one cache entry, and that the default is replacement. Knowing the rule exists and what it is given is enough at this level.

for a middle

Explain the three inputs — existing, incoming, this request's arguments — and name at least two ways plain concatenation breaks. Be able to write the reset branch for a cursor-less request.

for a senior

Diagnose from symptoms: duplicated rows point at missing de-duplication, a load more that never advances points at stale page metadata, an interleaved list points at out-of-order writes. Also own the memory cost of one unbounded entry.

for a principal

Argue for one merge implementation shared across every paginated field rather than per-screen rules, and decide where windowed reads and retention caps belong so no team re-derives them.

## Why a rule is needed at all Once the window arguments — the cursor, the offset, the page size — are excluded from a field's storage key, every page of the list writes to the **same** entry. That is the point of excluding them, but it creates a new problem immediately: the cache now has an existing value and an incoming value for one key and no idea what the result should be. Its default is last-write-wins, which turns "load more" into "replace the list with the newest twenty-five rows". A merge rule is what you supply so the cache does something better. ## The signature A merge rule is a function of three things and returns one: - the **existing** value stored under the key — absent on the very first write, which the rule must handle; - the **incoming** value, in the raw shape the field has in the response: a connection object carrying edges plus page metadata, or a plain list; - the **arguments and variables of this request**, which is the only thing that tells the rule *where* the incoming window belongs. That third input is what separates a correct rule from a naive one. `existing.edges + incoming.edges` ignores it entirely. ## The four ways blind append breaks **A page can arrive twice.** A retry, a fetch policy that serves the cache and then the network, a component remounting and re-issuing its query — any of these can deliver a window you already merged. Concatenation duplicates every row in it. In a benefits enrolment list, a double-click on "load more" fired the same `after` cursor twice and produced twenty-five enrolments listed twice in a row, with duplicate render keys behind them. De-duplicate by item identity, or write at a computed position so a repeat overwrites rather than extends. **Direction.** Backward pagination hands you the window *before* what you hold. It belongs at the front of the accumulated list, and appending it puts older rows below newer ones. **A cursor-less request is a reset, not a page.** After a filter change, a pull-to-refresh, or the first mount following an eviction, the client asks for the head of the list with no cursor. A rule that appends turns that into a list containing the first page twice. Treat the absence of a cursor argument as "replace". **Page metadata goes stale.** The list travels with information about where the window ended and whether more exists. If the merge keeps the stored copy instead of the incoming one, the end cursor never moves — so the next "load more" sends the cursor it already used, merges the same window again, and the list stops advancing. The symptom in the UI is a button that spins and does nothing, and it is one of the most common merge bugs there is. Take the metadata from the incoming page. ## Writing at a position rather than concatenating The more robust shape of the rule is not "append" but "write this window at the position its arguments imply". With offset arguments that position is explicit: the offset is the index. With cursor arguments it is implicit, but the request's cursor identifies the tail the window follows, so the rule can check that it still matches the tail of the accumulated list, and reset or drop the write when it does not. That check is what stops two in-flight "load more" requests from interleaving when they come back out of order. ```pseudocode merge(existing, incoming, args): if args.after is null and args.before is null: return incoming # a fresh head fetch is a reset base = existing or emptyList() known = idsOf(base) fresh = [e for e in incoming.edges if e.node.id not in known] items = args.before ? fresh + base : base + fresh return { edges: items, pageInfo: incoming.pageInfo } ``` ## The cost of one accumulating entry The merged entry is, by construction, everything fetched so far. Scrolling seventy-four pages of twenty-five statements leaves eighteen hundred references live for the session, and every write to that entry can wake every view reading it. Three usual mitigations: pair the merge with a read rule that returns only the window a caller asked for, cap how many pages the entry retains, and evict or reset the entry when the view unmounts or the filter changes. ## Specified or conventional? Entirely conventional. The GraphQL specification defines how a server executes a document and shapes a response; it has no concept of a client store, a storage key, or a merge function. Two client caches can differ on every detail above — whether a merge receives the field arguments, whether page metadata is merged separately, whether reads are windowed. Answer in terms of the mechanism, not in terms of one library's option names.

  • A list refetches the same window forever and never advances. What is the likely merge bug?
    The rule kept the stored page metadata instead of the incoming page's. The end cursor never moves, so every load more sends the cursor it already used, merges a window it already holds, and the list stops growing. It looks like a hung button. Take the end cursor and the has-next flag from the incoming value.
  • How should a merge rule treat a request that carries no cursor argument at all?
    As a reset, not as another page — return the incoming value and discard what was stored. That request shape occurs on a filter change, a pull-to-refresh, or the first mount after eviction, and appending it would put the head of the list into the entry twice.
  • The merged entry grows to thousands of items in one session. What do you do about it?
    That growth is inherent to a single accumulating entry, so bound it deliberately: pair the merge with a read rule that returns only the window a caller asked for, cap how many pages the entry retains, and evict or reset it when the view unmounts, the route changes or the filter changes.

saying these in an interview costs you the question

  • Just concatenates existing and incoming items
  • Says duplicates cannot happen, the server never resends
  • Keeps the stored page metadata after merging
  • Treats a cursor-less refetch as another page to append
  • Always appends, ignoring backward pagination
  • Thinks the GraphQL spec defines a merge function

context