Why does a normalized client cache miss when a read selects a field it never stored?
answer
- A read is executed, not looked up
- Completeness is all or nothing
- Absent differs from a stored null
- Same objects, different fields selected
- A rename orphans every stored entry
basics
~20 sA read runs the whole selection set against the entity store and must find every selected field. One field absent from an entity makes the read incomplete, so the client reports a miss and goes to the network rather than returning a partly filled object.
solid answer
~50 sReading from a normalized store is an execution, not a lookup: the client walks the selection set from a root entry, following references between entities, and collects a value for every field the document selects. Completeness is all-or-nothing by default — if any entity along the way is missing one selected field, there is no correct value to put in the response shape, so the read is declared a miss and the operation goes to the network. Crucially, *absent* and *null* are different states. A field stored with the value `null` is a hit that yields `null`; a field the store has never seen is a hole. That is why a list operation selecting four lean fields does not warm a detail operation selecting nine: the entities exist and are correctly keyed, but the extra five fields were never written. Overlap between operations has to be real field-for-field, not merely the same objects.
code
graphql · 19 linesquery BinDashboard {
binPallets(binCode: "A-14-3") {
__typename
id
binCode
quantity
}
}
query PalletDetail {
pallet(id: "PLT-8842") {
__typename
id
binCode
quantity
lastCountedAt
damageFlag
}
}go deeper
Remember the shape of the answer: the client must find every field the operation selects, and one missing field sends the whole operation to the network even though the object is already stored.
Walk the read algorithm out loud — root entry, field lookup, follow references, recurse — and be precise that a stored null is a hit while a never-written field is a hole.
Show you can reduce misses deliberately: shared named fragments, a fatter list payload where drill-in dominates, and a version-stamped persisted store dropped on schema change rather than left full of unusable entries.
Frame it as a client-server contract. Decide who owns the canonical field set per type across teams, and treat a field rename as a client-breaking change with a coordinated release plan, not a same-day server edit.
## The read path is where normalization surprises people A document cache answers one question: have I seen this exact request before? A normalized store answers a harder one: can I *reconstruct* the answer to this request from objects I already hold? Reconstruction is a small execution engine running against local data, and it either completes or it does not. The algorithm is simple to state. Start at the root entry for the operation type. For each field in the selection set, look it up on the current entity. If the field's value is a reference, jump to that entity and recurse with the nested selection set. If it is a list of references, do that for each element. Assemble everything into the exact response shape the document describes. Now the failure case. Suppose the current entity has no key for a selected field. There is no value to place in the result, and the client will not silently omit the field, because the caller asked for a result matching the document. So the default policy is to abandon the read and fetch from the server. That is the *partial cache miss*, and it is the single most common source of "why is my client still hitting the network when the data is obviously right there?" ## Absent is not null The distinction that separates a middle answer from a junior one: a stored `null` is data. If a response returned `lastCountedAt: null` and the store wrote that, a later read of `lastCountedAt` is a hit yielding `null`. A field that never appeared in any response for that entity is not data — the store cannot tell whether it is null, a value, or an error, so it must ask. Candidates who conflate the two will predict hits where the client will actually refetch. ## A worked example in a warehouse inventory graph Two screens read the same pallets. The bin dashboard runs a list operation selecting `__typename`, `id`, `binCode` and `quantity` for the 34 pallets in bin `A-14-3`. After it resolves, the store holds 34 correctly keyed `Pallet` entries. The pallet detail screen then reads pallet `PLT-8842`, selecting those four fields plus `lastCountedAt`, `damageFlag` and a nested `Sku` with two fields of its own. The entity is present. Its key is right. Four of its selected fields are stored. The read still misses, on `lastCountedAt`, and the client issues a network request for the whole detail operation. The engineering response is not to "fix the cache" — the cache is behaving correctly. It is to make the overlap real: define the fields a `Pallet` needs on a named fragment and select that fragment from both operations, so the list write stores exactly what the detail read will ask for. That trades a fatter list response for a warm detail screen, and whether it is worth it depends on how often users drill in. ## Why clients default to all-or-nothing You might ask why the client does not return the four fields it has and fetch only the missing three. Some client libraries can be configured to return an incomplete result and let the caller decide, and a few can issue a narrowed follow-up operation. But the conservative default exists for good reasons: a caller that asked for a shape and received a different one has to handle undefined fields everywhere, and a per-field diff against the schema is real complexity for a benefit that a fragment discipline delivers more cheaply. Treat the strict default as correct behaviour to design around, not a defect to work around. ## The failure mode this produces in production The nastiest version of this is a schema change. Suppose a field is renamed — `quantity` becomes `unitCount` — and the change ships. Every entity already in every client's store carries the old field key, and nothing carries the new one. From the store's point of view, no entity has ever seen `unitCount`, so every read selecting it misses, and a client that had been serving screens locally starts fetching all of them. On the web that is a bad afternoon. For a released mobile client whose normalized store is persisted to disk across launches, it is worse: users restore a store full of entries that can never satisfy a read again, so the app pays full network cost on every screen *and* carries the dead entries until something evicts them. That is why a persisted client store is stamped with a schema or build version and dropped when the stamp changes — an empty store is strictly better than a store full of unusable entries. It is also why a rename is treated as a breaking change on the client side even when the server keeps both fields alive during a transition: the old clients keep working, but no released client's stored data survives the switch to the new name. ## What to say when asked how you would reduce misses Three levers, in order of cost. First, colocate field requirements in named fragments so that operations touching the same type genuinely select the same fields. Second, accept a slightly larger list payload where drill-in is the dominant navigation. Third, where a screen can usefully render before all its data arrives, split the operation so the cheap, already-cached part reads locally and the expensive part loads separately. What you should not offer is loosening completeness globally: you would trade a predictable network call for unpredictable undefined fields in view code.
- The entity is in the store with the right key, yet the client still fetched. What went wrong?Nothing went wrong — identity and completeness are separate conditions. The key told the client where to look; the read then needed every selected field on that entity, and at least one had never been written by any earlier response. Keying is what makes the entry findable; the field set is what makes it usable. The fix is to align the selections, usually by having both operations select the same named fragment.
- Why does a persisted client store usually get thrown away after a schema change?Stored entities carry the field keys that were current when they were written. Rename or restructure a field and no stored entry can satisfy a read selecting the new name, so every read misses while the dead entries still occupy space. Stamping the persisted store with a schema or build version and dropping it on mismatch turns a slow, confusing degradation into one cold start, which is the better trade.
- Would returning the fields the store does have be an improvement?Only where the caller is written for it. Handing back a result that does not match the requested shape pushes undefined-field handling into every view, and it hides the fact that the data is stale or absent. Some clients expose that as an opt-in policy for a specific read, and a narrowed follow-up fetch for just the missing fields is another option, but a fragment discipline usually removes the problem more cheaply than either.
saying these in an interview costs you the question
- Thinks a hit needs only the same operation name
- Expects the store to return a partly filled object by default
- Treats a stored null as a missing field
- Assumes overlapping operations always warm each other
- Believes a renamed field leaves stored entries usable
- Says the cache should silently drop fields it lacks