How does a cache that stores one entry per request differ from one that normalizes entities, and when does the difference matter?
answer
- what is the unit of storage
- question-to-answer versus record-by-identity
- duplication is the cost of simplicity
- identity and merge policy are the price
- thin record overwriting a rich one
basics
~20 sA keyed cache stores each response under the request that produced it, so one record can sit in several entries, each refreshed separately. A normalized cache stores each record once by identity, so one update reaches every reader.
solid answer
~50 sA **keyed** store treats the response as an opaque document filed under the request: the list request has one entry, the detail request another, and a record in both is stored twice. It is simple and needs nothing from the payload; the price is coarse updating — a change means refreshing whole keys, and two entries can disagree on screen until both are refreshed. A **normalized** store splits responses into records identified by type and id, keeps one copy of each, and stores entries as reference lists. One record update corrects every screen showing it. Its price is real: each record needs a stable identity, relationships must be described somehow, thin records from list views can overwrite richer ones, and dropping an unreferenced record is harder than expiring a key. The difference shows when one entity is on screen in several shapes.
go deeper
Recall the two filing styles: one entry per request, or one copy per record with entries pointing at it. The second avoids the same record being stored twice.
Explain the trade cleanly. Duplication in a keyed store is a correctness risk, not just memory; normalization needs identity, a relationship description and a merge policy to pay off.
Recognise the symptoms that justify the move — one entity in two shapes on a screen, many overlapping lists, hand-written multi-key patching — and know the middle ground of grouped invalidation.
Weigh the machinery against the team. Normalization pushes a contract onto the API (stable ids, consistent shapes) and onto every new developer; decide whether that contract is worth owning app-wide or per surface.
## Two ways to file the same response Every cache of server-owned data must answer one question: what is the unit of storage? There are two answers in wide use, and they lead to different apps. In a **keyed** (document) store, the unit is the **request**. A response is filed whole under a key describing what was asked, and nothing inside it is interpreted. The cache is effectively a map from question to answer. In a **normalized** store, the unit is the **record**. Each response is decomposed: every object with a recognisable identity is stored once in a table keyed by type and id, and the request key stores only the shape of the answer — an ordered list of references plus whatever else the response carried, such as counts or a cursor. ## What changes for the reader | aspect | keyed store | normalized store | |---|---|---| | unit stored | one response per request key | one record per identity, plus reference lists | | duplication | the same record repeats across overlapping entries | stored once, referenced many times | | effect of one record changing | every key holding it must be refreshed or patched | every reader corrects at once | | requirements on the payload | none | a stable identity per record, and a description of relationships | | reading a record you have not asked for | impossible, even if it arrived inside another response | possible, because it is in the table | | dropping data | expire a key | decide when a record no longer referenced may go | | typical failure | two entries disagree on screen | a thin record overwrites a richer one | ## The problem normalization solves Put a list of orders and one order's detail panel on the same screen. In a keyed store they are two entries containing two copies of the same order. Change the order and the two copies go out of step; refresh only one key and the screen contradicts itself. This is not theoretical — a status shown in one place and stale in another is among the most reported cache bugs, and it comes from the storage unit, not from the network. A normalized store makes the order one object with one status. The list entry and the detail entry both point at it, so a single update makes both current, and no code has to know which keys contained it. ## The problems normalization creates - **Identity is mandatory.** A record with no stable id cannot be normalized. Responses that return unkeyed nested objects, or ids that are only unique within one response, force either a synthetic identity or an exception. - **Partial records collide.** A list view often returns a thin version of a record and the detail view a rich one. Merging naively, the thin version can overwrite fields the rich one had. The store needs a merge policy — fields present win, absent fields are left alone — and a way to know a record is incomplete so a screen needing the full shape still fetches it. - **Clean-up is harder.** Expiring a request key is easy; deciding that a record referenced by no live entry may be dropped is reference bookkeeping, and getting it wrong leaks memory or evicts something still on screen. - **The shape has to be described.** Something must say which nested fields are records and how they relate, whether that is inferred from a field naming convention, declared by hand, or derived from a schema the server publishes. That description is a maintenance surface a keyed store does not have. ## Choosing Start keyed. It is less machinery, and a great many screens never show the same record twice at once. Move toward normalization when the symptoms appear: the same entity rendered in two shapes on one screen, many overlapping lists of the same records, a growing pile of hand-written code that patches several keys after one change, or memory spent on repeated copies of the same records in paginated lists. A useful middle ground exists: keep a keyed store but make invalidation deliberate — group keys so that one logical change marks a known family stale, and let stale-while-revalidate reads hide the refetch. That buys much of the consistency without an identity requirement on every record, at the cost of more refetching than a normalized store would need.
- What goes wrong when a list response's thin records are merged into a normalized store?They can overwrite the richer copies a detail view had already stored, so a screen suddenly renders missing fields. The store needs a field-level merge that only writes what actually arrived, plus some notion of completeness so a screen requiring the full record still fetches it rather than rendering a half-filled one.
- Can a keyed store keep two screens consistent without normalizing?Largely, by grouping keys so one logical change marks a known family of keys stale and letting stale-while-revalidate reads cover the refetch. The screens converge instead of updating simultaneously, and you pay in requests, but you avoid requiring a stable identity and a relationship description for every record.
A keyed store is a filing cabinet of photocopied reports; a normalized store is a single card index the reports cite. Correcting one fact means hunting every photocopy, or editing one card.
saying these in an interview costs you the question
- Calls normalization strictly better with no cost named
- Assumes every record arrives with a stable identity
- Merges partial records over complete ones blindly
- Thinks duplication only wastes memory, never correctness
- Expects reference clean-up to be as simple as key expiry