A store selector rebuilds a filtered, sorted list on every notification — how do you memoise it and why might its cache never hit?
answer
- two stages, one cache
- input steps pluck, combiner creates
- identity comparison, not contents
- a manufactured input never matches
- one entry, many arguments, constant thrash
basics
~20 sSplit it: cheap input steps that pluck references, then one memoised combining step holding the filter and sort, cached on input identity. The cache never hits when an input step builds something fresh, or one entry serves differing arguments.
solid answer
~50 sA selector derives a reader's slice from store state, and it runs whenever that reader produces output — often on notifications about state it never touches. The fix is two stages: trivial input steps returning references (the collection, the criteria, the sort key) and one combining step holding the expensive filter-and-sort plus a cache of its last inputs and result, compared by identity. Immutable updates make that work, because an untouched collection is the same object. The classic reason such a cache never hits is an input step that manufactures a value — a filtered collection, an object assembled from fields, a substituted default — so identity differs every call. The second is one single-entry cache shared by readers passing different arguments, each evicting the other. Fix the first by constructing only after the cache; the second with a per-reader or keyed cache.
go deeper
Remember what a selector is: a function from store state to the slice one reader needs, which runs each time that reader produces output.
Explain the two-stage shape — reference-plucking input steps, then one cached combining step — and why the comparison is on identity rather than contents.
Recognise both classic defects from the symptom alone: an input step that builds something, and one single-entry cache shared across differing arguments. Verify with a hit counter.
Decide the caching topology: per-reader instances, a keyed cache with an eviction rule, or one derived map. At scale that is a memory budget, not only a speed choice.
A **selector** is a function from the whole of a store's state to the slice a particular reader wants. Cheap selectors just reach in and return a reference. The interesting ones **derive**: filter a collection by the current criteria, sort it, group it, total it, join two slices together. Deriving inside a selector is the right place for it — the rule lives once, next to the data, rather than being repeated in every component that needs the result. The problem is that a selector runs whenever its reader is asked to produce output, which can be on every store notification, including notifications caused by parts of the state the selector never touches. If the derivation is heavy, that is real repeated work on something that did not change. ## The two-stage shape The standard answer is to split the selector in two: - **Input steps** are trivial and pure reference plucking: the collection, the filter criteria, the sort key. They run on every call and cost nothing. - **A combining step** holds the expensive derivation and a small cache of its last inputs and last result. Called again, it compares the inputs it was handed — by identity, which is why the input steps must return references rather than build anything — and returns the stored result if they match. That is the meaning of *memoising on input identity*: the cache does not inspect the collection's contents, it asks whether it was handed the same collection object and the same criteria as last time. Immutable updates make this work, because a collection that was not written is the same object, and one that was written is a different object. ## Why the cache never hits Nearly every "my memoised selector still recomputes every time" report is an input step that manufactures a value: | Input step returns | Why the cache misses | |---|---| | a freshly built collection (a filtered, mapped or concatenated one) | new identity each call, even when the contents are identical | | an object or record assembled from several fields | assembled on the spot, so never the same object twice | | a default substituted when the slice is missing | a new empty collection or object each call | | a value read from somewhere the store does not own and that is replaced each turn | changes for reasons unrelated to the derivation | The rule that follows: **input steps return references, the combining step creates**. Anything that constructs belongs after the cache, never before it. If an input genuinely must be assembled, derive it in its own memoised step first, so the assembly has its own cache and a stable result to hand downstream. ## One cache, many callers A second failure is subtler. A memoised derivation typically keeps **one** remembered pair. Give that single cache several callers with different arguments — one component selecting items for user A and another for user B — and each call evicts the other's entry. Every call then misses, and the code is slower than the uncached version because it also pays the comparisons. The ways out: 1. **A cache per reader.** Create a selector instance per component instance, so each has its own entry. 2. **A cache keyed by the argument**, with a bounded number of entries and an eviction rule, so several arguments coexist. 3. **Move the argument into the derivation**, deriving a map from argument to result once, and letting each reader look up its key. Which one is right depends on how many distinct arguments there are and how big each result is: a cache that holds one entry per user of a large app has become a memory decision, not just a speed one. ## Practical rules - Memoise the expensive step only. A filter over a handful of items does not need a cache; a sort and group over thousands does. - Keep derivations out of the reader wherever the same derivation serves more than one reader — one cache beats several. - Remember that a derivation returning a new collection on a miss hands a new identity downstream, so anything that compares by identity sees a change on every genuine recomputation. That is expected; what is not expected is a new identity on a **hit**, which means the cache is not being consulted at all. - Verify a cache rather than trusting it. A counter incremented inside the expensive step, read after a few interactions, settles the question faster than reasoning about it. ## Interview framing Say the shape out loud: cheap input steps that pluck references, one memoised combining step holding the derivation, identity comparison between them. Then name the classic defect — an input step that builds something, so the comparison never matches — and the second one, a single-entry cache shared by callers with different arguments. Finishing with "and I would confirm the cache actually hits before believing in it" is the part that reads as production experience.
- Why is comparing selector inputs by identity rather than by contents the right default?Because a store that updates immutably already encodes change in identity: an untouched collection is the same object, a written one is a new object. An identity check is constant-time and gives a correct answer, whereas comparing contents costs time proportional to the data — sometimes as much as the derivation you are trying to skip.
- How do you handle a memoised derivation that many readers call with different arguments?Give each reader its own selector instance so each has its own entry, or key the cache by the argument with a bounded number of entries and an eviction rule, or move the argument into the derivation and derive a map from argument to result once. Which fits depends on how many distinct arguments exist and how large each result is — it becomes a memory decision, not only a speed one.
- How would you confirm a memoised selector actually hits?Count, do not reason. A counter incremented inside the expensive step, read after a handful of realistic interactions, settles it immediately: a count that tracks the number of calls means the cache never hits, while a count that only rises when the underlying data changes means it does. Reasoning about identity from the code misses manufactured inputs routinely.
saying these in an interview costs you the question
- Puts the filtering inside the input steps rather than after the cache
- Expects a memoised derivation to compare collection contents rather than identity
- Returns a fresh empty collection as a default from an input step
- Assumes one shared single-entry cache serves callers with different arguments
- Believes memoising a selector removes the notification that triggered it
- Trusts a cache is effective without ever checking that it hits