skip to content

Your service already runs the mapper's shared cache, and you are asked to add an application cache of assembled models. How do you decide, and what does running both cost?

level: principalimportance: should knowfreq 45%

answer

  1. profile before choosing a layer
  2. rows versus assembled answers
  3. the self-maintaining layer is cheaper
  4. two layers, two invalidation surfaces
  5. keep scopes disjoint, keep it reversible

basics

~20 s

Decide by where the cost is: row fetching favours the mapper's tier, which maintains itself; assembly favours an application entry, which the mapper cannot help with. Running both doubles the invalidation surface, so keep their scopes disjoint.

solid answer

~50 s

Start from the profile, not the preference. If the expensive part is fetching rows by identifier, the mapper's own tier already covers it and, crucially, **maintains itself** — the mapper knows when it wrote a row. If the expensive part is assembly (joins, aggregation, several stitched calls, serialization), no row cache can help, because it caches rows and the cost is in turning them into an answer. Adding a second layer is not free: every write path now has two families of entries to account for, entries above the mapper carry a shape that redeploys, and a confusing read can now be explained by either layer. If both are justified, keep their scopes **disjoint** — mapper tier for by-identifier row reads, application cache for expensive assembled answers — so one write never has to fan out unpredictably across both.

go deeper

for a junior

Note the basic split: the mapper's cache saves fetching rows, an application cache saves building the answer. They are not interchangeable.

for a middle

Be able to attribute a slow read to fetching versus assembly, and say which layer removes that specific cost and why the other cannot.

for a senior

Argue the operational cost — a second invalidation surface, shape versioning across deploys, and telling which layer served a stale value during an incident.

for a principal

Own the standing obligation and the exit: enforce the mechanism in one component, bound every entry's lifetime, keep scopes disjoint, and make disabling the cache a configuration change.

## Ask what the read is actually paying for The two layers remove different costs, so the decision is empirical before it is architectural. Profile one slow read and attribute its time: - **Round trips and row fetching dominate** — many small lookups by identifier, a graph walked one association at a time. The mapper's own shared tier is the fit. It is invisible to your code and it invalidates itself when the mapper writes. - **Assembly dominates** — wide joins, aggregation, a result stitched from several queries, then projected and serialized. A row cache changes nothing here; the rows were cheap and the work was afterwards. Only caching the finished answer removes it. - **Neither dominates and the read is simply frequent** — the honest answer is often that no cache is warranted yet, and an index or a narrower projection is the cheaper fix. ## What the second layer costs 1. **A second invalidation surface.** The mapper's tier is maintained by the component doing the writing. An application entry is maintained by your write paths, every one of them, forever — including the ones added next quarter by someone who does not know the cache exists. This is the dominant long-run cost and it is paid in code that is easy to forget rather than in latency. 2. **Shape coupled to deploys.** Row data is governed by the schema; an assembled model is governed by code that ships independently. Entries written by the previous version outlive it, so the key needs a shape version and a rolling deploy needs thinking about. 3. **Ambiguous debugging.** A user reports stale data. With one layer there is one explanation; with two there are three, counting the interaction. Both layers need to be observable — hit rate, age, and a way to say which one answered — or incident time grows. 4. **Compounded staleness.** A miss in the upper layer can be filled from a hit in the lower one, so an assembled answer can be built from row data that was itself already old. The worst case is the sum of the two windows, not the larger of them. ## Deciding, in order | Question | If yes | If no | |---|---|---| | Is the dominant cost assembly rather than fetching? | An application entry is the only layer that can help | Prefer the mapper's tier or a better query | | Must the cached answer be shared across processes? | Favours an external application store | An in-mapper tier may suffice | | Can every write path be found and updated? | Two layers are maintainable | Keep one, and let it be the self-maintaining one | | Is a bounded staleness window acceptable for this data? | Caching is on the table at all | Do not cache; fix the read | ## If you run both, make their scopes disjoint The failure mode of two overlapping caches is a write that must invalidate an unpredictable set of keys in both. Avoid it by construction: - Let the mapper's tier own **by-identifier row reads** and nothing else. - Let the application cache own **expensive assembled answers**, keyed by the inputs that produced them. - Do not cache the same read in both. A read that hits above never reaches the mapper, so the lower entry only adds memory and another way to be stale. - Prefer dropping an application entry over rewriting it on a write, so the next read reassembles from whatever the lower layers currently hold, rather than freezing a value the write path composed by hand. ## The organisational half of the answer The durable cost of this decision is not latency, it is the standing obligation. A cache above the mapper is a piece of shared state your team now owns: its freshness rules, its key discipline, its shape versioning, its behaviour during deploys and during a store outage. The right question in the design review is not "will this be faster" — it will — but "who maintains the invalidation when the person who wrote this has moved on". Two things make the obligation survivable. Enforce the mechanism in one place rather than by convention: a single component that builds keys, populates after commit and drops on write, so no service reinvents it. And give the cache a bounded lifetime so that every mistake in it is self-correcting, at a cost you chose deliberately rather than one you discovered during an incident. Finally, keep the reversal cheap. A cache added behind a clean boundary can be turned off under load or during an incident and the system merely gets slower; a cache the code has come to depend on for correctness cannot. Design for the first.

  • Why does the mapper's own tier carry a lower maintenance cost than one above it?
    Because the component that writes the row is the one holding the entry, so invalidation is internal and automatic. An application entry sits above the writer, which knows nothing about it, so every write path has to be taught — and can forget.
  • If both layers are enabled, what is the worst-case staleness of an assembled answer?
    Roughly the sum of the two windows: the upper entry can be filled on a miss from lower-layer data that was already old, and then held for its own lifetime. Reason about it as an accumulation, not as the larger of the two.
  • When would you argue against adding either cache?
    When the read is slow for a fixable reason — a missing index, an over-wide projection, a per-row fetch that should be one query — or when the data cannot tolerate any staleness. A cache over a fixable query buys latency and permanently hides the defect.
  • How do you keep the decision reversible?
    Put the lookup and populate behind one boundary the service calls, so removing the cache is a configuration change and the code path still works, only slower. Never let correctness depend on an entry existing, and prove it by running the system with the cache disabled.

saying these in an interview costs you the question

  • Adds a second cache without profiling where the time actually goes
  • Expects the mapper to invalidate entries written above it
  • Caches the same read in both layers for extra safety
  • Treats worst-case staleness as the larger window, not the sum
  • Leaves invalidation to each write path's own discipline
  • Lets correctness depend on the cache being present