How would you decide where a scaled-out meta-framework app's durable cache should live: a shared store, per-instance memory, or a CDN?
answer
- requirements first, mechanism second
- how wrong may two instances be
- what a miss costs, times instances
- can you actually purge it
- memory accelerates, it never decides
basics
~20 sPut cache authority where a correction must be observable and a miss is expensive. Anonymous URL-addressable documents belong in front of the app, fleet-consistent expensive results in a shared store, and per-instance memory is a latency layer, never the authority.
solid answer
~50 sStart from requirements rather than from stores. Ask how wrong two instances may be, how expensive a miss is and how many instances would pay it, how quickly a correction has to become visible, and what the site should do when the cache is unavailable. Those answers usually sort themselves: responses identical for everyone are cheapest to serve from a cache in front of the application, because it absorbs traffic before it arrives; results that are costly to produce and must agree across instances justify a shared store, accepting it as a dependency on the render path; per-instance memory is legitimate only as a short-lived layer in front of an authority, or for cheap entries where brief divergence is harmless. The decisive constraints are the purge path you can actually operate and what happens when the store is down.
go deeper
Know that a cache has to live somewhere and that the choice has consequences: inside a process, in a store the instances share, or in front of the application. You are not expected to make the call yet.
Be able to state what each placement is good at and what it costs, and to explain why memory cannot be the authority for anything that must be corrected on demand.
Derive the placement from requirements: consistency tolerance, miss cost, correction time, failure behaviour. Show you would design the shared store as fail-soft and limit duplicate renders on a cold entry.
Set the allocation rule for the organisation, with correction-time commitments, an explicit failure budget when the cache is down, ownership of the purge path, and a standing rule keeping per-person output out of shared placements.
## Decide from requirements, not from stores The failure mode of this decision is choosing a mechanism first and discovering the requirements afterwards. Five questions settle most of it: 1. **How much may two instances disagree, and for how long?** Some pages can differ for a minute without anyone caring. A price, a permission or a published correction cannot. 2. **What does a miss cost, and who pays it?** A page assembled from several slow upstream calls is worth keeping; a page rendered from data already in hand may be cheaper to rebuild than to fetch over the network. 3. **How concentrated is the traffic?** A small hot set is cached well almost anywhere, including per-instance memory, because every instance warms up quickly. A long tail warms up slowly and multiplies the cost of every instance keeping its own copy. 4. **How fast must a correction be visible?** This sets which lifetimes are acceptable and whether an on-demand purge path is mandatory. 5. **What happens when the cache is unavailable?** A shared store on the render path is a new dependency; if the fleet cannot serve without it, availability has just been coupled to it. ## What each placement is actually good at | Placement | Good at | Costs and limits | |---|---|---| | Cache in front of the app | absorbing traffic before it reaches any instance; one copy per URL | only for responses safe to reuse for everyone; purging it is a separate operation you must own | | A store shared by instances | consistency across the fleet; one render serves everyone; survives restarts | a network hop per read, a dependency on the request path, entries readable fleet-wide | | Per-instance memory | lowest latency; no dependency; trivially correct for one process | one copy per process, diverging; a purge reaches only one; cold after every replacement | Read down the middle column and the allocation follows: put anonymous documents where they can be served without touching the application, put expensive fleet-consistent results in the shared store, and keep memory as a small accelerator in front of an authority that a purge can actually reach. ## The constraints that decide close calls - **The purge path is the real constraint.** If there is no operable way to remove a copy on demand, the only remaining control is the lifetime, and the correction time becomes whatever that lifetime is. Teams routinely choose a placement they cannot invalidate and then discover their publishing requirement is impossible. - **Failure budget.** Design the shared store as fail-soft, reads degrading to a miss, and then ask honestly whether the fleet can carry full render load in that state. If not, the store is a single point of failure wearing a cache's clothes, and needs the same treatment as one. - **Blast radius of a mistake.** A fleet-wide store is readable by every instance, which makes a keying mistake on personalised content a disclosure rather than an inefficiency. That argues for keeping personalised rendering out of shared placements entirely and delivering the personal part separately. - **Cold-start economics.** Anything that clears on release or on scale-out has to be survivable at full render load, or it needs deliberate warming. - **Who operates it.** A shared store is capacity, credentials, monitoring and an on-call owner. That cost is real and is paid continuously, unlike the one-off cost of choosing it. ## A defensible allocation For most applications the answer is layered rather than singular: 1. Serve anonymous, URL-addressable documents from the cache in front, with lifetimes short enough to match the publishing requirement, plus an on-demand purge. 2. Keep expensive, fleet-consistent results in a shared store, fail-soft, with duplicate renders limited so one cold entry does not become one render per instance. 3. Allow a small per-instance layer with a lifetime measured in seconds, explicitly documented as allowed to lag, never as the thing a purge must reach. 4. Keep per-person output out of every shared placement, and assemble it separately. ## What makes this a lead's question There is no correct universal answer, and the strong response is not a store but a set of stated requirements, an allocation justified against them, and an owned failure mode. The weak response picks the most capable option everywhere, doubles the operational surface, and still cannot say how long a wrong page stays wrong.
- When is per-instance memory the right answer even in a large fleet?When entries are cheap to rebuild, the working set is small and hot, and brief divergence between instances is invisible to users, for example a short-lived layer in front of a shared store. It is also right when adding a network dependency to the render path would cost more availability than the cache buys in latency.
- How do you decide the lifetime for responses stored in front of the app?From the publishing requirement, not from a performance target. If a correction must be visible in five minutes and there is a reliable purge path, the lifetime can be long because purging is the control. Without a purge path, the lifetime is the control, so it has to be shorter than the correction time you promised.
- What would make you refuse a shared store despite consistency needs?If the fleet cannot serve acceptably when it is unavailable, and nobody will own its capacity and monitoring. A cache that becomes a single point of failure has traded a consistency problem for an availability one. The alternative is to reduce what needs fleet-wide consistency, for instance by moving the shared copy in front of the app where it is already replicated.
- How do you keep this decision from being relitigated per team?Write it down as an allocation rule with a reason attached: which classes of response go where, what the maximum correction time is, and what the fail-soft behaviour must be. Teams then argue about whether their case fits a class, which is a much shorter conversation than choosing a cache from scratch each time.
saying these in an interview costs you the question
- Choosing a store before stating any consistency requirement
- Treating a shared store as free consistency with no dependency cost
- Picking a placement with no operable purge path
- Assuming per-instance memory is acceptable because it is fast
- Ignoring whether the fleet can serve when the shared store is down
- Using one placement everywhere to avoid thinking about classes