In a meta-framework app running as several instances behind a load balancer, what does a shared cache store give you that per-instance memory does not?
answer
- where the bytes physically live
- one process, one private copy
- answers flip between refreshes
- one clock per copy
- state of the app, not the process
basics
~20 sA shared store keeps one copy of regenerated pages and cached data that every instance reads and writes, so all instances agree. Per-instance memory gives each process its own copy, and those copies drift apart.
solid answer
~40 sMost meta-frameworks default to caching rendered pages and fetched data in the memory of the process that produced them. With one instance that is invisible. Scale out and each process builds its own entries, starts its own lifetime clock for each one, and never sees what the others hold, so the same URL can come back fresh on one refresh and stale on the next depending on which instance answered. A shared store moves those bytes out of the process into something every instance reads and writes, keyed the same way by all of them. Then a cached page is one copy with one lifetime, a regeneration serves the whole fleet, and an instruction to drop an entry is observed by everyone rather than by whichever process happened to receive it.
go deeper
Remember the one sentence: by default the cache lives inside the process, so more processes means more copies. Be able to name the symptom, a page that is fresh on one refresh and stale on the next.
Explain the mechanics: separate entries, separate lifetime clocks, a purge landing in one process, and a cold start per instance. Then describe what moving the entries to a shared store changes for each of those.
Show that you treat the store as a dependency on the request path: bounded latency, reads degrading to a miss, writes allowed to fail, and nothing user-specific stored where the whole fleet can read it.
Frame it as where cache authority belongs. Argue which classes of entry justify a shared dependency at all, and what the fleet should do when that store is unavailable or slow.
## A cache is bytes in a specific place It is easy to talk about *the* cache as if an application had one. It does not. A cache is a set of bytes sitting somewhere concrete, and the first question to ask of any caching feature is **where those bytes physically live**. Meta-frameworks that can prerender pages, regenerate them later, and cache the data used to build them have to put those results somewhere, and the cheapest place available to a server is **memory inside the process that rendered them**. That is the common default, it is fast, it needs no configuration, and it is completely invisible while the application runs as exactly one process. It stops being invisible the moment a second process exists — a second container, a second replica, a second machine — because a load balancer now spreads requests across processes that cannot see each other's memory. ## What per-instance memory actually means When each instance keeps its own cache: - An entry written by instance A **cannot be read** by instance B. B renders the route itself and keeps a second, independent copy. - **Lifetimes are per copy.** If a cached page may live ten minutes, each instance starts its own ten minutes when it first rendered that route, so the copies go stale at different wall-clock times. - An instruction to drop or refresh an entry is executed by **whichever instance received that request**. The rest keep serving what they hold. - Replacing or restarting an instance erases only that instance's entries, so the fleet is never uniformly warm or uniformly cold. Because the balancer is free to send consecutive requests from the same user to different instances, the user is effectively **sampling a different cache on every request**. The symptom is not a permanent wrong answer, which would be easy to spot; it is an answer that flips, in proportion to the number of instances. ## What a shared store changes A shared store is a store outside all of the instances that each of them talks to over the network: read an entry by key, write an entry, remove entries. The framework keeps deciding *what* to cache and *how the key is computed*; the store only keeps bytes. Once the entries live there: | Property | Per-instance memory | Shared store | |---|---|---| | Who can read an entry | only the process that wrote it | every instance | | Lifetime of an entry | one clock per process | one clock for the app | | Cost of a cold entry | one render per instance | one render for the fleet | | Effect of dropping an entry | clears one process | clears the copy everyone reads | | Effect of replacing instances | everything is lost | entries outlive the release | The last row is the double-edged one: outliving a release is exactly why a shared store is useful and also how a page built by the previous build can still be served after the new one is live. ## What it costs, and what it does not fix A shared store is not free: 1. **A network hop joins the render path.** A slow store can cost more than re-rendering a cheap page, so entries have to be worth fetching. 2. **A new dependency sits in front of every request.** Reads should degrade to a miss and a fresh render rather than to an error page, and writes should be allowed to fail quietly. 3. **Payloads must survive leaving the process**, so whatever is cached has to be serialisable — rendered markup, serialised data, headers — not live objects. 4. **Everything stored is readable by every instance**, so a response that depends on who is asking does not belong in a store shared by the whole fleet. And it does not fix the layers above it. A CDN or other shared cache in front of the app keeps **its own copies**, which no change inside the application touches; the store also does not stop several instances from missing the same cold entry at the same moment and all rendering it at once. ## The rule to carry away Cache state that has to be consistent for users should be a property of **the application**, not of a process. Per-instance memory is legitimate as a small, short-lived layer in front of a shared store, where being a little behind for a second is harmless; it is not legitimate as the authority for anything a user is meant to see change.
- Why does per-instance memory usually look correct in a test environment?Test environments commonly run a single instance, so there is only one cache, one lifetime clock and one place a purge can land. The behaviour is genuinely consistent there. Divergence needs at least two processes to appear, which is why this class of bug is first seen after scaling out or after enabling a second replica in production.
- Does a shared store mean instances should keep nothing in memory?No. An instance may keep a very short-lived copy in front of the shared store to save the network hop on hot entries. The rule is about authority: the shared store holds the version everyone must agree on, and the in-process copy is allowed to be a second or two behind, never to be the thing a purge has to reach.
- If every instance computes the same cache key, is that enough to share a cache?No. Identical keys are necessary but not sufficient; they only mean two processes would look under the same name. Sharing requires one storage location that both processes actually read from and write to. Matching keys over separate memories simply produce two independent entries that happen to be called the same thing.
Six clerks each keeping a private notebook of answers, versus all six reading one wall board. The notebooks quietly disagree; the board cannot.
saying these in an interview costs you the question
- Assuming the balancer pins a user to one instance, so caches agree
- Thinking a cache lifetime is global when each process runs its own clock
- Believing one regeneration updates every instance at once
- Treating identical cache keys as proof the cache is shared
- Calling the CDN the shared cache and ignoring the app's own copies
- Dismissing divergence because a restart clears memory anyway