skip to content

How would you decide which responses may enter a framework's in-process response cache, and how would you bound their staleness when several instances run?

level: principalimportance: should knowfreq 38%

answer

  1. eligibility before tuning
  2. deny by default, opt in per route
  3. each process holds its own copy
  4. no cross-process purge exists
  5. short lifetime or version-stamped keys

basics

~20 s

Admit only responses no per-caller input touched, that are expensive to produce and tolerant of being out of date. Since each instance holds its own copy with no cross-process purge, bound staleness with a short lifetime or version-stamped keys.

solid answer

~40 s

Eligibility is a policy decision, not a tuning knob, so I make it deny-by-default: a route opts in, and may only opt in if its response is a pure function of inputs already in the key. Then I ask whether it is worth it - production genuinely expensive, the same key requested often enough to hit, the data tolerant of the staleness window. The harder half is invalidation. An in-process cache lives in one instance's memory, so a write handled by one instance cannot purge the others, and "purge on write" quietly becomes "purge wherever the write landed". I prefer designs needing no purge: short lifetimes that cap staleness, or version-stamped keys that simply miss. I budget the memory explicitly and expect a cold cache after every deployment.

go deeper

for a junior

Focus on the first question only: is this response the same for everybody? If who is asking changes the answer, it does not belong in a cache shared by all callers.

for a middle

Be able to explain that the cache lives inside one process, so several instances hold several copies and an entry cleared in one is untouched in the others.

for a senior

Show the operating view: a memory budget, hit-rate metrics, a kill switch, and the knowledge that a rolling deployment empties everything and the cold path has to survive it.

for a principal

Argue eligibility as policy with a named staleness budget per route, choose invalidation you can actually guarantee across instances, and be willing to conclude the cache is not worth the correctness risk it introduces.

## The eligibility question An in-process response cache is the cheapest cache available - no network hop, no serialization, a lookup in local memory - and the most dangerous, because it sits inside the request path and holds rendered output that may not be shareable. The policy I want is **deny by default with explicit opt-in**, and a short list of conditions the route must meet: 1. **Identity-free.** The body must not depend on who is asking, or on anything not present in the key. A response derived from the caller's session, permissions or account belongs somewhere else. 2. **Expensive enough.** The saving is the handler's cost minus a memory lookup. If the handler is a cheap read, the cache adds risk and memory for a saving nobody can measure. 3. **Repeated enough.** Hit rate is decided by how many distinct keys the traffic produces. A long tail of unique keys fills memory and never hits. 4. **Staleness-tolerant.** Someone has to say out loud how out of date this may be, in seconds. If the answer is "it must never be stale", the route is not eligible. 5. **Bounded in size.** A cache holding a few large rendered bodies competes with the memory the request path itself needs; the entry size times the key count has to fit inside a stated budget. A useful discipline is to write the eligibility statement into the route as a comment or declaration, because the thing that breaks it later is not the cache - it is a handler that starts reading per-caller data a year on, under a caching declaration nobody re-read. ## Why invalidation is the hard half Run more than one instance and the same logical entry exists once per process, each populated at a different moment, each invisible to the others. That has consequences people discover late: | Situation | What actually happens | |---|---| | A write arrives | only the instance that handled it can purge its own copy | | An explicit purge endpoint is called | one instance is cleared, the rest are untouched | | Two callers hit two instances | they can see two different versions of the same page | | A deployment rolls | every cache starts empty and the load lands on the origin data | | An instance is added under load | it serves nothing from memory until it warms | So a design that relies on "delete the entry when the data changes" is only honest with a single process, and single-process assumptions have a habit of outliving their truth. The approaches that survive: - **A short lifetime as the contract.** The entry expires on its own, and the stated staleness bound is that lifetime. Simple, provable, and it makes the trade explicit rather than implicit. - **Version-stamped keys.** Include a token that changes when the underlying data does, read from a cheap shared source. A change produces new keys; old entries are never purged, just abandoned to expiry. This converts invalidation into a read, which every instance performs independently. The cost is that the token lookup is now on the hot path, and if it is not cheap the cache has bought little. - **Fan-out notifications.** Tell every instance to drop a key. This works, but it adds an always-on delivery path whose own failures produce stale data silently, which is a real operational surface to own. ## The decision I would actually make I treat the in-process response cache as a narrow instrument: a small number of expensive, shared, slow-moving responses, with a lifetime short enough that nobody needs a purge, and no route added without someone naming the staleness they accept. When the pressure is really on the data behind the handler rather than the rendering, caching further down usually beats caching the rendered response, because the entries are smaller, more reusable across routes, and far less likely to carry someone's personal data. Things I want in place before switching it on anywhere: - a **memory budget** with entry-count and size limits, so the cache cannot starve the request path; - **hit rate and entry-count metrics**, because a cache nobody measures is a liability with no upside evidence; - a **kill switch** that empties and disables it without a deployment, because the first sign of a bad key is a user report, and the response to that must be seconds, not a release; - a **review rule** that re-checks eligibility whenever a cached route's handler changes; - an expectation, written down, that **deployment empties every cache**, so capacity planning assumes the cold path can carry full traffic. ## The honest summary The cache is free performance only while the key is right, the data is shared, and somebody owns the staleness number. The moment any of those three stops being true it becomes a correctness problem that hides behind a performance win, which is why the eligibility decision belongs to policy and review rather than to whoever is optimizing a slow endpoint that week.

  • Why do version-stamped keys often beat an explicit purge for an in-process cache?
    Because a purge has to reach every process, and the delivery path can fail silently. A key carrying a version token turns invalidation into a read each instance performs for itself: a new version simply produces new keys, and abandoned entries fall out on expiry with nobody coordinating anything.
  • What does a rolling deployment do to this cache, and why does it matter for capacity?
    Every replaced instance starts empty, so for a period the full traffic hits the uncached path at once. If the origin data cannot carry that, the cache has become load-bearing - meaning you are running above your real capacity and the deployment itself is the outage trigger.
  • When would you cache the data behind the handler instead of the rendered response?
    When the expense is the reads rather than the rendering, when the same data feeds several routes or formats, or when the rendered output is personalized while the underlying data is not. The entries are smaller, reused more widely, and far less likely to hold anything caller-specific.

saying these in an interview costs you the question

  • Treats in-process caching as a tuning switch rather than a policy decision
  • Assumes purging one instance clears the entry everywhere
  • Has no stated staleness budget for a cached route
  • Ignores that the cache competes for the same memory as the request path
  • Plans capacity as if caches were never cold after a deployment
  • Adds routes to the cache without rechecking eligibility when handlers change