A product-catalog service starts timing out under load. The API gateway in front of it is configured to serve the last successfully cached response for each product page when that happens, with no expiry check during the outage. What risks does this introduce, and how would you bound them?
answer
- TTL ceiling on staleness
- stale-while-revalidate / stale-if-error
- volatile fields (price/stock) are the danger zone
- thundering herd on recovery
- tag/track staleness as a metric
basics
~20 sServing old cached data keeps the site up but customers might see wrong prices or 'in stock' items that sold out. Fix it by capping how old the cached data can be and clearly marking it as possibly outdated.
solid answer
~40 sServing stale cache trades correctness for availability: the page stays up, but price, stock, and promotional data can be minutes-to-hours wrong, which is fine for a description or image but risky for stock/price. I'd bound the risk by putting a hard TTL on how stale a fallback response is allowed to be (e.g. refuse to serve anything older than 15 minutes and show a placeholder instead), by excluding highly volatile fields from the cached fallback (or overlaying an 'availability may have changed' notice), and by tracking staleness age as a metric so an extended outage doesn't go unnoticed. I'd also make sure the cache refill path doesn't itself get overwhelmed once the primary recovers (thundering herd).
go deeper
Should understand that 'stale' means the data might be out of date and give one example of when that's risky (stock count) vs safe (product photo).
Should know to cap staleness with a TTL ceiling, separate volatile from stable fields, and add basic monitoring for how often/long stale data is served.
Should proactively design the staleness bound per data type, anticipate the thundering-herd recovery problem, and know of/apply a pattern like stale-while-revalidate or stale-if-error.
Should set organization-wide policy for which data classes are eligible for stale-serving, tie staleness metrics into SLOs/error budgets, and weigh legal/financial exposure (e.g. mispriced checkout) in the design.
## What serving stale cache actually does Serving stale cache as a fallback means: - on a normal request, the system fetches fresh data and stores (or refreshes) a copy in a cache, or a separate 'last known good' store; - on a failed or timed-out call to the primary source, instead of propagating the error, the code reads from that last-known-good copy and returns it as if current, usually tagging it internally (and sometimes visibly) as stale. This differs from ordinary caching-for-performance in one crucial way: an ordinary cache respects a `TTL` and refuses/refreshes once expired, while a stale-cache fallback is deliberately used past its normal freshness window, specifically because the alternative — an error — is judged worse than showing old data. ## Why teams do it The reason to do this is that many kinds of data are 'good enough slightly wrong' — a product description, an image, a category listing rarely changes minute to minute, so continuing to show yesterday's version is strictly better for the user than a blank error page. It converts a hard outage of the origin into a soft, mostly invisible degradation for read-heavy paths, and it buys the origin system time to recover without every client request piling on with retries. ## The trade-off The core trade-off is **freshness versus availability**, and it is not free even when the data 'looks' harmless. - Prices, stock counts, and promotional flags are the **classic danger zone** — an item can sell out or a price can change between when the cache was captured and when a stale copy is served, and if the checkout flow trusts that same stale number, customers can be charged incorrectly or oversold. - There's also an **operational cost**: keeping a 'last known good' copy around indefinitely means more storage, and the cache warming/refill logic adds complexity that itself needs testing. - And once the primary recovers, if every client's cache is due to refresh simultaneously, the origin can be hit by a **'thundering herd'** of near-simultaneous refill requests right as it's coming back up, causing a second wave of failure. ## Failure modes in production 1. **Unbounded staleness** — in production, the most common failure: without a hard ceiling on how old a stale-fallback response is allowed to be, an extended outage (hours or days) results in customers silently browsing an entire storefront frozen in time, with no visible signal anything is wrong — support tickets pile up before anyone notices the metric. 2. **A second failure mode is stale data leaking into a path that assumed freshness** — e.g. a stale 'in stock: yes' flag reaching the add-to-cart button, resulting in overselling that has to be resolved with refunds. 3. **A third is the thundering-herd recovery problem**, where the fallback mechanism itself becomes the cause of a second outage when the primary comes back online and every stale entry tries to refresh at once — mitigated with jittered/staggered refresh and stale-while-revalidate patterns rather than a synchronized cold refetch. ## How the caching layer bounds it CDN and HTTP caching define this formally via the `stale-while-revalidate` and `stale-if-error` `Cache-Control` directives from RFC 5861: `stale-if-error` explicitly tells a cache 'if the origin errors or times out, you may serve this stale response for up to N seconds instead of surfacing the error,' which is exactly the fallback-on-timeout behavior described in the scenario, with the bound (N seconds) baked into the contract. Content platforms like news sites and e-commerce catalogs commonly configure this at the CDN/edge layer so that when an origin has a bad deploy or a database blip, readers still get yesterday's homepage instead of a 502, while the staleness window is capped to something like 5 to 30 minutes so the blast radius of wrong data stays small and self-correcting once the origin heals.
- How is serving stale cache different from a normal TTL-based cache?A normal cache respects its TTL and refuses/refreshes stale entries as a routine part of every request. A stale-cache fallback deliberately serves data past that TTL, but only as an exception path triggered by a dependency failure — it's an availability safety valve, not the everyday behavior.
- What would you do differently for a field like 'items in stock' versus 'product description' when deciding what to serve from a stale cache?Product description is low-risk to serve stale almost indefinitely since it rarely changes and being wrong has low cost. Stock count is high-risk — I'd either exclude it from the stale fallback (show 'availability unknown' instead) or re-validate it at the moment of add-to-cart/checkout even if the rest of the page is served stale.
- How do you prevent the thundering-herd problem when the origin recovers after an extended outage and thousands of stale cache entries all need refreshing?Stagger the refresh with jitter so requests don't all hit the origin in the same second, and/or use a stale-while-revalidate pattern where one request refreshes the entry in the background while others continue getting the still-slightly-stale cached copy. Rate-limiting refill concurrency at the cache layer also protects the recovering origin.
Like keeping yesterday's newspaper on the porch to hand out when today's delivery truck breaks down — fine for yesterday's weather summary, risky if someone uses it to check today's stock prices.
saying these in an interview costs you the question
- Treats stale-cache fallback as identical to normal caching with no special bound
- No mention of a staleness ceiling/expiry on the fallback itself
- Doesn't distinguish volatile fields (price/stock) from stable ones (description/images)
- Ignores the thundering-herd risk when the primary recovers
- No plan to detect/alert on how long the fallback has been serving stale data