skip to content

What do the HTTP response directives `Cache-Control: stale-while-revalidate=60` and `Cache-Control: stale-if-error=86400` instruct a cache to do, and what risk does each accept in exchange for what benefit?

level: seniorimportance: should knowfreq 34%

answer

  1. swr = serve stale now, refresh behind the scenes
  2. sie = serve stale when origin errors or times out
  3. max age becomes max-age + swr
  4. must-revalidate cancels both
  5. stale hides outages — alert on stale-serve rate

basics

~20 s

stale-while-revalidate=60 lets a cache serve an expired response immediately for up to 60 more seconds while it refreshes in the background, so no user waits on the origin. stale-if-error=86400 lets it serve an expired response for up to a day when revalidation fails or the origin returns 5xx. Both trade bounded staleness for latency and availability.

solid answer

~60 s

Both are RFC 5861 extensions that let a cache serve content past its freshness lifetime, under different triggers. **`stale-while-revalidate=N`**: for N seconds after expiry, the cache returns the stale copy *immediately* and kicks off an asynchronous revalidation. The user never pays the origin round trip; the next user gets the fresh copy. This eliminates the latency cliff at expiry and, at the edge, the thundering herd where every expiring key hits the origin at once. **`stale-if-error=N`**: if revalidation fails — a connection error, a timeout, or a 500/502/503/504 — the cache may serve the stale copy for up to N seconds past expiry instead of propagating the failure. It turns a short origin outage into invisible, slightly-old content. The risk in both cases is bounded staleness that you have now made *explicit policy*: someone will see data up to N seconds old. For a price, an inventory count or a permissions decision that may be unacceptable; for an article, a nav menu or a product page it is almost always the right trade. Neither applies if the response carries `must-revalidate`, which forbids stale serving outright.

code

http · 4 lines
http
HTTP/1.1 200 OK
Cache-Control: public, max-age=60, stale-while-revalidate=30, stale-if-error=86400
ETag: "catalog-7781"
Content-Type: application/json

go deeper

for a junior

Know the two behaviours in one line each: serve stale while refreshing, and serve stale when the origin is broken.

for a middle

Add the triggers and the arithmetic — maximum age is max-age plus the stale-while-revalidate window.

for a senior

Discuss stampede smoothing, per-node behaviour on a CDN, which content classes must never serve stale, and the monitoring gap stale-if-error creates.

for a principal

Treat the numbers as a stated staleness contract for downstream consumers and decide per content class where availability outranks accuracy.

## The problem at the moment of expiry With plain `max-age=60`, the request that arrives at second 61 is the unlucky one: it blocks on a full origin round trip while every earlier request was served from cache in microseconds. At scale this is worse than unfair — if many cache entries share a lifetime, they expire together and produce a synchronised burst of origin traffic, the cache stampede. And if the origin happens to be down at that moment, a perfectly good stored copy is discarded in favour of an error page. RFC 5861 adds two directives that address these separately. ## stale-while-revalidate=N For N seconds after the freshness lifetime ends, the cache is permitted to serve the stored (now stale) response **immediately** and start revalidation in the background. The requesting client waits for nothing. When the revalidation completes, the stored response is replaced and its freshness clock restarts, so subsequent users see fresh content. Effects worth naming: - **Latency becomes uniform.** No user ever pays the origin round trip for this resource, so the p99 stops spiking at expiry boundaries. - **Origin load is smoothed.** In a shared cache one background revalidation serves the whole population behind that node, instead of every concurrent request racing to the origin. - **The effective maximum age becomes `max-age + stale-while-revalidate`.** That is the number to state when someone asks how stale content can get. If nobody requests the resource during the stale window, the entry simply expires out of the window and the next request revalidates synchronously. A subtlety: the guarantee holds per cache node. A CDN with many edge nodes and a cold node still incurs a synchronous fetch, so this is a complement to request coalescing at the edge, not a replacement. ## stale-if-error=N This one triggers only on failure. If the cache tries to revalidate and gets a connection error, a timeout, or a 500, 502, 503 or 504, it may serve the stale stored response for up to N seconds past expiry rather than surfacing the error. It is availability insurance, and the N is usually much larger than for `stale-while-revalidate` — hours or a day — because the point is to ride out an incident. The judgement call is which failure is worse: a user seeing yesterday's content or a user seeing an error page. For marketing pages, documentation, catalogue listings and navigation, stale wins overwhelmingly. For anything where acting on stale data has consequences — prices at checkout, stock levels at the point of commitment, an authorization decision — the error is the honest answer, and hiding the outage may be actively harmful. A second-order risk: serving stale during an outage removes the user-visible symptom, which can delay detection if your alerting is based on user-facing errors rather than on origin health and revalidation-failure rates. Instrument the cache's stale-serve counter and alert on it. ## How they combine and interact A typical edge policy reads `Cache-Control: public, max-age=60, stale-while-revalidate=30, stale-if-error=86400`: fresh for a minute, then served instantly while refreshing for thirty seconds, and held for a day if the origin is unavailable. Interactions to remember: - **`must-revalidate` disables both.** It explicitly forbids serving a stale response, so combining them is contradictory; `must-revalidate` wins. - **`no-store` makes both irrelevant**, since nothing is stored. - **`no-cache` makes both irrelevant for reuse**, since every use requires successful revalidation. - **Support is uneven.** Both are widely implemented in CDNs; browser support for `stale-while-revalidate` exists but is less consistent, and `stale-if-error` less so still. Treat them as best-effort optimisations rather than guarantees you can build correctness on — and note that a CDN can implement equivalent behaviour through its own configuration regardless of the header. ## How to reason about the value of N Pick `stale-while-revalidate` from the shape of your traffic: it needs to be long enough that a background revalidation completes comfortably, so a few multiples of your p99 origin latency. Pick `stale-if-error` from your incident profile: long enough to cover a realistic recovery, short enough that genuinely abandoned content does not linger. And state the resulting worst-case age explicitly in the service's documentation, because you have converted an accident into a policy and someone downstream needs to know the number.

  • With `max-age=60, stale-while-revalidate=30`, what is the oldest content a user can be served?
    Just under 90 seconds old. The response is fresh for 60 seconds, and for the next 30 it may be served stale while a background revalidation runs. Past 90 seconds the stale window has closed and the next request must revalidate synchronously before anything is returned.
  • Why can `stale-if-error` make an incident harder to detect?
    It removes the user-visible symptom: during an origin outage, users keep getting successful responses from the cache, so error-rate alerts based on client-facing traffic stay flat. Detection has to come from origin health checks and from the cache's own revalidation-failure and stale-serve metrics, which need to be exported and alerted on deliberately.
  • What happens if a response carries both `must-revalidate` and `stale-while-revalidate`?
    `must-revalidate` wins and no stale content is served. It explicitly forbids reusing a stale response without successful revalidation, so the stale-serving extension has no permission to act on. Sending both signals a confused policy and should be resolved in favour of whichever guarantee the resource actually needs.

saying these in an interview costs you the question

  • Thinking `stale-while-revalidate` extends freshness — the response is genuinely stale, just servable
  • Using `stale-if-error` on prices, stock levels or authorization decisions
  • Combining either with `must-revalidate` and expecting stale serving to happen
  • Assuming full browser support and building correctness on it
  • Not accounting for `max-age + stale-while-revalidate` when quoting a worst-case data age

context