skip to content

Invalidation Strategy

Keeping caches honest after writes: picking TTLs by resource volatility, purging CDN entries, and deciding between versioned URLs and validators. Interviewers ask because invalidation is where caching designs actually fail in production.

part ofAPI stylesoverview, primer and where to startread it →
on this pageshow

questions

4

How do you choose a TTL for an HTTP API response — the number you put in Cache-Control: max-age or s-maxage — for resources that change at very different rates?

level: middleimportance: must knowfreq 55%

answer

  1. Tolerance for staleness, not change rate
  2. s-maxage long (purgeable) + max-age short
  3. stale-while-revalidate kills stampedes
  4. Immutable URL → max-age one year
  5. TTL is the safety net for a missed purge

basics

~20 s

Pick the TTL from how stale the data may safely be, not from how often it changes. Volatile or personalised resources get seconds or no shared caching; stable reference data gets minutes to hours; immutable, uniquely-addressed content gets a year. Use s-maxage to give shared caches a different, usually longer, value.

solid answer

~50 s

Start from the business question — **how stale may this be before someone is harmed?** — not from the change rate. A resource that changes every second can still tolerate a 10-second TTL if a 10-second-old view is acceptable. A workable ladder: - **Immutable, content-addressed representations** (a hash or version in the URL): `max-age=31536000, immutable`. Nothing to invalidate because the URL changes instead. - **Stable reference data** (currency lists, catalogue taxonomies): minutes to hours, with revalidation after expiry so a refresh is cheap. - **Frequently-read, slowly-changing collections**: seconds to a minute at the shared cache, which absorbs most traffic while keeping visible lag small. - **Per-user or authorization-sensitive data**: `private` or `no-store`; never let a shared cache hold it. Use `s-maxage` to give the CDN a longer TTL than browsers, since you can purge the CDN but cannot purge a browser. Pair a short TTL with `stale-while-revalidate` so expiry serves a slightly old copy instantly and refreshes in the background instead of stampeding the origin.

code

http · 3 lines
http
HTTP/1.1 200 OK
Content-Type: application/json
Cache-Control: public, max-age=10, s-maxage=600, stale-while-revalidate=60

go deeper

for a junior

Know that max-age sets how long a response may be reused, that user-specific data should be private or no-store, and that different endpoints deserve different values.

for a middle

Explain the s-maxage versus max-age split, choose TTLs from staleness tolerance, and know what stale-while-revalidate buys you.

for a senior

Argue the long-TTL-plus-purge strategy, treat the TTL as a bound on missed-purge damage, and cite hit ratio, origin rate and observed post-write staleness as the measures.

for a principal

Frame TTL as a per-resource contract about staleness that is documented for consumers, with platform defaults, guardrails against caching authorization-dependent data, and an explicit position on freshness versus origin cost.

## The real question is tolerance, not volatility Engineers instinctively set TTL from how often data changes. That is the wrong input. The right input is **how long a stale answer remains acceptable to the consumer of this API**. A stock level that changes many times per second may still be perfectly serviceable at 15 seconds old on a product page, while a feature flag that changes twice a year may need to take effect within seconds when it does change. Volatility tells you how *often* you will serve something stale; tolerance tells you whether that matters. So the design conversation is: what breaks if a client sees data N seconds old? For a public catalogue, nothing. For a permissions or entitlement response, quite a lot. For a payment status, the user reloads and complains. Write the tolerance down per resource; the TTL falls out of it. ## The freshness directives you are choosing between - `max-age=N` — how long any cache may serve the response without checking back. - `s-maxage=N` — overrides `max-age` for **shared** caches (CDN, reverse proxy) only. This split is the single most useful lever in API caching, because shared caches are purgeable and private browser caches are not. - `private` — only a single user's cache may store it; shared caches must not. - `no-store` — do not persist at all. - `stale-while-revalidate=N` — after expiry, a cache may serve the stale copy for up to N seconds while it refreshes in the background. - `stale-if-error=N` — a cache may serve stale content for up to N seconds when the origin is failing. A short `max-age` plus a generous `stale-while-revalidate` is often the best shape for an API: users almost always get an instant answer, the origin sees one refresh per key per window rather than a stampede, and worst-case staleness stays bounded and explicit. ## A ladder by resource class **Immutable representations.** If the URL uniquely identifies one byte-for-byte version — a content hash, a build id, an immutable revision — cache for a year with `immutable`. There is no invalidation problem because you never need to invalidate; you publish a new URL. This is the cheapest correct caching in existence and worth engineering the URLs for. **Reference and configuration data.** Country lists, category trees, price books. Minutes to hours. These are read constantly and changed rarely, so the hit ratio is enormous and the risk is a delayed rollout of a change — which is exactly what an explicit purge fixes. **Hot collections and detail resources.** Product listings, public profiles, search facets. Seconds to a minute at the shared cache. The economics are dominated by request volume: even a 10-second TTL on a page served ten thousand times a second collapses origin load by orders of magnitude. **User-specific and authorization-dependent data.** Mark `private`, or `no-store` where the content is sensitive. If the response varies by who is asking, a shared cache that ignores that is a data-leak bug, not a performance tradeoff. Where you do cache per-user, be extremely careful about which request attributes the cache key includes. **Write results and status endpoints.** Generally `no-store`. A cached "pending" answer that outlives the transition is a support ticket. ## Long shared TTL plus purge: the strategy that scales The most effective pattern in production APIs is to set an **aggressively long `s-maxage`** and rely on **explicit purge at write time** rather than on expiry. Expiry-driven caching forces you to trade freshness against hit ratio, and you must pick a compromise that is simultaneously too stale for correctness and too short for efficiency. Purge-driven caching breaks the trade: the CDN holds content until you tell it not to, so the hit ratio is near-total, and the propagation delay after a write is the purge latency rather than the TTL. The cost is that you must actually issue the purges, for every representation a write affects, reliably. That is real work, and a purge you forget is unbounded staleness — which is why the TTL still matters as a **safety net**. A long-but-finite `s-maxage` bounds the damage of a missed purge; `s-maxage` of a year with no purge on a mutable resource is a bug waiting to happen. ## Practical guidance Set TTLs per route rather than globally; a single platform default guarantees that some resources are too stale and others are needlessly uncached. Keep the browser's `max-age` short and the CDN's `s-maxage` long, because you can purge one and not the other. Publish the intended staleness in your API documentation so consumers can reason about it. And measure: cache hit ratio, origin request rate, and observed staleness after writes are the three numbers that tell you whether the choice was right.

  • Why set s-maxage much higher than max-age rather than using one value for both?
    Because shared caches are controllable and private caches are not. A CDN or reverse proxy can be purged the moment data changes, so it can safely hold content far longer, absorbing most traffic. A browser cache cannot be reached, so anything it holds is stale for the full duration no matter what happens at the origin. Splitting the two gives you a high hit ratio without long uncontrollable staleness on clients.
  • What does stale-while-revalidate change about origin load at expiry?
    Without it, every request arriving after expiry either blocks on revalidation or is forwarded to the origin, so a popular key produces a burst of simultaneous origin requests — a stampede. With it, the cache immediately serves the stale copy and refreshes once in the background, so the origin sees roughly one request per key per window and clients never wait for the refresh, at the cost of a bounded, explicitly chosen extra staleness.
  • If you plan to purge on every write, why not set an effectively infinite TTL?
    Because purges fail. A dropped purge call, a missed related resource, or an outage in the purge path leaves content cached with no expiry to rescue it, and the staleness is unbounded until someone notices. A long-but-finite TTL is the backstop that converts a missed purge from a permanent bug into a bounded one, and it is why long TTL plus purge, not infinite TTL, is the production pattern.

A newspaper versus a departures board: the paper is stale by design and nobody minds, while a board two minutes behind causes missed flights — same volatility, wildly different tolerance.

saying these in an interview costs you the question

  • Deriving TTL from how often data changes rather than how stale it may safely be.
  • Applying one global TTL to every endpoint in the API.
  • Letting user-specific or authorization-dependent responses be cached by shared caches.
  • Assuming a long TTL is safe purely because a purge is issued on write, with no finite backstop.
  • Choosing a very short TTL to stay fresh and then being surprised by origin stampedes at expiry.

context

open as a page

When would you invalidate cached API content by changing the URL — putting a version or content hash in the path — rather than relying on the origin being revalidated? What do you give up with each approach?

level: middleimportance: should knowfreq 40%

basics

~20 s

Change the URL when the representation is immutable and the client learns the new URL from somewhere else — then cache forever with no invalidation at all. Rely on revalidation when the identity must stay stable, at the cost of a round trip per check and an origin that must stay reachable.

open as a page

Explain surrogate-key (cache-tag) based purging in a CDN such as Fastly or Varnish, and how it differs from purging by URL.

level: seniorimportance: should knowfreq 42%

basics

~20 s

The origin tags each response with surrogate keys naming the data it contains. The CDN indexes cached objects by tag, so one purge of a key evicts every representation containing that data — regardless of URL, query string or variant — instead of you enumerating URLs.

open as a page