skip to content

Caching and Conditional Requests

Designing resources so they can actually be cached: cacheability as a REST constraint, choosing ETag or Last-Modified per resource, conditional GET as a polling contract, and how you invalidate. Interviewers ask to see whether caching was designed in or bolted on afterwards.

part ofAPI stylesoverview, primer and where to startread it →
on this pageshow

questions

12

REST lists 'cacheable' as one of its architectural constraints. What does that constraint actually require of an API's responses, and which responses in a typical HTTP API can be cached at all?

level: juniorimportance: must knowfreq 48%

answer

  1. every response labelled cacheable or not
  2. GET/HEAD cacheable; PUT/PATCH/DELETE never
  3. POST only with explicit freshness — in practice, no
  4. caching pays back the cost of statelessness
  5. cacheable is not the same as public

basics

~20 s

It requires every response to say, implicitly or explicitly, whether it may be reused and for how long. In practice GET and HEAD responses are the cacheable ones; POST responses only with explicit freshness; PUT, PATCH and DELETE responses are never cached.

solid answer

~60 s

The constraint is about **labelling**: each response must declare whether a client or an intermediary may store and reuse it, and for how long. That label is what lets caches exist between client and origin without either side coordinating. In practice: **GET** and **HEAD** responses are cacheable by default for the status codes HTTP defines that way (200, 203, 204, 206, 300, 301, 308, 404, 405, 410, 414, 501). **POST** responses are cacheable only if you explicitly give them freshness information, and almost nothing does. **PUT**, **PATCH** and **DELETE** responses are never cacheable — and a successful unsafe request invalidates stored entries for that URI. The payoff is latency, origin offload and partial failure tolerance: a cache can answer while the origin is slow or down. The cost is staleness, which is exactly why the constraint says the response must be *labelled* — the API author decides how much staleness is acceptable per resource, rather than every client guessing. Design consequence: put anything you want cached behind a GET with a stable URL.

go deeper

for a junior

State the constraint — responses must be labelled cacheable or not — and that GET/HEAD are the cacheable methods while PUT/PATCH/DELETE are not.

for a middle

Explain implicit versus explicit labelling, the invalidating effect of successful unsafe requests, and why POST-as-read kills cacheability.

for a senior

Frame the staleness budget as a per-resource product decision and separate 'cacheable' from 'shareable', with the wrong-answer failure mode in mind.

for a principal

Connect the constraint to statelessness and layering — caching is what makes a stateless, layered system affordable — and weigh latency and offload against staleness and invalidation debt.

## The constraint in the architectural style REST is defined as a set of constraints on a distributed hypermedia system: client–server, stateless, **cacheable**, uniform interface, layered system, and optional code-on-demand. The cache constraint states that responses must be implicitly or explicitly labelled cacheable or non-cacheable, so that a client — or any intermediary in the layered path — may reuse a stored response for equivalent later requests. It exists because it composes with the other constraints. Statelessness means every request carries its own context, so a response is meaningful on its own and can be stored and replayed. The layered-system constraint means the client cannot tell whether it is talking to the origin or to a cache, which is what makes a CDN transparent. Caching is the payment for statelessness: statelessness costs repetition, caching gives it back. ## What is cacheable in HTTP terms - **GET and HEAD** — the safe, read-only methods. Their responses are the normal cacheable ones, for statuses HTTP defines as cacheable by default (200, 203, 204, 206, 300, 301, 308, 404, 405, 410, 414, 501). Everything else needs explicit freshness information before it may be stored. - **POST** — technically cacheable, but only when the response carries explicit freshness information, and reuse is subtle enough that most caches simply do not do it. Treat POST as uncacheable in design. - **PUT, PATCH, DELETE** — never cacheable. Moreover, a successful unsafe request invalidates cached entries for the request URI, which is the mechanism that keeps a client's own cache honest after it writes. This is the practical reason to keep reads on GET. The moment a read is expressed as `POST /search` because the query is long, you have opted the entire endpoint out of every cache in the path — browser, corporate proxy, CDN, client library. Sometimes that is the right call; it should be a decision, not an accident. ## Labelling as an API-design act "Implicitly or explicitly labelled" is the load-bearing phrase. Implicit labelling means relying on defaults: caches applying heuristic freshness to a 200 GET with a `Last-Modified` date. Heuristics are a guess made by software you do not control, so mature APIs label explicitly — the API author is the only party who knows how stale a given resource may safely be. That per-resource staleness budget is a product decision: - a currency conversion table: minutes are fine; - a public product catalogue: seconds to minutes at the edge; - a user's account balance: no shared caching at all; - an immutable, versioned artifact: effectively forever. The corollary is that "cacheable" is not a synonym for "public". A response can be storable only by the client that received it, or storable by any shared cache. Getting that distinction wrong is how one user's data reaches another, which is why per-user data and public data usually belong on different resources rather than mixed into one response. ## What you gain and what you pay Gains: lower latency (an edge hit is milliseconds versus a cross-region origin call), origin offload (the same handful of hot resources served once and reused N times), and resilience — a cache holding a still-usable copy can keep serving while the origin is degraded. Costs: staleness, and the operational burden of dealing with it when something must change immediately — the invalidation problem, which is a design topic in its own right. There is also a correctness burden: a response that is cacheable but should not be, or a cache key missing a dimension the response varies on, produces wrong answers rather than slow ones, and those bugs are hard to reproduce because they depend on who else asked recently. ## In review When reviewing an API for this constraint, ask three questions per endpoint: is it a GET with a stable, meaningful URL; does the response state how long it may be reused; and is the data in it safe for whoever is allowed to store it? An endpoint that fails all three is not participating in the constraint at all, and no amount of infrastructure will make it cacheable later.

  • Why does expressing a read as POST /search hurt more than just losing a browser cache?
    It opts the endpoint out of every cache in the layered path — client library, corporate proxy, reverse proxy and CDN — because POST responses are not stored in practice. It also removes the shareable URL, so the query cannot be bookmarked, linked or trivially reissued, and unsafe methods invalidate rather than populate caches.
  • How does the cache constraint interact with statelessness?
    They are complements. Statelessness makes each request self-describing, which is exactly what allows a response to be stored and reused for an equivalent later request without server-side context. In return, caching offsets the repetition that statelessness imposes, which is why the two are listed as constraints of the same style.

saying these in an interview costs you the question

  • Saying every HTTP response is cacheable by default regardless of method
  • Believing POST responses are routinely cached by intermediaries
  • Treating 'cacheable' as equivalent to 'safe to store in a shared cache'
  • Assuming caching is purely infrastructure and needs no API-design input
  • Ignoring that a successful PUT or DELETE invalidates cached entries for that URI

context

open as a page

You are designing a JSON API and must decide, for each resource, whether responses carry an ETag validator, a Last-Modified date, both, or neither. How do you make that call?

level: middleimportance: must knowfreq 52%

basics

~20 s

Use an ETag when you have a cheap exact version (row version, sequence, content hash) or when changes can happen within the same second. Use Last-Modified when a meaningful modification time already exists and human readability or heuristic freshness helps. Emit both when free; emit neither for tiny or per-request-volatile responses.

open as a page

How do you choose a TTL for an HTTP API response — the number you put in Cache-Control: max-age or s-maxage — for resources that change at very different rates?

level: middleimportance: must knowfreq 55%

basics

~20 s

Pick the TTL from how stale the data may safely be, not from how often it changes. Volatile or personalised resources get seconds or no shared caching; stable reference data gets minutes to hours; immutable, uniquely-addressed content gets a year. Use s-maxage to give shared caches a different, usually longer, value.

open as a page

An API endpoint returns data that differs per authenticated user. What has to be true before any shared cache such as a CDN or reverse proxy may store that response, and how do you design so one user's data can never be served to another?

level: seniorimportance: must knowfreq 46%

basics

~20 s

A shared cache must not store an authenticated response unless the response explicitly permits shared storage. Safer than tuning directives: separate public resources from per-user ones, put the user identity in the URL path, and mark per-user responses as storable only by that user's own client.

open as a page

Reviewing a client integration, you notice every API call appends a unique query parameter such as ?_=1739382736 or ?nocache=<uuid>. What effect does that have on caches, and when is a changing URL the right tool rather than a mistake?

level: middleimportance: should knowfreq 36%

basics

~20 s

A unique parameter makes a unique cache key, so nothing is ever reused: hit rate is zero and every request reaches the origin. A changing URL is right only in the inverse case — content-addressed or versioned URLs that change when the content changes, allowing very long lifetimes.

open as a page

When would you invalidate cached API content by changing the URL — putting a version or content hash in the path — rather than relying on the origin being revalidated? What do you give up with each approach?

level: middleimportance: should knowfreq 40%

basics

~20 s

Change the URL when the representation is immutable and the client learns the new URL from somewhere else — then cache forever with no invalidation at all. Rely on revalidation when the identity must stay stable, at the cost of a round trip per check and an origin that must stay reachable.

open as a page

You are designing a JSON REST API and want conditional GET support so repeat readers can be answered with HTTP 304 Not Modified. What does the server have to implement for those 304s to actually be cheap, where do the savings land, and when would you decide the feature is not worth adding at all?

level: seniorimportance: should knowfreq 45%

basics

~20 s

A 304 pays off only if deciding "unchanged" is far cheaper than building the response — derive the validator from a stored version column or a hash saved at write time, and check it before rendering. Savings are egress bytes and client parsing; origin CPU only if you short-circuit.

open as a page

A public API has clients that poll an endpoint every minute looking for changes. How would you design that endpoint and its rate-limit policy so polls that find nothing new are cheap for both sides?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Give the endpoint a stable, cheaply computed ETag, require clients to poll with If-None-Match, answer unchanged polls with 304, and do not charge those 304s against the rate limit. Publish the polling interval, and offer webhooks or a change-feed for clients that need lower latency.

open as a page

Explain surrogate-key (cache-tag) based purging in a CDN such as Fastly or Varnish, and how it differs from purging by URL.

level: seniorimportance: should knowfreq 42%

basics

~20 s

The origin tags each response with surrogate keys naming the data it contains. The CDN indexes cached objects by tag, so one purge of a key evicts every representation containing that data — regardless of URL, query string or variant — instead of you enumerating URLs.

open as a page

You own a read-heavy public API fronted by a CDN, and the origin is saturated even though most requests ask for the same handful of things. How would you restructure the API's resources so a much larger share of traffic can be served from the edge?

level: principalimportance: should knowfreq 30%

basics

~20 s

Split personalized and volatile data out of the hot public resources, make responses byte-stable, cut cache-key cardinality by canonicalizing and constraining query parameters, and set a per-resource staleness budget. Then measure hit ratio per route and iterate on the worst offenders.

open as a page

HTTP entity tags come in a strong form like "abc" and a weak form written W/"abc". As an API designer, when would you deliberately issue a weak ETag, and what do you give up by doing so?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

A strong ETag promises byte-for-byte identical content; a weak one promises only semantic equivalence. Issue weak tags when the bytes may vary harmlessly (compression, field ordering, cosmetic fields) but the meaning does not. You give up range requests and optimistic-concurrency preconditions, which need strong comparison.

open as a page