skip to content

REST lists 'cacheable' as one of its architectural constraints. What does that constraint actually require of an API's responses, and which responses in a typical HTTP API can be cached at all?

level: juniorimportance: must knowfreq 48%

answer

  1. every response labelled cacheable or not
  2. GET/HEAD cacheable; PUT/PATCH/DELETE never
  3. POST only with explicit freshness — in practice, no
  4. caching pays back the cost of statelessness
  5. cacheable is not the same as public

basics

~20 s

It requires every response to say, implicitly or explicitly, whether it may be reused and for how long. In practice GET and HEAD responses are the cacheable ones; POST responses only with explicit freshness; PUT, PATCH and DELETE responses are never cached.

solid answer

~60 s

The constraint is about **labelling**: each response must declare whether a client or an intermediary may store and reuse it, and for how long. That label is what lets caches exist between client and origin without either side coordinating. In practice: **GET** and **HEAD** responses are cacheable by default for the status codes HTTP defines that way (200, 203, 204, 206, 300, 301, 308, 404, 405, 410, 414, 501). **POST** responses are cacheable only if you explicitly give them freshness information, and almost nothing does. **PUT**, **PATCH** and **DELETE** responses are never cacheable — and a successful unsafe request invalidates stored entries for that URI. The payoff is latency, origin offload and partial failure tolerance: a cache can answer while the origin is slow or down. The cost is staleness, which is exactly why the constraint says the response must be *labelled* — the API author decides how much staleness is acceptable per resource, rather than every client guessing. Design consequence: put anything you want cached behind a GET with a stable URL.

go deeper

for a junior

State the constraint — responses must be labelled cacheable or not — and that GET/HEAD are the cacheable methods while PUT/PATCH/DELETE are not.

for a middle

Explain implicit versus explicit labelling, the invalidating effect of successful unsafe requests, and why POST-as-read kills cacheability.

for a senior

Frame the staleness budget as a per-resource product decision and separate 'cacheable' from 'shareable', with the wrong-answer failure mode in mind.

for a principal

Connect the constraint to statelessness and layering — caching is what makes a stateless, layered system affordable — and weigh latency and offload against staleness and invalidation debt.

## The constraint in the architectural style REST is defined as a set of constraints on a distributed hypermedia system: client–server, stateless, **cacheable**, uniform interface, layered system, and optional code-on-demand. The cache constraint states that responses must be implicitly or explicitly labelled cacheable or non-cacheable, so that a client — or any intermediary in the layered path — may reuse a stored response for equivalent later requests. It exists because it composes with the other constraints. Statelessness means every request carries its own context, so a response is meaningful on its own and can be stored and replayed. The layered-system constraint means the client cannot tell whether it is talking to the origin or to a cache, which is what makes a CDN transparent. Caching is the payment for statelessness: statelessness costs repetition, caching gives it back. ## What is cacheable in HTTP terms - **GET and HEAD** — the safe, read-only methods. Their responses are the normal cacheable ones, for statuses HTTP defines as cacheable by default (200, 203, 204, 206, 300, 301, 308, 404, 405, 410, 414, 501). Everything else needs explicit freshness information before it may be stored. - **POST** — technically cacheable, but only when the response carries explicit freshness information, and reuse is subtle enough that most caches simply do not do it. Treat POST as uncacheable in design. - **PUT, PATCH, DELETE** — never cacheable. Moreover, a successful unsafe request invalidates cached entries for the request URI, which is the mechanism that keeps a client's own cache honest after it writes. This is the practical reason to keep reads on GET. The moment a read is expressed as `POST /search` because the query is long, you have opted the entire endpoint out of every cache in the path — browser, corporate proxy, CDN, client library. Sometimes that is the right call; it should be a decision, not an accident. ## Labelling as an API-design act "Implicitly or explicitly labelled" is the load-bearing phrase. Implicit labelling means relying on defaults: caches applying heuristic freshness to a 200 GET with a `Last-Modified` date. Heuristics are a guess made by software you do not control, so mature APIs label explicitly — the API author is the only party who knows how stale a given resource may safely be. That per-resource staleness budget is a product decision: - a currency conversion table: minutes are fine; - a public product catalogue: seconds to minutes at the edge; - a user's account balance: no shared caching at all; - an immutable, versioned artifact: effectively forever. The corollary is that "cacheable" is not a synonym for "public". A response can be storable only by the client that received it, or storable by any shared cache. Getting that distinction wrong is how one user's data reaches another, which is why per-user data and public data usually belong on different resources rather than mixed into one response. ## What you gain and what you pay Gains: lower latency (an edge hit is milliseconds versus a cross-region origin call), origin offload (the same handful of hot resources served once and reused N times), and resilience — a cache holding a still-usable copy can keep serving while the origin is degraded. Costs: staleness, and the operational burden of dealing with it when something must change immediately — the invalidation problem, which is a design topic in its own right. There is also a correctness burden: a response that is cacheable but should not be, or a cache key missing a dimension the response varies on, produces wrong answers rather than slow ones, and those bugs are hard to reproduce because they depend on who else asked recently. ## In review When reviewing an API for this constraint, ask three questions per endpoint: is it a GET with a stable, meaningful URL; does the response state how long it may be reused; and is the data in it safe for whoever is allowed to store it? An endpoint that fails all three is not participating in the constraint at all, and no amount of infrastructure will make it cacheable later.

  • Why does expressing a read as POST /search hurt more than just losing a browser cache?
    It opts the endpoint out of every cache in the layered path — client library, corporate proxy, reverse proxy and CDN — because POST responses are not stored in practice. It also removes the shareable URL, so the query cannot be bookmarked, linked or trivially reissued, and unsafe methods invalidate rather than populate caches.
  • How does the cache constraint interact with statelessness?
    They are complements. Statelessness makes each request self-describing, which is exactly what allows a response to be stored and reused for an equivalent later request without server-side context. In return, caching offsets the repetition that statelessness imposes, which is why the two are listed as constraints of the same style.

saying these in an interview costs you the question

  • Saying every HTTP response is cacheable by default regardless of method
  • Believing POST responses are routinely cached by intermediaries
  • Treating 'cacheable' as equivalent to 'safe to store in a shared cache'
  • Assuming caching is purely infrastructure and needs no API-design input
  • Ignoring that a successful PUT or DELETE invalidates cached entries for that URI

context