A CDN caching static assets is configured to include a custom request header in its cache key, and shortly after, some users start receiving other users' cached error pages or unexpected content for a shared static URL. What kind of failure is this, and how does it happen mechanically?
answer
- cache key vs actual response variance mismatch
- unkeyed header poisons shared cache entry
- cache 2xx only, skip error responses
- public bucket + CDN = broad exposure
basics
~20 sThis is cache poisoning: the CDN stored a wrong or attacker-influenced response under a cache key that many different users then hit, so instead of one broken response for one weird request, everyone gets served the bad cached copy.
solid answer
~50 sThis is a cache poisoning / cache key misconfiguration failure. It happens when the CDN's cache key doesn't fully account for something that changes the response — e.g. it caches by URL alone but the origin varies its response based on a header that isn't part of the key, so a single crafted or unusual request gets its response cached at a shared edge node and then served to every subsequent unrelated user hitting that same cached key. It's dangerous with static content hosting specifically because the entire point is maximum reuse across users — a single poisoned entry can be served broadly and persist until TTL expiry or purge. The fix is ensuring Vary headers and cache-key configuration exactly match what the origin actually varies on, and caching only successful responses so an error doesn't get amplified to everyone.
go deeper
Should recognize that a CDN serves the same cached thing to many people, so a wrong cached copy affects many users, not just one.
Should connect the bug to a mismatch between cache key and what actually varies the response, even without using the term 'unkeyed input'.
Should articulate the cache-key-granularity trade-off, propose caching only 2xx and correct Vary configuration as concrete fixes, and reason about blast radius.
Should discuss this as a systemic risk of the pattern's core value proposition (broad reuse = broad blast radius), design origin/CDN configuration defensively from the start, and account for full remediation (purge + origin fix + bucket policy audit) after an incident, not just symptom suppression.
## The mismatch that causes it Cache poisoning in a CDN-fronted static hosting setup is a mismatch between what actually determines a response's content and what the CDN uses as its **cache key**. A cache key is the identifier the CDN uses to decide whether two requests should get the same cached response — normally just the URL path (and sometimes select query parameters). The origin, meanwhile, may (correctly or by misconfiguration) produce different responses for requests that share that same cache key but differ in something the key ignores — most commonly a request header. - If the origin's response genuinely depends on, say, a `Host` header, a locale header, or a debug header, but the CDN's cache key is just the path, then the first request to reach a given edge PoP **wins**: whatever response the origin returned for that particular request's headers gets cached under the shared key. - Every subsequent request to that same path at that PoP — regardless of its own headers — receives the cached copy from the first request. - If an attacker can influence what that first cached response looks like, that attacker-shaped response becomes what every other user sees for a completely normal request to that URL, until the cache entry expires or is purged. ## Why static hosting is exposed to it This exists as a risk specifically because of what makes static hosting valuable in the first place: the entire design goal is that one cached response serves many unrelated users' requests. That's a feature under normal operation and becomes the exact mechanism of harm once the wrong response gets cached, because the blast radius of one bad cache write is every subsequent user hitting that same key, not just the one attacker. ## The trade-off: granularity against hit ratio The trade-off underlying this class of bug is between cache-key granularity and hit ratio. | Cache key | What you get | |---|---| | includes every header the origin might theoretically vary on | maximally safe but would fragment the cache into near-uselessness — you'd essentially never get a hit, defeating the purpose of the CDN | | too coarse (URL-only) | maximally cache-efficient but unsafe if the origin's response ever depends on something outside that key | The correct fix is precision, not maximalism in either direction: the CDN's `Vary` handling and cache-key configuration should include exactly the request dimensions the origin's response actually depends on — no more, no less — and, ideally, the origin for pure static assets shouldn't vary its response on arbitrary headers at all, since static assets are the one category of content where 'same URL, same bytes, for everyone' should hold by design. ## Failure modes in production Failure modes in production tend to surface in a few recognizable shapes. 1. **The classic Host-header cache poisoning pattern**, well documented across various CDN/reverse-proxy setups historically: an origin that reflects or uses the `Host` header in some part of its response combined with a CDN that caches by path only, letting an attacker send a request with a manipulated `Host` header that gets baked into a cached response served to everyone. 2. **Caching an error response.** Another is if an edge case — a malformed request, an unusual `Accept` header, an internal server error — produces an error page and that error page gets cached under the same key as the normal successful response, because the CDN wasn't configured to skip caching non-2xx responses, subsequent normal users suddenly get served the error page instead of the real asset, looking like a broad outage traceable to one bad request. 3. **A misconfigured object storage bucket policy** — a related but distinct failure mode. If static assets are hosted in a bucket with overly permissive public-read settings, or if non-public files are accidentally placed in the same publicly-readable prefix as legitimately public assets, those files become world-readable through the CDN just like any other static asset, since the CDN faithfully caches and serves whatever origin returns for a given path. ## The defense pattern A well-known real-world class of incident matching this pattern is web cache poisoning research documented by security researchers against various CDN configurations, where unkeyed inputs — headers, cookies, or response headers the CDN didn't consider part of the cache key — were used to get malicious content cached and served broadly. The practical defense pattern for static content hosting specifically is to keep static asset responses header-independent wherever possible, explicitly configure the CDN to cache only successful (2xx) responses, and treat any `Vary` usage as something to justify and audit rather than something incidental.
- Why is caching only 2xx responses (and not, say, 404s or 500s) an important CDN configuration for static hosting?If a transient origin error or a request for a not-yet-uploaded file gets cached as if it were a normal response, every subsequent request for that same path — even after the real file exists or the transient error clears — keeps getting served the cached error until TTL expiry or a purge. Restricting caching to successful responses, or giving error responses a very short TTL, avoids turning a brief origin hiccup into an extended outage for all users of that URL.
- How does a Vary header relate to this failure mode?Vary tells caches which request headers the response actually depends on, so a correctly configured cache will key on the URL plus those specific header values instead of collapsing all variants into one shared entry. Missing or incorrect Vary configuration is exactly what lets an origin's header-dependent behavior produce a cache-poisoning-shaped bug, since the CDN then can't tell it needs to keep responses separate.
- Why is a public object storage bucket misconfiguration especially dangerous when it's fronted by a CDN, compared to the bucket being exposed directly?The CDN caches and widely distributes whatever it's given, so a leaked file isn't just reachable once from origin — it gets replicated across edge caches globally and can keep being served from cache even after the bucket policy is fixed and the file removed from origin, until every cached copy's TTL expires or an explicit purge is issued everywhere. The CDN turns a single point of exposure into a much harder-to-fully-remediate one.
Like a photocopier that makes one copy for whoever asks first and then hands that exact same copy to everyone else who asks for 'the same document' afterward — if the first request was a forged or broken version, that's what everyone gets until someone notices and reprints.
saying these in an interview costs you the question
- Thinks cache poisoning only means DNS cache poisoning and misapplies that concept here
- Suggests maximal cache-key granularity (include every header) as a fix without noting the hit-ratio cost
- Doesn't mention that the fix is aligning Vary/cache-key config to actual origin response variance
- Assumes purging the CDN alone fully remediates a leaked-file incident without mentioning removing it from origin/bucket policy too