A gateway is configured to cache GET responses and gzip-compress payloads on behalf of every backend service behind it. What has to be true of a response for gateway-level caching to be safe, and what commonly goes wrong when a team turns it on without checking?
answer
- cache key must include what varies the response
- Vary header + Cache-Control drive eligibility
- personalized data + shared cache key = leak
- compression is low-risk, size/CPU trade-off
- default deny caching for authenticated responses
basics
~20 sThe gateway can save a copy of a response and reuse it for later identical requests, saving backend work, but only if the response is truly the same for everyone who asks; if the response depends on who's asking (like personal data), caching it for all users is a serious bug.
solid answer
~50 sGateway-level response caching stores a previous response, keyed by request URL/method/headers, and serves it directly for matching future requests without hitting the backend, using cache-control directives (max-age, no-store, private) and often a vary header to decide what's cacheable and for whom. It's safe for content that's identical across users, like public product listings or reference data. It becomes dangerous the moment a response contains user-specific or session-specific data, because a naive cache key (ignoring the auth header or user ID) can serve one user's personal data to a completely different user who happens to hit the same URL, a serious privacy/security leak. Gzip/Brotli compression offloaded to the gateway is lower-risk: it just reduces payload size for compressible content types, at the cost of a small CPU overhead per request and needing to respect client Accept-Encoding negotiation; the main pitfall there is compressing already-compressed content (like images) for no benefit or accidentally breaking streaming responses.
go deeper
Should know caching reuses a stored response to save backend work, and that this only makes sense for content that's the same for everyone.
Should mention Cache-Control directives and be able to say personalized content shouldn't be cached the same way as public content.
Should explain the Vary header's role, describe the cross-user data leak failure mode concretely, and distinguish compression's low-risk profile from caching's correctness risk.
Should push for opt-in-by-default caching policy for authenticated endpoints as an organizational safeguard, and connect this to broader defense-in-depth thinking about how a single misconfigured caching rule can become a fleet-wide privacy incident.
## Two concerns, two risk profiles Response caching and compression are two of the classic cross-cutting concerns pushed into a gateway alongside TLS termination and authentication, because both are generic transformations of an HTTP response that don't require any domain knowledge from the backend that produced it. But they carry very different **risk profiles**, and conflating them is a common source of production incidents. ## How the gateway decides what is cacheable Caching at the gateway works by storing a copy of a previous response and serving it directly for subsequent matching requests, bypassing the backend entirely for cache hits. The gateway decides what's cacheable using HTTP cache-control semantics: | Cache-control semantics | Effect on the cache | |---|---| | `Cache-Control: max-age=300` | a response with it is eligible to be served from cache for 300 seconds | | `no-store` or `no-cache` | explicitly opts a response out | | `Vary` header | tells the cache which additional request headers (like `Accept-Language` or `Authorization`) must match for a cached response to be reused, rather than treating the URL alone as the cache key | The entire safety of gateway caching hinges on getting the **cache key** right: it must include everything that makes a response different between two requesters. This is trivial for genuinely public, identical-for-everyone content, like a product catalog page, a list of public blog posts, or static reference data, where caching is a clean win: backend load drops dramatically for popular endpoints, and latency for cache hits drops to whatever the gateway's own response time is, often single-digit milliseconds versus a full backend round trip. ## Where caching turns dangerous The danger appears the instant a cached response is user-specific. If a backend returns "your account balance" or "your order history" at a URL like `/api/account`, and the gateway caches that response keyed only on the URL (ignoring the `Authorization` header or session cookie that determines whose account it is), the next different user who requests the same URL, even one with a completely different identity, gets served the first user's cached response. This is not a theoretical risk; it is a well-documented class of production incident, sometimes called a **cache poisoning** or **cross-user data leak**, and it has affected major platforms in the wild, generally because a CDN or gateway cache was configured to cache by URL alone for an endpoint that should have varied by identity or been marked `no-store` entirely. The fix is disciplined: any response that varies by caller must either be - excluded from caching (`Cache-Control: private, no-store`), or - have its cache key correctly include the identifying header or token, and gateways/CDNs should default to not caching authenticated responses unless explicitly opted in, rather than the other way around. ## Compression — lower risk, different pitfalls Compression, offloaded the same way, carries much lower risk because it doesn't affect who sees what, it only affects how many bytes travel over the wire. The gateway checks the client's `Accept-Encoding` header, compresses the response body (typically gzip or the more modern Brotli) if the content type is compressible (text, JSON, HTML; not already-compressed formats like JPEG or a zip file), and sets `Content-Encoding` accordingly. This centralizes what would otherwise be redundant compression logic in every backend, and moves the CPU cost of compression to the gateway tier, which can be scaled independently. The main pitfalls are **efficiency rather than correctness**: - compressing content that's already compressed (images, pre-gzipped assets) wastes CPU for no size benefit and should be skipped by content-type - very small responses can end up larger after compression overhead than uncompressed, so gateways typically set a minimum size threshold before bothering - compression interacting badly with streaming or chunked responses can introduce buffering delays if not configured to compress incrementally ## What actually burns teams In production, teams that get burned by this pattern almost always got burned by caching, not compression, and specifically by caching an authenticated or personalized endpoint without noticing, often because a blanket "cache everything under this path prefix" rule at the gateway or CDN layer didn't distinguish between `/api/products` (safe) and `/api/products/recommended-for-you` (unsafe, personalized). A concrete real-world example of this pattern done right is a CDN like Cloudflare or Fastly sitting in front of an API, configured to cache public catalog endpoints aggressively with long max-age while explicitly setting `Cache-Control: private, no-store` on any endpoint returning account-specific or session-specific data, backed by automated tests or CDN-level rules that catch a newly added personalized endpoint that forgot to opt out of the default caching policy.
- Why is a Vary header important for safe gateway-level caching of responses that differ by client?Vary tells the cache which additional request headers, like Accept-Language or Authorization, must match between the original cached request and a new request before the cached response can be reused. Without it, the cache key might only be the URL, so requests that differ in a way that changes the response (different language, different logged-in user) would incorrectly get served an unrelated cached copy.
- Why does compression carry much lower risk than caching when both are offloaded to the gateway?Compression only changes the byte-level encoding of a response for the one client receiving it in that exact request-response cycle, it never stores or reuses data across different callers, so it can't leak one user's data to another the way a badly-keyed cache can. Its failure modes are efficiency problems, like wasting CPU compressing already-compressed content, not correctness or privacy problems.
- What's the safest default policy for whether an authenticated API response should be cacheable at the gateway?Default to not caching (Cache-Control: private, no-store) unless a team has explicitly reviewed the endpoint and confirmed the response is identical regardless of caller identity, rather than defaulting to caching everything and hoping personalized endpoints remember to opt out. Opt-in caching for authenticated content is much safer than opt-out, because a forgotten opt-out is a silent data leak, while a forgotten opt-in is just a missed performance win.
Like a copy shop that photocopies a flyer once and hands out copies to everyone who asks for it, which is efficient for a public flyer, but disastrous if it accidentally photocopies someone's personal medical form and hands that same copy to the next person who walks in asking for "the form."
saying these in an interview costs you the question
- thinks caching and compression carry the same risk profile
- doesn't mention cache key correctness (URL alone is not enough for personalized data)
- unaware that caching an authenticated endpoint incorrectly can leak one user's data to another
- assumes compression is risk-free with no content-type or size considerations
- can't explain what Cache-Control or Vary headers do