skip to content

An Amazon API Gateway REST API stage has response caching enabled with a 300-second TTL, but two problems appear: the cache hit rate is near zero, and occasionally one customer receives another customer's response. What determines what gets cached and returned, and how would you fix both symptoms?

level: seniorimportance: nice to knowfreq 38%

answer

  1. only designated parameters form the key
  2. missing input means shared entry
  3. unique input means no hits
  4. TTL does not fix a leak
  5. Cache-Control: max-age=0 needs permission

basics

~20 s

API Gateway stage caching keys entries on the method request parameters you explicitly designate as cache key parameters. Omit the parameter that varies and every caller shares one entry — the cross-customer leak. Include a parameter that is unique per request and nothing ever hits.

solid answer

~50 s

Caching is enabled per stage on a REST API, and the cache key is built only from the method request parameters you mark as cache key parameters — path parameters, query strings or headers. Nothing else is implicit, and that is the root of both symptoms. If the identifying header, such as the one carrying the tenant or user, is not in the key, every caller collides on one entry and receives whatever the first caller's response was. If something high-cardinality is in the key, such as a request ID or a full authorization token, each request produces a unique key and the hit rate collapses. The fix is to make the key exactly the set of inputs the response actually varies on. Also consider whether the endpoint should be cached at all: per-user responses usually should not be, and clients holding `execute-api:InvalidateCache` can bypass the cache with `Cache-Control: max-age=0` unless the stage is configured to reject that.

code

bash · 3 lines
bash
curl -i https://abc123.execute-api.eu-west-1.amazonaws.com/prod/reports \
  -H 'x-api-key: EXAMPLEKEY123' \
  -H 'Cache-Control: max-age=0'

go deeper

for a junior

Know that API Gateway caching is turned on per stage with a TTL, and that a cached response means the backend integration is not called at all.

for a middle

Explain that the cache key is built only from the method request parameters explicitly designated as cache key parameters, and reason about what collides or fragments when that set is wrong.

for a senior

Diagnose both directions from symptoms — a shared entry across tenants versus a key too unique to ever hit — and treat the cross-tenant case as a disclosure incident rather than a tuning problem.

for a principal

Own the placement decision: hourly provisioned regional cache versus a CDN, which responses are cacheable at all, and who is permitted to invalidate, since an open invalidation path removes the protection under exactly the load that motivated it.

## What API Gateway stage caching is On a REST API you can enable a cache **per stage**, choosing a cache capacity and a default TTL (300 seconds by default, up to 3600). Responses to method requests are stored in that cache; a subsequent matching request is answered from it without invoking the integration. It is billed by the hour for the provisioned capacity — not per request — so an unused or ineffective cache is pure cost. HTTP APIs do not have this feature at all. The intent is backend protection: a cached response is one the integration never runs, which is exactly the relief a throttle would otherwise have to provide by rejecting traffic. ## The cache key is explicit, and only explicit This is the fact both symptoms come from. API Gateway does **not** infer the cache key from the request. You designate specific **method request parameters** — path parameters, query-string parameters or headers — as cache key parameters, and the key is built from the method, the resource path and those designated values. Anything you do not designate is invisible to the cache. ### Symptom one: cross-customer responses If a response varies by caller — because the integration reads a tenant header, a user identifier, or an authorizer context — and that input is not a cache key parameter, then every caller maps to the same key. The first response is stored and served to everyone else until the TTL expires. This is not a subtle bug; it is a data-disclosure incident, and it is the standard failure mode of switching caching on without auditing what the responses depend on. There are two correct fixes and one wrong one: - Add the identifying parameter to the cache key, so each tenant gets its own entry. Note that this multiplies the number of entries by the number of tenants, which changes the capacity you need. - Do not cache the endpoint at all. Per-user, per-request data is often simply not cacheable, and disabling caching on that method while keeping it on genuinely shared ones is frequently the right call. - The wrong fix is shortening the TTL. A 5-second TTL still leaks; it just leaks less often, which makes the incident harder to reproduce. ### Symptom two: no hits The mirror image. If a cache key parameter carries something unique per request — a request ID, a timestamp, a cache-busting query string a client appends, or a bearer token that rotates — every request computes a new key, stores a new entry, and hits nothing. You pay for the cache, evict constantly, and the integration still runs on every call. The diagnosis is to compare the key inputs against what genuinely changes the response body. High-cardinality inputs that do not change the response must come out of the key. A token that identifies the caller but whose value rotates is the classic offender: it varies far more often than the response does, so a stable tenant identifier is the better key input. ## Invalidation, and who is allowed to do it A client can ask API Gateway to bypass and refresh a cached entry by sending `Cache-Control: max-age=0`. That capability is gated by the `execute-api:InvalidateCache` IAM permission, and the stage has a "require authorization" setting controlling what happens when an unauthorized caller sends the header — the request can be failed with 403, or the header ignored (optionally with a warning header). Leaving invalidation open to unauthenticated callers hands them a switch that turns your cache off under load, which is precisely when you need it. You can also flush the whole stage cache administratively, which is the blunt instrument after a bad deploy. ## Deciding whether to cache here at all An honest senior answer usually widens the question. API Gateway's stage cache is provisioned capacity billed by the hour, sits per stage, and offers a fairly coarse key model. For public, cacheable GET traffic a CDN in front of the API often does the job better and cheaper. The gateway cache earns its place when the traffic is regional and authenticated, when the origin is expensive per call, and when the response genuinely varies on a small, low-cardinality set of inputs you can enumerate. If you cannot enumerate those inputs confidently, that uncertainty is itself the argument against enabling it.

  • Why is shortening the TTL an unacceptable fix for the cross-customer leak?
    Because the key is still wrong. Any TTL above zero means some requests are answered from another caller's entry, so a shorter window reduces the frequency of the disclosure without removing it — and makes it far harder to reproduce and detect. The key must include the varying input, or the method must not be cached.
  • How do you stop clients from bypassing the cache whenever they feel like it?
    Cache invalidation via the `Cache-Control: max-age=0` header is gated by the `execute-api:InvalidateCache` permission, and the stage can be set to require authorization for it — failing unauthorized attempts with 403 rather than honouring them. Without that, any caller can force a miss on every request and defeat the protection the cache was providing.
  • When would you put a CDN in front instead of using the stage cache?
    When the cacheable traffic is public, globally distributed and keyed on the URL, a CDN gives edge locality and per-request pricing rather than an hourly provisioned cache in one region. The gateway cache is a better fit for authenticated, regional traffic whose response varies on a small set of enumerable inputs and whose origin is expensive per call.

saying these in an interview costs you the question

  • Assuming the whole request forms the cache key
  • Fixing a cross-tenant leak by lowering the TTL
  • Putting a rotating token in the cache key
  • Thinking the stage cache is billed per request
  • Leaving client-driven invalidation open to anyone

context