skip to content

A client calling an Amazon API Gateway REST API starts receiving HTTP 429 responses. Which two different API Gateway conditions produce a 429, how do you tell them apart, and how should the client react to each?

level: juniorimportance: should knowfreq 68%

answer

  1. two different causes, one status code
  2. one clears in a second
  3. the other waits for the period
  4. THROTTLED versus QUOTA_EXCEEDED gateway responses
  5. backoff needs jitter, not a loop

basics

~20 s

API Gateway returns 429 for two reasons: a rate or burst throttle was hit, or a usage-plan quota is exhausted. A throttle clears within seconds, so retry with exponential backoff and jitter; a quota only resets at its day, week or month boundary.

solid answer

~50 s

API Gateway has two independent limiters that both answer with `429 Too Many Requests`. The first is throttling — a token bucket with a steady-state rate and a burst allowance, applied at the account, stage, method or usage-plan level. The second is a usage-plan quota, a fixed number of requests an API key may make per day, week or month. They are distinguishable: throttling emits the `THROTTLED` gateway response (body `Too Many Requests`), while quota exhaustion emits `QUOTA_EXCEEDED` (body `Limit Exceeded`), and both response types can be customised per API. The reaction differs completely. A throttle is transient — the bucket refills within a second — so the client should retry with exponential backoff plus jitter, never in a tight loop. A quota is not transient: retrying is pointless until the period rolls over, so the client should stop, surface the condition, and the owner should move that key to a larger plan.

go deeper

for a junior

Know that 429 means the caller is being limited, not that the service is broken, and that the correct client behaviour is a delayed retry rather than an immediate one.

for a middle

Be ready to explain that throttling is a token bucket that refills in under a second while a quota is a counter tied to a day, week or month, and to name the two gateway response types that distinguish them.

for a senior

Show how you would make the difference visible in production: custom gateway response bodies, access-log fields that separate the causes, and a client retry policy with capped exponential backoff and jitter.

for a principal

Own the contract question — what your API promises callers when it refuses them, how limit exhaustion is communicated before it becomes failure, and whether the quota is a commercial lever or a stability control.

## Why one status code covers two situations HTTP has a single status for "you are sending too much": `429 Too Many Requests`. Amazon API Gateway uses it for two mechanisms that behave nothing alike, and conflating them is the most common reason a client's retry logic makes an outage worse instead of better. ## Mechanism one: throttling (a token bucket) Throttling limits the *speed* of requests. Every throttle in API Gateway is a pair of numbers: a **rate** (steady-state requests per second) and a **burst** (how many requests may arrive at once). Internally this is a token bucket — the burst value is the bucket's capacity, the rate value is how fast tokens are put back. Each request removes one token; a request that finds the bucket empty is rejected with 429 without ever reaching your integration. Because tokens refill continuously, a throttle-induced 429 is over almost immediately: the very next second there is capacity again. Throttles exist at several levels — the account limit for the region, the stage's default, a per-method or per-route override, and per-API-key limits inside a usage plan — and the request is rejected by whichever bucket empties first. ## Mechanism two: a usage-plan quota (a counter) A quota is a *volume* limit, not a speed limit: an API key attached to a usage plan may make, say, 1,000,000 requests per `MONTH` (the supported periods are `DAY`, `WEEK` and `MONTH`). This is a counter that only resets when the period rolls over. A caller can be nowhere near any rate limit and still be refused, because it has simply used up its allowance. Waiting a few seconds changes nothing; waiting until the first of the month does. ## Telling the two apart API Gateway models them as two different **gateway responses**, which is the customisation point you would use to make the difference obvious to callers: - `THROTTLED` — default body `{"message": "Too Many Requests"}` - `QUOTA_EXCEEDED` — default body `{"message": "Limit Exceeded"}` Both default to status 429. Because you can override the status code, body and headers of each gateway response per API, a good API design makes them self-describing — for example adding a machine-readable error code, and for the quota case a hint about when the period resets. Do not assume a `Retry-After` header is present: API Gateway does not add one by default, so a client must supply its own backoff policy. From the server side, the distinction is also visible in access logs (`$context.status` with the response body) and in CloudWatch, where both show up in the API's `4XXError` metric — so a 429 spike looks like a client-error spike unless you log enough context to separate them. ## What the client should do For a throttle: retry, but politely. Exponential backoff with **jitter** is the standard answer, because synchronised retries from many clients re-create the same burst that caused the throttle (the thundering-herd problem). Cap the number of attempts and the total delay. Most AWS SDKs already implement this for AWS API calls, but a plain HTTP client calling *your* API does not — you own that logic. ```bash # naive: hammers the same empty bucket until curl -sf https://api.example.com/v1/orders; do :; done # better: back off with jitter between attempts for attempt in 1 2 3 4 5; do curl -sf https://api.example.com/v1/orders && break sleep $(( (2 ** attempt) + (RANDOM % 3) )) done ``` For a quota: do not retry at all. Fail the operation, log it distinctly, and alert a human — the fix is commercial (a bigger plan) or behavioural (the client is looping), not technical. ## What the API owner should do Treat 429s as signal, not noise. A steady trickle of throttle 429s from one key usually means that customer needs a higher tier or is missing client-side batching. A sudden broad wave across all callers usually means a shared limit — the stage default or the account limit for the region — is the binding constraint, and raising a per-method number will not help. Quota 429s are a billing and communication event: customers should learn they are near the limit before they hit it, from usage reporting, not from a wall of failures.

  • Why does adding jitter to the retry delay matter more than the backoff curve itself?
    Because plain exponential backoff keeps clients synchronised. If a thousand callers are throttled in the same second and all wait exactly two seconds, they re-arrive together and empty the bucket again. Random jitter spreads the retries across the window, so capacity is consumed smoothly instead of in a repeating spike.
  • Both conditions land in the same CloudWatch metric. How would you separate them operationally?
    Enable stage access logging and capture `$context.status` along with identifying fields such as the API key ID and resource path, then query the logs to split 429s by response body or by key. A metric filter or Logs Insights query over that field gives you a per-cause count without needing a metric API Gateway does not emit.
  • A caller says it gets 429 even though it sends only a handful of requests per second. What would you check?
    Whether the limit it is hitting is not its own. A stage or account throttle is shared across every caller of that API and every API in the region, so a noisy neighbour can throttle a quiet client. Check whether the 429s correlate with total traffic rather than that client's traffic, and whether its usage-plan quota is simply exhausted.

saying these in an interview costs you the question

  • Assuming a 429 always means the backend is down
  • Retrying a quota 429 in a tight loop
  • Believing a quota resets after a short wait
  • Expecting API Gateway to send Retry-After by default
  • Treating 429 as a client bug and raising limits blindly

context