skip to content

Throttling, Quotas & Caching

Protecting the backend behind the front door: account, stage, and per-route throttles, usage plans that give each API key a quota, and optional response caching. A strong answer names the 429 the client sees and where the token bucket actually lives.

part ofAWSoverview, primer and where to startread it →
on this pageshow

questions

5

Amazon API Gateway lets you set throttling limits at the account, stage, method or route, and usage-plan levels. Explain what the rate and burst values in a throttle pair actually control, and which of those layers decides whether a given request is rejected.

level: middleimportance: must knowfreq 60%

answer

  1. burst is capacity, rate is refill
  2. most specific bucket checked first
  3. first empty bucket rejects
  4. account limit is region-wide, shared
  5. ceiling, not a reservation

basics

~20 s

Each API Gateway throttle is a token bucket: burst is the bucket's capacity, rate is how many tokens per second refill it. A request is rejected by the first empty bucket it meets, checked from the most specific usage-plan limit outward to the region-wide account limit.

solid answer

~50 s

Every throttle in API Gateway is a rate/burst pair implementing a token bucket. `burst` is the bucket size — how many requests can arrive simultaneously before rejection — and `rate` is the steady-state refill in requests per second. For a REST API the limits are evaluated from most specific to least: per-client per-method limits in the usage plan, then the client's overall usage-plan limits, then the method-level override on the stage, then the stage default, then the account-level limit for the region. The first bucket with no token wins and the caller gets 429; the integration is never invoked. Two consequences matter. The account limit is shared by every API in that account and region, so one API can throttle another. And these are ceilings, not reservations — setting a per-method limit caps that method, it does not guarantee it capacity when a shared bucket upstream is already empty.

code

bash · 6 lines
bash
aws apigateway update-stage \
  --rest-api-id abc123 \
  --stage-name prod \
  --patch-operations \
    op=replace,path=/*/*/throttling/rateLimit,value=200 \
    op=replace,path=/*/*/throttling/burstLimit,value=400

go deeper

for a junior

Recall that a throttle has two numbers — a per-second rate and a burst allowance — and that exceeding either returns 429 without the backend ever being called.

for a middle

Explain the token bucket precisely: burst as capacity, rate as refill, and the evaluation order from usage-plan limits outward to the account limit for the region.

for a senior

Demonstrate sizing judgment — derive the rate from what the backend can sustain, size burst for real traffic shape, and diagnose which shared bucket is binding when a per-method limit clearly is not.

for a principal

Own the isolation argument: limits are ceilings, so tenant or workload isolation comes from account, API and usage-plan boundaries, and you should be able to justify that structure against its operational cost.

## The unit of throttling: a rate/burst pair Wherever API Gateway lets you throttle, it asks for the same two numbers, and they are not two ways of saying the same thing. - **rate** — the sustainable requests per second. This is the refill speed of a token bucket. - **burst** — the bucket's capacity: the number of requests that may be served in a very short spike before the bucket is empty. Each arriving request takes one token. If a token is available, the request proceeds to the integration. If not, API Gateway rejects it at the front door with `429 Too Many Requests` (the `THROTTLED` gateway response), and you are not billed for a backend invocation because there was none. This is why a client can be under the rate limit on a per-minute average and still be throttled: it sent its whole minute's traffic in one tenth of a second, exceeding burst. Conversely a client that paces itself will happily sustain the rate indefinitely without ever emptying the bucket. ## Where the buckets live For a REST API there are four places a throttle can be configured, and they nest: 1. **Account level, per region.** As of 2025 the default is 10,000 requests per second steady-state with a 5,000-request burst, shared across all APIs in that account and region, and adjustable through a quota-increase request. 2. **Stage level.** A default rate/burst for every method on that stage. 3. **Method level.** An override for one method on that stage, useful for protecting an expensive or fragile route. 4. **Usage plan.** Per-API-key limits — both an overall rate/burst for the key and optional per-method overrides for that key. HTTP APIs have the same idea with fewer layers: a default route throttle per stage plus per-route overrides, under the same account limit. They do not have usage plans. ## Which layer decides API Gateway evaluates them from the most specific to the least specific: - the per-client per-method limit in the usage plan, then - the client's overall usage-plan limit, then - the method-level limit on the stage, then - the stage default, then - the account-level limit for the region. The **first empty bucket rejects the request**. That ordering is the whole answer to "which limit applies": not the smallest number, not the last one configured — the first one, walking outward from the caller's own allocation, that has no token left. ```bash # Set a default method-level throttle on every method of a stage aws apigateway update-stage \ --rest-api-id abc123 --stage-name prod \ --patch-operations \ op=replace,path=/*/*/throttling/rateLimit,value=200 \ op=replace,path=/*/*/throttling/burstLimit,value=400 ``` ## The two consequences interviewers are listening for **Shared limits mean shared blast radius.** The account limit is per account and region, not per API. A batch job hammering an internal API can consume the region's tokens and throttle an unrelated customer-facing API in the same account. That is a strong argument for separating high-volume or untrusted workloads into their own account, and for not treating a raised account quota as an alternative to per-caller limits. **Limits are ceilings, not reservations.** Configuring a 500 rps method limit does not reserve 500 rps for that method. It only means that method will be cut off above 500. If the stage default or the account bucket is already empty, the request is refused before the method limit is ever consulted. If you genuinely need isolation between tenants or between critical and bulk traffic, the mechanisms are separate APIs, separate stages, per-key usage plans, or separate accounts — not a bigger number on one method. ## Sizing them Start from what the backend can actually absorb, not from what the client would like. The throttle exists so that the front door fails cheaply instead of letting the integration collapse expensively. Choose the rate at or slightly below the backend's sustainable capacity, then set burst high enough to absorb normal spikiness — traffic is rarely smooth — but low enough that a spike cannot outrun the backend's ability to recover. Watch the API's `4XXError` metric and access logs after any change: a rising 429 share against flat traffic means the limit is now the binding constraint.

  • A method has a 500 rps limit but callers are throttled at around 100 rps. What is going on?
    A broader bucket is emptying first. The stage default or the account limit for the region is being consumed by other methods or other APIs in the same account, so requests are rejected before the method-level limit is reached. Raising the method number changes nothing; you need to find which shared limit is binding, using total traffic rather than that method's traffic.
  • How would you give one important tenant capacity that noisy tenants cannot consume?
    Not with a bigger method limit, since limits are ceilings rather than reservations. Real isolation comes from separating the traffic: a dedicated usage plan and API key with its own rate and burst, a separate stage or API, or in the strongest case a separate AWS account so the region-wide account throttle is not shared at all.
  • Why can a client that averages well under the rate limit still get throttled?
    Because the burst value governs simultaneity, not the average. A client that sends its whole second's worth of traffic in a few milliseconds empties the bucket before it can refill, and the excess is rejected even though the one-second average is inside the rate. The fix is client-side pacing or a larger burst allowance.

saying these in an interview costs you the question

  • Thinking rate and burst are two names for one setting
  • Believing the smallest configured limit always applies
  • Assuming a per-method limit reserves capacity
  • Forgetting the account limit is shared across all APIs in the region
  • Claiming throttled requests still reach the backend

context

open as a page

You are launching a public REST API on Amazon API Gateway with a free tier and a paid tier. How do usage plans give each customer its own rate limit and monthly allowance, and what are the practical limits of that mechanism?

level: middleimportance: must knowfreq 52%

basics

~20 s

An API Gateway usage plan binds a rate/burst throttle and a request quota per day, week or month to API keys, and associates them with specific API stages. Each customer gets its own key, so tiers are just plans with different numbers. Usage plans exist for REST APIs only.

open as a page

A client calling an Amazon API Gateway REST API starts receiving HTTP 429 responses. Which two different API Gateway conditions produce a 429, how do you tell them apart, and how should the client react to each?

level: juniorimportance: should knowfreq 68%

basics

~20 s

API Gateway returns 429 for two reasons: a rate or burst throttle was hit, or a usage-plan quota is exhausted. A throttle clears within seconds, so retry with exponential backoff and jitter; a quota only resets at its day, week or month boundary.

open as a page

A public API on Amazon API Gateway can be protected by stage throttles, per-key usage-plan quotas, and an AWS WAF web ACL with a rate-based rule associated with the stage. How would you decide which of these layers to use so that one abusive caller cannot degrade everyone else?

level: principalimportance: should knowfreq 35%

basics

~20 s

Each layer limits a different thing: stage throttles cap total load but cannot tell callers apart, usage-plan quotas cap a known customer's volume, and a WAF rate-based rule caps an anonymous source by IP or another aggregation key. Anonymous abuse needs WAF; identified abuse needs usage plans.

open as a page

An Amazon API Gateway REST API stage has response caching enabled with a 300-second TTL, but two problems appear: the cache hit rate is near zero, and occasionally one customer receives another customer's response. What determines what gets cached and returned, and how would you fix both symptoms?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

API Gateway stage caching keys entries on the method request parameters you explicitly designate as cache key parameters. Omit the parameter that varies and every caller shares one entry — the cross-customer leak. Include a parameter that is unique per request and nothing ever hits.

open as a page