skip to content

A public API on Amazon API Gateway can be protected by stage throttles, per-key usage-plan quotas, and an AWS WAF web ACL with a rate-based rule associated with the stage. How would you decide which of these layers to use so that one abusive caller cannot degrade everyone else?

level: principalimportance: should knowfreq 35%

answer

  1. ask what identifies the abuser
  2. aggregate throttle punishes everyone equally
  3. keys only work for known callers
  4. WAF is the anonymous-traffic layer
  5. 429 from limits, 403 from WAF

basics

~20 s

Each layer limits a different thing: stage throttles cap total load but cannot tell callers apart, usage-plan quotas cap a known customer's volume, and a WAF rate-based rule caps an anonymous source by IP or another aggregation key. Anonymous abuse needs WAF; identified abuse needs usage plans.

solid answer

~50 s

Pick the layer by what identifies the abuser. A stage or method throttle is aggregate — it protects the backend from total load but treats one abusive caller and a thousand legitimate ones identically, so it fails the isolation requirement on its own. A usage plan gives each API key its own bucket and quota, which is the right tool once callers are identified, but it needs a key and therefore does not apply to anonymous traffic. An AWS WAF web ACL associated with the stage adds a rate-based rule that counts requests over a rolling five-minute window per source IP or another aggregation key and blocks with 403, which is what handles anonymous abuse, plus request inspection that rate limits cannot do. In practice you run all three: WAF for anonymous and malicious traffic, usage plans for per-customer fairness and commercial caps, and a stage throttle as the last-resort ceiling that keeps the backend alive. Note WAF associates with REST API stages, not HTTP APIs.

code

json · 16 lines
json
{
  "Name": "PerIpRateLimit",
  "Priority": 0,
  "Statement": {
    "RateBasedStatement": {
      "Limit": 2000,
      "AggregateKeyType": "IP"
    }
  },
  "Action": { "Block": {} },
  "VisibilityConfig": {
    "SampledRequestsEnabled": true,
    "CloudWatchMetricsEnabled": true,
    "MetricName": "PerIpRateLimit"
  }
}

go deeper

for a junior

Know that these are different tools: a throttle limits how fast the API is called overall, a usage plan limits a specific customer's key, and AWS WAF can block traffic based on the request itself.

for a middle

Explain what each layer uses to identify a caller and what it returns — 429 from throttles and quotas, 403 from a WAF block — and why an aggregate throttle cannot deliver per-caller fairness.

for a senior

Show the composition and its operational reality: WAF outermost for anonymous abuse, usage plans for identified callers, throttles as the backend ceiling, plus count-mode rollout and logs that let you answer why a customer saw errors.

for a principal

Own the tradeoff explicitly — the cost of each layer against the failure mode it uniquely covers, the aggregation key's false-positive risk, and how the abuse-protection requirement feeds back into the REST-versus-HTTP API choice.

## The question behind the question "One caller must not degrade everyone else" is an isolation requirement, and it immediately rules out the layer most teams reach for first. A stage throttle is a single shared bucket: when it is emptied, every caller is refused, including the ones behaving well. Aggregate throttling protects the *backend*; it does nothing for *fairness*. Choosing correctly means asking, for each layer, **what does it use to tell callers apart, and what does it do when it decides one is over the line.** ## Layer by layer **Stage and method throttles.** Identifier: none. They see requests to an API, not callers. Rejection is 429 at the front door before the integration runs, which is cheap and fast. Their role is the last-resort ceiling — the number that says "beyond this the backend cannot cope, so refuse rather than collapse". Every design should have one, and no design should rely on it for fairness. **Usage plans.** Identifier: the API key, and therefore a known customer. Each key gets its own rate/burst bucket and its own `DAY`, `WEEK` or `MONTH` quota, so a customer looping on your API exhausts only its own allocation. This is the real isolation mechanism for identified traffic, and it doubles as the commercial tiering model. Its boundary conditions: it requires a key, so it cannot see anonymous callers; enforcement is best-effort rather than exact; and usage plans exist for REST APIs only, not HTTP APIs. **AWS WAF.** Identifier: attributes of the request itself. A web ACL associated with a REST API stage can carry a rate-based rule that counts requests over a rolling five-minute window, aggregated by source IP, by a forwarded IP, or by other request-derived keys, and blocks the offender — with 403, not 429. That is the only one of the three that can act on an anonymous caller, and it is also the only one that can inspect *what* is being sent rather than just how much: bad user agents, injection patterns, geographies, managed rule groups for known-bad sources. ```json { "Name": "PerIpRateLimit", "Priority": 0, "Statement": { "RateBasedStatement": { "Limit": 2000, "AggregateKeyType": "IP" } }, "Action": { "Block": {} }, "VisibilityConfig": { "SampledRequestsEnabled": true, "CloudWatchMetricsEnabled": true, "MetricName": "PerIpRateLimit" } } ``` ## How they compose The layers are not alternatives; they sit at different distances from the backend and answer different failure modes. 1. **WAF, outermost.** Sheds traffic that should never have been served at all — scrapers, credential-stuffing bursts, obvious junk — and does it per source rather than in aggregate. 2. **Usage plans, per identified caller.** Enforces fairness and the commercial contract among legitimate customers. 3. **Stage and account throttles, innermost.** The absolute ceiling that keeps the integration alive when the outer layers are wrong. A useful test of the design: for each of "a paying customer accidentally loops", "an anonymous scraper hammers a public endpoint", and "traffic is simply larger than the backend can serve", name which layer catches it. If any of the three has no owner, there is a gap. ## The costs you must weigh None of this is free, which is why it is a judgment call rather than a checklist. WAF is priced per web ACL, per rule and per million requests, so putting it in front of a low-value internal API is often not worth it. Rate-based rules act on a rolling window with a short evaluation delay, so they stop sustained abuse rather than a single-second spike — the throttle covers that gap. Aggregating by IP misfires when many legitimate users share an address behind NAT or a corporate proxy, so the aggregation key deserves thought. Usage plans impose an onboarding and key-management burden. And note the platform constraint: WAF web ACLs associate with regional REST API stages, not with HTTP APIs, which is another input into the REST-versus-HTTP decision when abuse protection is a requirement. ## The response the caller sees Worth stating explicitly because it shows you have operated this: a throttle or quota rejection is 429, while a WAF block is 403 with no explanation. Clients therefore cannot distinguish "slow down" from "you are blocked", which matters for support and for client retry behaviour — and it is why your own runbooks need the WAF and API Gateway logs side by side to answer "why did this customer get errors".

  • Why is an aggregate stage throttle alone insufficient for the stated isolation requirement?
    Because it is one shared bucket with no notion of caller. When an abusive client empties it, the 429s land on everyone, so the abuser has successfully degraded the service for legitimate users. It protects the backend from collapse but delivers no fairness; isolation requires per-caller accounting through usage plans or a per-source WAF rule.
  • What is the risk of aggregating a WAF rate-based rule by source IP?
    Many legitimate users can share one address behind corporate NAT, a mobile carrier gateway or a proxy, so a per-IP limit blocks a whole population at once, while a distributed attacker spread over many addresses stays under it. Where a better signal exists — a header the client cannot forge, or an identity resolved upstream — aggregating on that is usually more accurate.
  • How would you validate this design before an incident rather than during one?
    Deploy WAF rules in count mode first and inspect the sampled requests and CloudWatch metrics to see who would have been blocked, then switch to block once the false-positive rate is acceptable. In parallel, load-test past the stage throttle to confirm the backend actually survives at the ceiling, and rehearse the runbook that reads WAF and API Gateway logs together.
  • When would you argue against adding WAF at all?
    When the API is not publicly reachable or the traffic is entirely authenticated and metered, the anonymous-abuse case WAF uniquely covers does not exist, and its per-ACL, per-rule and per-request pricing buys little. Usage plans plus a sane stage ceiling already deliver isolation there, and adding a layer you will not tune is cost plus a false sense of coverage.

saying these in an interview costs you the question

  • Treating a stage throttle as per-caller protection
  • Assuming WAF returns 429 like a throttle
  • Expecting usage plans to limit anonymous callers
  • Rate-limiting by IP without considering shared NAT
  • Adding WAF rules in block mode without a count phase

context