You are launching a public REST API on Amazon API Gateway with a free tier and a paid tier. How do usage plans give each customer its own rate limit and monthly allowance, and what are the practical limits of that mechanism?
answer
- per-key bucket and per-key counter
- quota periods are day, week, month
- method must require the key
- key identifies, does not authenticate
- REST APIs only, not HTTP APIs
basics
~20 sAn API Gateway usage plan binds a rate/burst throttle and a request quota per day, week or month to API keys, and associates them with specific API stages. Each customer gets its own key, so tiers are just plans with different numbers. Usage plans exist for REST APIs only.
solid answer
~50 sA usage plan is the metering object in API Gateway: it carries a throttle (rate and burst), an optional quota of N requests per `DAY`, `WEEK` or `MONTH`, and a list of API stages it applies to. You attach API keys to it — one key per customer — and each key gets its own bucket and its own counter, so a free plan and a gold plan are the same construct with different numbers. Methods must be marked as requiring an API key, and callers send it in the `x-api-key` header. The caveats matter as much as the mechanism. Usage plans and API keys are a REST API feature; HTTP APIs do not have them. An API key is an identifier for metering, not a credential — you still need a real authorizer for authentication. Quota and throttle enforcement is best-effort in a distributed system, so a quota is a commercial guardrail, not an exact billing meter.
code
bash · 5 linesaws apigateway create-usage-plan \
--name Gold \
--throttle rateLimit=100,burstLimit=200 \
--quota limit=1000000,period=MONTH \
--api-stages apiId=abc123,stage=prodgo deeper
Know that customers are identified by an API key sent in the x-api-key header, and that the usage plan attached to that key is what carries its rate limit and request quota.
Explain the wiring end to end — plan holds throttle, quota period and stages; key attaches to plan; method must require a key — and note that HTTP APIs have none of this.
Bring the operational caveats: keys are identifiers not credentials, enforcement is best-effort so it cannot be a billing source, and a hard monthly cut-off needs a notification and upgrade path around it.
Frame it as a product boundary — what tiering the API promises commercially, whether metering belongs in the gateway or your own layer, and how key issuance, rotation and revocation are operated at customer scale.
## What a usage plan actually is Stage and method throttles limit *the API*. They cannot tell one caller from another, so a single aggressive customer consumes the shared bucket and everyone else is throttled. A **usage plan** is API Gateway's answer to that: a named object holding - a **throttle** — rate and burst, applied per API key rather than per API, - an optional **quota** — a request count per `DAY`, `WEEK` or `MONTH`, - the set of **API stages** the plan applies to, with optional per-method throttle overrides. You then attach **API keys** to the plan. Each key gets its own token bucket and its own quota counter. Tiering is therefore trivial: create `Free` and `Gold` plans with different numbers and move a customer's key between them. ```bash aws apigateway create-usage-plan \ --name Gold \ --throttle rateLimit=100,burstLimit=200 \ --quota limit=1000000,period=MONTH \ --api-stages apiId=abc123,stage=prod ``` ## Wiring it up Three things must line up or the plan silently does nothing: 1. The **method must require an API key**. If it does not, requests arrive anonymously, no key is identified, and only the stage and account throttles apply. 2. The **caller must send the key**, by default in the `x-api-key` header. The API key source can instead be an authorizer, which is how you attach metering to an identity your authorizer resolved rather than to a header the client controls. 3. The **key must be attached to a plan** that includes the stage being called. A key with no plan is accepted as a key but metered by nothing. When the key's bucket is empty the caller gets the `THROTTLED` gateway response; when its quota is spent it gets `QUOTA_EXCEEDED`. Both are 429 by default. ## The limits of the mechanism **REST APIs only.** HTTP APIs — the cheaper, lower-latency flavour — support route and stage throttling but have no API keys and no usage plans. If per-customer quotas are a product requirement, that requirement alone can decide REST over HTTP, and it is a decision worth stating explicitly in a design discussion because the cost difference is real. **A key is not authentication.** An API key is a plain string identifying a caller for metering. It is not signed, not scoped, and confers no permissions by itself. Anyone holding it can use it. Authentication and authorization are a separate concern — the key answers "whose quota do I decrement", not "is this caller allowed here". **Enforcement is best-effort.** Throttle and quota accounting is distributed, so a caller may occasionally slip slightly past a boundary. That is fine for protecting a backend and for a commercial guardrail; it is not an exact meter. If you bill per request, bill from your own usage records or from access logs, not from the assumption that the quota was enforced to the request. **Quotas are blunt.** A customer that exhausts a monthly quota on day three is simply refused for twenty-seven days. Real products soften this with usage notifications, overage tiers, or a plan change workflow — none of which API Gateway provides. Usage data is retrievable per key and per day, which is what you build those workflows on. **Key distribution is your problem.** Creating, rotating and revoking keys per customer is an onboarding workflow you have to build and secure; a leaked key is a stolen quota until it is revoked. ## How to reason about it in an interview The strong answer distinguishes the three jobs that are easy to blur together: *authenticating* a caller (an authorizer), *identifying whose allowance to charge* (the API key and its usage plan), and *protecting the backend from total load* (stage, method and account throttles). Usage plans do the middle job well and the third job per-caller — but they do not do the first at all, and relying on a key as a secret is the single most common mistake with this feature.
- A customer's key is attached to a plan, yet its calls are never metered against it. What do you check first?Whether the method actually requires an API key. If it does not, the request is served anonymously and no key is resolved, so the plan's per-key throttle and quota never apply. After that, confirm the caller really sends the key in `x-api-key`, and that the plan's api-stages list includes the exact API and stage being called.
- Why should billing not be derived from the usage-plan quota?Because enforcement is best-effort across a distributed system, so the boundary is approximate rather than exact. The quota is a guardrail that stops runaway consumption, not an accounting record. Bill from access logs or your own usage pipeline, and use the quota only to cap exposure.
- Your product requires per-customer quotas but you wanted an HTTP API for its lower cost. What are the options?Either accept a REST API for that surface, since HTTP APIs have no API keys or usage plans, or move metering into your own layer — an authorizer or the application counting per-tenant usage against a store and rejecting over-limit callers. The second option is more code and more latency, so the honest comparison is that engineering cost against the per-request price difference.
saying these in an interview costs you the question
- Treating an API key as a secret credential
- Assuming HTTP APIs support usage plans
- Forgetting to mark the method as requiring a key
- Billing customers straight from quota enforcement
- Believing one usage plan can meter callers with no key