A public API needs to enforce three limits simultaneously: 1,000 requests/day per API key, 20 requests/second per API key (burst control), and 50,000 requests/second across all clients combined (protecting the backend). Why can't a single counter satisfy all three, and what problems come from combining per-key and global limits?
answer
- different windows, different keys
- daily quota = contract, per-sec = burst, global = capacity
- check all, reject on first fail
- signal which limit was hit
- tiered/reserved capacity for fairness
basics
~20 sEach limit protects a different thing over a different time scale, so each needs its own counter and window; a single number can't answer 'am I under my daily cap AND my per-second burst cap AND is the whole system under its global cap' at once.
solid answer
~50 sPer-key limits (daily quota, per-second burst) protect fairness and enforce a pricing/usage contract for that individual client, while a global limit protects shared backend capacity regardless of who's calling. These operate on different keys (per-API-key vs a single system-wide key) and different windows (a day vs a second), so they must be tracked as separate counters checked independently — a request is only admitted if it passes all applicable checks, typically evaluated cheapest-first (global check can be a fast local approximation, per-key checks hit the shared store). The main complication is that combining them creates a fairness problem: if the global limit is the binding constraint, a legitimate client well within their own per-key quota can still get 429'd because other tenants exhausted the shared budget, which is confusing unless the API clearly signals which limit was hit (e.g., a distinct error code or header per limit type).
go deeper
Should recognize that limits can exist 'per client' and 'for everyone at once' as two different things.
Should explain why the same request needs to be checked against multiple independent counters with different keys and windows.
Expected to design the check ordering, response signaling (which limit was hit), and articulate the fairness problem a global limit introduces for well-behaved clients.
Expected to design tiered/reserved capacity allocation to resolve the fairness gap and reason about the operational cost of maintaining multiple budget dimensions at scale.
## Three limits, three different jobs Multi-dimensional rate limiting exists because different limits protect different things, at different time scales, against different failure modes, and conflating them into one counter throws away information needed to make the right accept/reject decision. - A **daily per-key quota** (1,000/day) is fundamentally a usage/business-contract mechanism: it enforces what the client has purchased or been allocated, and its violation means "this specific client has used more than their allowance," independent of whether the system is under any load at all — you'd reject the 1,001st request even at 3 a.m. with a totally idle backend. - A **per-second per-key burst limit** (20/sec) protects against a single client's traffic spikes overwhelming the connection they're using or triggering downstream rate limits the provider itself is subject to; it operates on a much shorter window and is really about smoothing, not total volume. - A **global cross-client limit** (50,000/sec system-wide) protects the shared backend's actual physical capacity — database connections, thread pools, CPU — and is completely indifferent to which client sent what; its purpose is preventing aggregate overload regardless of how fairly that load is distributed among clients. ## How the checks compose Mechanically this means three independent counters (or counter families) with three different keys and three different windows, one keyed by each of: | Dimension | Keyed by | |---|---| | daily quota | API-key+day | | burst | API-key+second | | global | a single global constant+second | A request must pass all applicable checks to be admitted — the limiter evaluates each one, typically ordered cheapest/most-likely-to-reject first as an optimization (e.g. check a fast local approximate global counter before round-tripping to a shared store for the precise per-key check), and rejects on the first one exceeded, ideally returning a response that identifies which limit was hit (a header like `X-RateLimit-Scope: global` vs per-key, or distinct error codes) so the caller isn't left guessing why it was throttled when its own dashboard shows it's nowhere near its daily quota. ## Why one counter cannot do all three The reason a single counter cannot satisfy this is that the three limits have **incompatible semantics**: - A counter keyed per-API-key can never enforce a global ceiling, because by construction it only ever sees one client's traffic and has no visibility into the other nine hundred clients also calling concurrently. - A counter keyed globally can never enforce per-client fairness or a per-client contractual quota, because it can't attribute usage back to who caused it. - A per-second counter and a per-day counter fundamentally answer different questions — a client bursting to exactly 20/sec sustained for an hour blows through neither instantaneously, but easily exhausts the daily 1,000-request quota in under a minute, so the two windows catch different failure patterns and neither subsumes the other. ## The fairness cost of a global limit The trade-off introduced by combining per-key and global limits is **fairness versus protection**. Global limits are necessary because backend capacity is a shared, finite resource that no amount of per-key fairness can conjure more of — if every one of a thousand well-behaved clients simultaneously sends traffic within their own per-key allowance, the sum can still exceed what the backend can handle. But once a global limit is the binding constraint, individual well-behaved clients get rejected through no fault of their own, purely because other tenants happened to be active at the same moment; this is confusing and, if not clearly signaled, looks like a bug in the client's own usage tracking ("my dashboard says I'm at 200 of 1,000 daily requests, why am I getting 429s?"). Well-designed APIs handle this with either: - differentiated response bodies/headers per limit dimension, - or priority tiers — paying customers get a reserved slice of global capacity so a free-tier surge can't degrade paid traffic — which introduces its own complexity of maintaining separate global budgets per tier rather than one shared pool. ## What to watch for in production The production failure mode to watch for is **under-communicating which limit was hit**: a flat 429 with no distinguishing signal forces client-side engineers to guess whether they need to slow their own burst rate, wait for the next day's quota reset, or simply retry shortly because the whole system was momentarily hot — three completely different remediations that look identical from a bare status code.
- Should the global limit or the per-key limit be checked first?Order doesn't matter for correctness since a request must pass all checks to be admitted either way, but it matters for efficiency: checking a cheap, often-local approximate global counter first lets you reject early under a system-wide overload without paying the cost of a precise per-key lookup, while checking per-key limits first is preferable when you want to give clients an accurate, specific rejection reason before a global check might mask it.
- How would you give paying customers protection from free-tier traffic spikes exhausting the shared global limit?Partition the global budget into reserved slices per tier — e.g., 80% of global capacity reserved for paid traffic, 20% for free tier — so a free-tier surge can only exhaust its own slice, not the paid slice. This adds complexity (separate budgets to track and rebalance) but directly fixes the fairness gap of one shared pool.
- Why might a per-second burst limit and a daily quota both be necessary even though the daily quota is the larger, more restrictive number over time?They protect against different failure shapes: a client could stay well under the daily total while still sending an instantaneous spike large enough to overwhelm a connection or downstream dependency in a single second, which the daily quota alone would never catch since it only aggregates over 24 hours.
It's like an apartment building with three separate rules at once: each tenant has a monthly water allowance (per-user daily quota), a max flow rate their pipes can handle at once (per-user burst limit), and the building's total water main capacity shared by everyone (global limit) — hitting any one of the three stops the tap, for different reasons.
saying these in an interview costs you the question
- Thinks one counter can serve all three purposes
- Doesn't distinguish contractual/business limits from capacity-protection limits
- Assumes a client hitting the global limit is doing something wrong
- No plan for telling the client which limit it hit