An API Gateway needs to stop any single client from making more than 100 requests per minute. Walk through how a token-bucket rate limiter enforces that, and why teams often layer both per-client and global limits at the gateway.
answer
- token bucket allows bursts up to capacity
- fixed window boundary burst problem
- per-client vs global limit layering
- 429 + Retry-After
- shared counter store needed across gateway instances
basics
~20 sThe gateway keeps a running count of how many requests each caller has made recently. Once a caller crosses the limit, the gateway rejects further requests for a while instead of forwarding them, protecting the backend services from being overwhelmed.
solid answer
~40 sA token bucket gives each client (identified by API key, IP, or user ID) a bucket that holds up to N tokens and refills at a fixed rate (e.g., 100 tokens/minute). Each request consumes one token; if the bucket is empty, the request is rejected with 429 and a Retry-After header. This allows short bursts up to the bucket size while enforcing a steady average rate, which is why it's preferred over a naive fixed-window counter that lets clients burst up to 2x at window boundaries. Gateways layer a per-client limit (fairness — one noisy client can't starve others) with a global/service limit (protects a downstream service's actual capacity regardless of how many distinct clients are calling it), because a per-client limit alone doesn't prevent many well-behaved clients from collectively overwhelming a backend.
go deeper
Knows the gateway can reject a client that sends too many requests too fast, returning some kind of error.
Can explain the token-bucket mechanism (capacity + refill rate) and knows to return 429 with Retry-After.
Reasons about layering per-client and global limits, and knows shared-state counters (e.g., Redis) are required once the gateway is horizontally scaled.
Designs tiered limiting strategy (per-tenant, per-route, global) balancing fairness against backend capacity, and anticipates retry-storm and traffic-spike failure modes at rollout.
## The mechanism: a token bucket per key Rate limiting at the API Gateway is the mechanism that caps how many requests a given caller (or the system as a whole) can make in a time window, rejecting the excess before it ever reaches a backend service. The most common algorithm is the **token bucket**: the gateway maintains, per rate-limit key (an API key, client IP, authenticated user ID, or a combination), a counter representing available 'tokens' up to some maximum capacity, plus a fixed refill rate — e.g., a bucket holding 100 tokens that refills at 100 tokens per 60 seconds, refilled continuously or in small increments rather than in one lump sum. Each incoming request checks the bucket: - **if at least one token is available**, it's decremented and the request proceeds; - **if the bucket is empty**, the gateway immediately returns HTTP 429 Too Many Requests along with a `Retry-After` header telling the client how long to wait, without ever forwarding the request downstream. ## Why a bucket beats a naive fixed window The key property that makes token bucket popular over a naive fixed-window counter (reset the count to zero every 60 seconds on the clock) is that it tolerates bursts smoothly: a client that's been idle can burst up to the bucket's full capacity, then settles into the steady refill rate, whereas a fixed window lets a client send its full quota right at the end of one window and its full quota again immediately at the start of the next, briefly doubling the effective rate right at the boundary. A **sliding-window** algorithm improves on fixed windows similarly, at the cost of more memory/computation to track a rolling history rather than a single counter. ## Why it lives at the edge Rate limiting exists at the gateway, rather than in each backend service, for the same reason auth is centralized there: it's a cross-cutting concern that's cheaper and more consistent to enforce once, at the edge, before wasted work happens downstream. Enforcing it at the gateway also means rejected requests never consume backend CPU, database connections, or thread-pool slots — the protection happens before the expensive part of the request lifecycle, which is the whole point of load shedding at the edge rather than deep inside the call graph. ## The trade-off: one dimension is never enough The trade-off that shows up almost immediately in real systems is that a single limiting dimension is never enough. - A pure **per-client limit** (say, 100 req/min per API key) protects fairness between tenants — no single noisy client can starve others — but does nothing to protect a downstream service's actual capacity: if a backend can sustainably handle 500 requests/second and there are 1,000 distinct well-behaved clients each individually under their 100/min limit, they can still collectively swamp it. - So production gateways typically layer multiple limits: **per-client limits** for fairness, and a **global or per-route limit** that caps total throughput to a value the backend is known to handle, sometimes combined with **per-tenant-tier limits** (a paid tier gets a higher bucket than a free tier). Layering adds operational complexity — now there are multiple counters to track, multiple limits to tune, and multiple places a request can be rejected for different reasons, which needs to be surfaced clearly to the caller. ## Failure modes in production Failure modes in production include: 1. **Rate-limit state not being shared across gateway instances** in a horizontally scaled fleet, so a client that should be capped at 100 req/min effectively gets 100 req/min per gateway instance unless the limiter uses a shared, fast store like **Redis** for counters — a very common bug when teams first scale their gateway out. 2. Another failure mode is **limits set too conservatively** during a traffic spike (a marketing launch, a retry storm from a struggling downstream dependency) causing legitimate traffic to be throttled and amplifying a partial outage into a full one from the client's perspective. 3. A third is **thundering-herd retries**: many clients receiving 429s simultaneously and retrying at the same fixed interval, which just recreates the traffic spike at the next window boundary unless clients implement jittered backoff honoring `Retry-After`. ## Where you have seen it A concrete real-world example: **Kong Gateway's** rate-limiting plugin implements exactly this pattern, offering both local (per-node) and cluster-wide (Redis-backed) counting so limits hold consistently across a horizontally scaled gateway fleet. **AWS API Gateway** has built-in usage plans with per-API-key throttle and burst limits (token-bucket semantics, configurable rate and burst capacity) applied before requests reach the integrated Lambda or backend, so the Lambda's concurrency is protected from being overwhelmed by a single misbehaving caller.
- Why does a fixed-window counter allow an effective 2x burst that a token bucket avoids?With a fixed window, a client can send its full quota in the last second of one window and its full quota again in the first second of the next window, since the counter simply resets at the boundary — those two bursts land close together in real time. Token bucket avoids this because tokens refill continuously/incrementally rather than resetting in a lump, so there's no boundary to exploit.
- You scale the gateway from 1 instance to 5 behind a load balancer, and suddenly clients can send 5x their configured limit. What went wrong and how do you fix it?Each gateway instance was tracking its own local counter, so a client landing on different instances across requests effectively got a separate bucket per instance. The fix is to back the counters with a shared, low-latency store (typically Redis) so all instances see and decrement the same counter for a given rate-limit key.
- If a client is legitimately hitting its rate limit, what should the gateway return besides a 429 status code, and why does it matter?A Retry-After header telling the client how many seconds to wait before retrying. It matters because without it, clients tend to retry immediately or on a fixed interval that isn't synchronized with the actual bucket refill, which can create a retry storm that keeps hitting the limit and wastes both client and gateway resources.
Like a nightclub bouncer with a wristband stamp: you get a limited number of re-entry stamps per night (your bucket), refilled slowly, and the bouncer also caps total people inside regardless of how many individuals still have stamps left (the global limit protecting the room's actual capacity).
saying these in an interview costs you the question
- Thinks rate limiting only needs a single global counter with no per-client dimension
- Doesn't know why fixed-window counters can allow boundary bursts
- Assumes rate-limit counters automatically stay consistent across multiple gateway instances
- No mention of what response code/header to return to a throttled client
- Believes rate limiting is enforced inside each backend service rather than at the edge