skip to content

questions

6

What is rate limiting, and what should an API return when a client exceeds its limit?

level: juniorimportance: must knowfreq 80%

answer

  1. 429 = slow down, not broken
  2. Retry-After tells when
  3. protects finite capacity
  4. per-key threshold
  5. X-RateLimit-* headers

basics

~10 s

Rate limiting caps how many requests a client can send in a time window. Once they go over it, the server rejects the extra requests with HTTP 429 and tells them when to retry.

solid answer

~40 s

Rate limiting protects a service from being overwhelmed by capping how many requests a client (or the whole system) can make in a time window, e.g. 100 requests/minute per API key. Once a client crosses that threshold, the server rejects further requests with HTTP 429 Too Many Requests instead of processing them or letting them cascade into a crash downstream. A well-behaved API also returns a Retry-After header (seconds, or an HTTP-date) telling the client when it's safe to retry, plus commonly X-RateLimit-Limit/Remaining/Reset headers so clients can self-throttle proactively instead of guessing and hammering the endpoint. This protects finite downstream resources (databases, third-party APIs, CPU), keeps usage fair across tenants, and gives well-behaved clients a clear backpressure signal instead of silent drops or timeouts.

go deeper

for a junior

Should know 429 is the status code, that it signals the client (not the server) is at fault for the rejection, and that some header tells the client when to retry — doesn't need algorithm detail.

for a middle

Should additionally know why per-client keying matters (API key vs IP), name X-RateLimit-Remaining/Reset headers, and explain the thundering-herd risk of retrying without honoring Retry-After.

for a senior

Expected to reason about choosing the right key granularity per endpoint (login vs read-only), the trade-off between strict limits (false positives) and loose ones (residual risk), and design a coherent 429 response contract across an API surface.

for a principal

Expected to treat rate limiting as part of a broader capacity/reliability strategy — connecting it to SLOs, cost control, tiered pricing, and abuse defense — and to make the fail-open/fail-closed call for when the limiter itself is degraded.

## What the mechanism actually does Rate limiting is a mechanism that caps the number of operations a given identity is allowed to perform within a defined time window. That identity may be: - a user - an API key - an IP address - or the system as a whole Mechanically, every incoming request carries or implies a **key**: - the API key in a header - the authenticated user ID - the source IP for anonymous traffic The server, or a dedicated limiter sitting in front of it, looks up a counter or budget tied to that key and compares it against a configured threshold, for example 100 requests per minute. 1. If the request is within budget, the limiter records the usage and lets the request through. 2. If the client has exceeded the threshold, the limiter short-circuits the request before it reaches business logic and responds with HTTP status **429 Too Many Requests**. This is deliberately distinct from a 5xx server error: 429 tells the caller "the server is fine, but you personally must slow down," which is actionable information a generic 500 or a dropped connection is not. ## Why it exists The reason rate limiting exists is **capacity protection and fairness**. Backend capacity is finite — database connection pools, CPU, memory, and quotas on third-party services called on a client's behalf all have hard ceilings. Without a limiter, one misbehaving client (a buggy retry loop, a scraper, a compromised script) can consume all of that capacity and degrade the service for every other client — the classic "noisy neighbor" problem in a multi-tenant system. Rate limiting also defends against abuse patterns such as: - credential-stuffing login attempts - content scraping - brute-force enumeration and it lets a business meter and monetize usage, since many SaaS APIs tie limits directly to pricing tiers. Beyond protecting the provider, a well-designed 429 is also a service to the caller: it converts an ambiguous failure (timeout, connection reset) into an explicit, machine-readable signal the caller can act on immediately rather than discover through trial and error. ## The retry signal The single most important piece of that signal is the `Retry-After` header, which tells the client either a number of seconds to wait or an HTTP-date after which it may try again. Complementary headers most production limiters also return, though none are a single universal standard: | Header | What it carries | |---|---| | `X-RateLimit-Limit` | the ceiling | | `X-RateLimit-Remaining` | budget left in the current window | | `X-RateLimit-Reset` | when the window rolls over | Together these let a well-behaved client throttle itself proactively, spacing requests out before ever hitting a 429, rather than reactively hammering the endpoint and backing off only after rejection. ## The core trade-off The core trade-off is **protecting the server versus rejecting legitimate traffic**. - **Set the threshold too low, or key it too coarsely** — by IP address, say, when many real users sit behind the same corporate NAT or mobile carrier gateway — and false positives appear: real users get 429s during ordinary usage spikes, like a mobile app's retry burst after regaining connectivity. - **Set it too high or key it too finely**, and a single compromised identity can still degrade shared resources before the limiter engages. Choosing the key and the threshold is therefore a product and capacity-planning decision as much as an engineering one, and it usually needs different values per endpoint — a login endpoint needs a much tighter limit than a read-only search endpoint because the cost and abuse profile differ. ## Failure modes Failure modes cluster around **omission of the retry signal** and **inconsistency across the fleet**. 1. **A bare 429 with no `Retry-After`** leaves poorly written clients — including some naive retry libraries — looping immediately, turning a single overloaded moment into a self-inflicted thundering herd that keeps the service pinned at capacity. 2. **A limiter implemented per application instance with in-memory counters behind a load balancer** is another common failure: each instance enforces "100/minute" independently, so a client fanned out across ten instances effectively gets 1,000/minute, defeating the limit's purpose — the classic argument for either sticky routing or a shared, centralized counter store. 3. **Clock skew** between distributed limiter instances or between client and server can also cause a `Retry-After` value to be honored too early or too late. ## Where it shows up A concrete, widely known example: - **GitHub's REST API** returns rate-limit status via X-RateLimit-Limit/Remaining/Reset headers on every response and rejects further calls once exhausted until the reset time. - **Stripe's API** similarly returns 429 with `Retry-After`, and its official client libraries implement exponential backoff that specifically honors that header rather than retrying on a fixed interval — the behavior any well-designed client should replicate.

  • What's the difference between rate limiting a client and a service returning 503?
    429 says the caller specifically exceeded their allotted usage and should retry later per Retry-After; 503 says the service overall is unavailable regardless of who's asking, often used for general overload or maintenance. Conflating them loses information the caller needs to react correctly — a 429 client should back off just its own traffic, while a 503 client should treat the whole endpoint as down.
  • Why key rate limits by API key or user ID rather than IP address when possible?
    Many real users can share one IP (corporate NAT, mobile carrier CGNAT, university networks), so IP-based limiting punishes innocent users sharing an address with a heavy or abusive one. API key or authenticated user ID ties the limit to the actual billing/usage identity, which is both fairer and harder to spoof than an IP.
  • Should a client retry immediately after getting a 429 without a Retry-After header?
    No — retrying immediately risks worsening the exact overload condition that caused the 429, especially if many clients do it simultaneously (thundering herd). Absent Retry-After, a client should fall back to its own exponential backoff with jitter.

Like a nightclub bouncer with a one-in-one-out rope: once capacity is hit you're turned away and told 'try again in 10 minutes' rather than let in and having the place collapse.

saying these in an interview costs you the question

  • Says 429 means the server crashed
  • Recommends retrying immediately in a tight loop after 429
  • Doesn't mention Retry-After or any retry signal at all
  • Assumes rate limiting only exists to stop malicious attackers, not fair-use/capacity

context

open as a page

A service enforces 'max 60 requests per minute per client' using a fixed window counter that resets every minute on the wall clock (e.g. at :00 of each minute). Why can a client legitimately send close to 120 requests in a short burst spanning a minute boundary, and how does a sliding window approach address this?

level: middleimportance: must knowfreq 82%

basics

~20 s

Fixed windows reset abruptly at a clock boundary, so a client can max out the limit right before the reset and again right after — nearly double the intended limit in a short burst. Sliding window looks at a rolling time range instead of a fixed clock tick, so it never lets that double-up happen.

open as a page

Compare the token bucket and leaky bucket algorithms for limiting a single client's request rate. How does each one handle a burst of traffic that arrives all at once, and when would you pick one over the other?

level: middleimportance: must knowfreq 88%

basics

~10 s

Token bucket lets a client save up unused capacity and spend it in a burst. Leaky bucket smooths everything out to a fixed steady rate no matter how bursty the input is.

open as a page

You're implementing a rate limiter shared across many stateless application server instances, backed by Redis, using a simple GET-then-SET pattern: read the current counter, check it against the limit, and if under, increment and write it back. Under concurrent load, why does this let clients exceed their limit, and how do you fix it?

level: seniorimportance: must knowfreq 75%

basics

~20 s

Two requests can both read the counter before either writes it back, so both see 'under limit' and both proceed — a race condition called a check-then-act bug. The fix is to make the check-and-increment a single atomic operation, e.g. Redis INCR or a Lua script.

open as a page

A public API needs to enforce three limits simultaneously: 1,000 requests/day per API key, 20 requests/second per API key (burst control), and 50,000 requests/second across all clients combined (protecting the backend). Why can't a single counter satisfy all three, and what problems come from combining per-key and global limits?

level: seniorimportance: should knowfreq 60%

basics

~20 s

Each limit protects a different thing over a different time scale, so each needs its own counter and window; a single number can't answer 'am I under my daily cap AND my per-second burst cap AND is the whole system under its global cap' at once.

open as a page

You're designing a rate limiter for an API gateway deployed across three geographic regions, each with many stateless gateway instances, enforcing one global per-API-key limit. A single centralized Redis in one region adds cross-region latency to every request and becomes a single point of failure. Walk through the design trade-offs: what happens if you use per-region local counters instead of a shared store, and what happens if the central Redis becomes unreachable?

level: principalimportance: should knowfreq 45%

basics

~30 s

Local per-region counters are fast but only see local traffic, so a client hitting all three regions can get up to 3x its real limit unless you divide the budget per region. A central Redis is accurate but adds latency and is a single point of failure — you must decide in advance whether to fail open (allow everything, risk overload) or fail closed (reject everything, cause an outage) when it's unreachable.

open as a page