In nginx, how does `limit_req_zone` with `rate=10r/s` actually meter requests, and what do the `burst` and `nodelay` parameters on `limit_req` change about which clients get rejected?
answer
- leaky bucket, not a per-second quota
- rate becomes a minimum interval
- burst queues, it does not raise the rate
- nodelay releases the queue immediately
- rejected excess answers 503 by default
basics
~20 snginx's limit_req is a leaky bucket: rate=10r/s admits one request every 100 ms per key, not ten at once. burst=N queues that many excess requests instead of rejecting them; nodelay forwards the queued ones immediately while slots still refill at the configured rate.
solid answer
~50 s`limit_req_zone` declares a shared memory zone keyed by a variable (usually `$binary_remote_addr`) plus a rate. nginx implements it as a leaky bucket, so `rate=10r/s` really means *one request every 100 ms* for that key — a client that fires ten requests in the same millisecond has nine of them treated as excess, and by default excess is rejected with 503. `burst=N` gives each key a queue of N excess requests: instead of being rejected they are held and released at the configured pace, so a bursty but low-average client survives. `nodelay` keeps the same queue accounting but forwards those queued requests **immediately**, and the occupied slots then drain back at the zone's rate — you get one instant burst of N followed by strict rate limiting. Since 1.15.7 `delay=M` gives you two stages: the first M are instant, the rest of the burst is throttled. Rejections are logged at `error` level and the status is configurable with `limit_req_status`.
code
nginx · 15 lineshttp {
limit_req_zone $binary_remote_addr zone=api:10m rate=10r/s;
limit_req_status 429;
limit_req_log_level warn;
server {
listen 80;
server_name api.example.com;
location /api/ {
limit_req zone=api burst=20 nodelay;
proxy_pass http://backend;
}
}
}go deeper
Know that limit_req_zone declares the counter and rate while limit_req applies it, and that the key is normally $binary_remote_addr. Be able to say that excess requests are rejected by default.
Explain the leaky-bucket model out loud: rate becomes a minimum interval per key, burst adds a queue, nodelay releases that queue immediately while slots refill at the rate. Mention the 503 default.
Show you have tuned this against real traffic — that burst without nodelay turns errors into client timeouts nobody can trace, and that every rejection writes an error-log line. Talk about zone sizing and LRU eviction.
Own the trade between blocking abuse and rejecting real users. Argue for measuring the current distribution before enforcing, for a status code callers can act on, and for knowing that per-instance zones multiply across a fleet.
## The two directives Rate limiting in nginx needs two things: a place to keep counters, and a rule that uses them. ```nginx http { limit_req_zone $binary_remote_addr zone=api:10m rate=10r/s; server { location /api/ { limit_req zone=api burst=20 nodelay; proxy_pass http://backend; } } } ``` `limit_req_zone` lives in the `http` context and takes three things: the **key** (any nginx variable — here the client address in binary form), the **zone** name and size, and the **rate**. `limit_req` then applies that zone in `http`, `server` or `location` context. The zone is a *shared memory* region: all worker processes read and write the same counters, so the limit is per nginx instance, not per worker. Size matters — nginx documents roughly 16,000 64-byte states per megabyte for `limit_req`, so `10m` holds on the order of 160,000 distinct keys. When the zone fills, nginx evicts the least recently used entries, and if it still cannot make room it returns an error for the request. `$binary_remote_addr` is preferred over `$remote_addr` because it stores four bytes for IPv4 (sixteen for IPv6) instead of a text string, which materially changes how many clients fit in the zone. ## Leaky bucket, not "ten per second" The single most common misreading is treating `rate=10r/s` as a per-second quota that resets on a clock boundary. It does not. nginx converts the rate into a minimum interval between accepted requests for that key — 10r/s is one request per 100 ms, 30r/m is one request per two seconds. There is no window that resets; the bucket leaks continuously. So with a bare `limit_req zone=api;` and no burst, a browser that opens a page and fires eight parallel XHRs to the same host has *one* request served and seven rejected, even though the client's average rate over the next second is well under ten. That is why a naive rate limit tends to break real traffic while barely inconveniencing an attacker, whose requests are evenly paced by a script. ## What burst does `burst=N` attaches a queue of N slots to each key. Excess requests occupy a slot instead of being rejected, and nginx releases them at the zone's rate. Requests arriving when the queue is already full are rejected. The cost is **latency**. With `rate=10r/s burst=20` and 21 simultaneous requests, the last queued one waits about two seconds before nginx even proxies it. For a browser this shows up as a page that loads its assets in visible stages; for an API client with a two-second read timeout it shows up as a timeout that no backend log explains, because the backend never saw the request. ## What nodelay does `nodelay` changes only the *scheduling*, not the accounting. Burst slots are still taken and still drain back at the zone rate, but a request that gets a slot is forwarded straight away. The practical shape is: a client may spend its whole burst instantly, and then is held to the steady rate until slots free up. This is almost always what you want in front of a browser-facing or API endpoint: real clients are bursty and idle, abusers are sustained. Without `nodelay` you convert a rejection into a delay, which usually just moves the failure into someone's client timeout. Since nginx 1.15.7 you can split the difference: ```nginx limit_req zone=api burst=20 delay=8; ``` The first 8 excess requests go through immediately; requests 9–20 of the burst are queued and paced; beyond that, rejected. ## Rejection behaviour Rejected requests get **503 Service Unavailable** by default. `limit_req_status 429;` changes that, which is worth doing so callers and your own dashboards can tell a limit from a genuine outage. `limit_req_log_level warn;` (default `error`) controls the log line, and every rejection *is* logged — a badly tuned limit is a fast way to fill a disk. Several `limit_req` directives may apply at once; all configured limits are checked and a request is rejected if it trips any of them. That is how you stack a per-IP limit with a coarser per-server limit. ## Reading it back in practice When someone reports intermittent 503s, the first check is the error log for `limiting requests, excess: ... by zone "..."`. The excess figure tells you how far over the bucket the client was, which is what you tune `burst` against. If the reports are of *slowness* rather than errors, suspect a `burst` without `nodelay`.
- With `burst=20` and no `nodelay`, a client reports timeouts rather than errors. Why?Because burst converts rejection into delay. The 20th queued request at `rate=10r/s` waits about two seconds before nginx forwards it, which exceeds a typical client read timeout. The client sees a timeout, and the backend has no log line at all because the request never reached it. Adding `nodelay`, or lowering `burst`, turns the silent delay back into a fast, visible rejection.
- Why is `$binary_remote_addr` preferred over `$remote_addr` as the zone key?It stores the address in binary — 4 bytes for IPv4, 16 for IPv6 — rather than as text, so each state in the shared memory zone is smaller and a given zone size tracks far more distinct clients. nginx's own documentation uses it in every example for this reason. The metering behaviour is identical; only the memory footprint differs.
- What happens when the shared memory zone runs out of space?nginx evicts the least recently used states to make room. If it still cannot allocate, the request is rejected with an error and a log entry. In practice a too-small zone silently weakens the limit under a wide-source attack, because entries for individual clients are evicted before their buckets have drained. Size the zone against your expected distinct-key count, not your request rate.
saying these in an interview costs you the question
- Saying rate=10r/s allows ten simultaneous requests each second
- Thinking burst raises the sustained rate rather than queuing excess
- Assuming the counter resets on a clock-second boundary
- Believing nodelay disables the limit or prevents all rejections
- Expecting a 429 by default instead of 503