You own the nginx edge in front of an API with a login endpoint, a search endpoint and a bulk export. How would you design the rate-limit zones, keys and rates, and roll them out without cutting off real users?
answer
- one zone per abuse shape
- the key decides who suffers together
- several limit_req directives all apply
- measure in dry run before enforcing
- per-instance zones multiply across a fleet
basics
~20 sUse one zone per abuse shape rather than one global limit: a slow per-IP zone for login, a per-identity zone for search, and a concurrency cap for export. Measure with dry-run and logging before enforcing, and divide rates by the number of nginx instances.
solid answer
~50 sStart from what each endpoint's abuse looks like, because a single global rate is always wrong for something. Login is credential stuffing: a very low per-IP `limit_req_zone` in requests per minute, small burst, no forgiveness. Search is expensive per call but legitimately bursty: key it on an authenticated identity built with `map` rather than on address, so shared corporate egress is not one bucket, with `burst` plus `nodelay`. Bulk export is an occupancy problem, so it wants `limit_conn` and `limit_rate`, not a request rate. Multiple `limit_req` directives can apply to one location and all of them are checked, so a coarse per-server ceiling can sit under the specific ones. Roll it out in three steps: log the key and outcome, run with `limit_req_dry_run on;` and read `$limit_req_status` to see who *would* have been rejected, then enforce. Two things people forget: zones are per nginx instance, so N nodes multiply your effective limit; and setting `limit_req_status 429` lets clients and dashboards tell a limit from an outage.
code
nginx · 31 lineshttp {
map $http_x_api_key $api_client {
default $binary_remote_addr;
~^(?<id>\S+)$ $id;
}
limit_req_zone $binary_remote_addr zone=login:10m rate=10r/m;
limit_req_zone $api_client zone=search:20m rate=20r/s;
limit_conn_zone $api_client zone=exportconn:10m;
limit_req_status 429;
limit_conn_status 429;
log_format ratelimit '$remote_addr key=$api_client '
'limit=$limit_req_status status=$status';
server {
access_log /var/log/nginx/rl.log ratelimit;
location = /login { limit_req zone=login burst=3;
proxy_pass http://auth; }
location /search/ { limit_req zone=search burst=40 nodelay;
limit_req_dry_run on;
proxy_pass http://app; }
location /export/ { limit_conn exportconn 2;
limit_rate 1m;
proxy_pass http://app; }
}
}go deeper
Know that separate zones exist for separate purposes and that the zone key decides what is being counted. Be able to point at login as the endpoint that needs the strictest limit.
Explain why the key choice matters — address versus authenticated identity — and that several limit_req directives compose rather than override. Mention zone sizing and LRU eviction.
Describe a safe rollout: log the key and outcome, run in dry run and read $limit_req_status, then enforce one zone at a time, with 429 so the signal is legible to clients and dashboards.
Own the architectural boundary: the edge limit is an approximate capacity guard because zones are per instance, while a precise per-customer quota belongs behind a shared counter. Justify the collateral damage each key choice accepts.
## Start from the abuse shape, not the endpoint list A single `limit_req` at the `http` level is the configuration everybody writes first and nobody keeps. It is simultaneously too strict for a page that fires twelve parallel calls and too loose for a login form. The design question is: for each class of traffic, what is the scarce resource and what does misuse look like? | Endpoint | Scarce resource | Abuse shape | Right tool | |---|---|---|---| | `/login` | credential guesses | many attempts, low volume each | very low per-IP request rate | | `/search` | backend CPU per query | bursty legitimate use, scripted scraping | per-identity request rate with burst | | `/export` | connections and bandwidth | few requests, held open for minutes | concurrency and bandwidth cap | ```nginx http { limit_req_zone $binary_remote_addr zone=login:10m rate=10r/m; limit_req_zone $api_client zone=search:20m rate=20r/s; limit_req_zone $server_name zone=global:10m rate=2000r/s; limit_conn_zone $api_client zone=exportconn:10m; limit_req_status 429; limit_conn_status 429; limit_req_log_level warn; } ``` ## Choosing the key is choosing who suffers together The key defines the fairness unit and therefore the collateral damage. - **`$binary_remote_addr`** works for anonymous traffic and is the only option before authentication — which is why it is right for login. Its cost is that a corporate NAT or a mobile carrier gateway is one bucket for thousands of people. - **An identity you derive** is right for authenticated traffic. Build it with `map` so a missing credential falls back to the address rather than collapsing everyone into an empty key: ```nginx map $http_authorization $api_client { default $binary_remote_addr; ~^Bearer\s+(?<tok>\S+)$ $tok; } ``` Be deliberate here: a raw credential becomes a zone key held in shared memory and, if you log the variable, written to disk. Prefer a non-secret identifier the edge can see, or a hash. - **`$server_name`** is for a fleet-wide ceiling that protects the platform rather than any user. Several `limit_req` directives may be active in one location, and nginx checks all of them — so the specific and the coarse limits compose rather than override. ## Sizing the zones A zone holds one state per distinct key. nginx documents roughly 16,000 64-byte states per megabyte for `limit_req`. When the zone is full, the least recently used states are evicted — which quietly weakens the limit under a wide-source attack, because an attacker rotating addresses evicts everyone's counters, including their own. Size against your expected *distinct key* population with headroom, not against request volume, and prefer `$binary_remote_addr` over `$remote_addr` for the smaller state. ## Rolling it out without an outage The failure mode of a rate-limit launch is discovering your assumed rate was three times too low, at peak, in production. Sequence it: **1. Observe.** Add the key and outcome to the access log format so you can build the actual distribution before choosing a number. Pick a rate at some percentile of observed real behaviour, not at what feels tidy. **2. Dry run.** Since nginx 1.17.6, `limit_req_dry_run on;` evaluates the limit and logs the decision without rejecting anything, and the `$limit_req_status` variable reports `PASSED`, `DELAYED`, `REJECTED`, `DELAYED_DRY_RUN` or `REJECTED_DRY_RUN`. `limit_conn_dry_run` does the same for concurrency. Run a full traffic cycle — a weekday peak and a batch window — and count who *would* have been rejected. This is the step that separates a design from a guess. ```nginx log_format ratelimit '$remote_addr $uri key=$api_client ' 'limit=$limit_req_status status=$status'; ``` **3. Enforce, narrowly first.** Turn dry run off for the highest-confidence zone (usually login, where a wrong number harms attackers more than users), watch the rejection rate, then widen. **4. Give clients something actionable.** `limit_req_status 429;` distinguishes a limit from a backend outage in every dashboard and client retry policy you own. Leaving the 503 default guarantees someone eventually pages an on-call engineer for a working system. ## The multiplication problem The zone is shared memory inside one nginx process group. Behind a load balancer with six nginx nodes, a client whose requests are spread across them can achieve six times the configured rate, and a client pinned to one node gets exactly the configured rate. Neither is what the config says. The options, in increasing cost: - **Divide** the rate by the node count. Simple, and wrong whenever the fleet scales or traffic distributes unevenly. - **Accept** the imprecision. Rate limiting at the edge is a blunt instrument; if it is a coarse safety valve rather than a quota, approximate is fine — say so explicitly. - **Move the precise limit inward** to a service backed by a shared counter, and keep the nginx limit as the cheap outer guard that sheds obvious floods before they cost anything. This layering is usually the right principal-level answer: the edge protects capacity, the application enforces the contract. ## What to say about tuning over time A rate limit is not a fire-and-forget config. Alert on rejection rate as a proportion of traffic — a jump means either an attack or a limit that has aged past a legitimate client's growth, and you want to know which within minutes rather than from a support ticket. Keep the zones and rates in the same review path as the rest of the edge config, and re-run the dry-run measurement whenever the client mix changes materially.
- Two `limit_req` directives from different zones apply to one location. Which one governs?Both. nginx checks every configured limit and rejects a request that exceeds any of them, so the effective behaviour is the most restrictive. That is deliberate and useful: a tight per-identity limit can sit alongside a coarse per-server ceiling that protects the platform when many identities misbehave at once. There is no override or last-wins semantics.
- Why not simply divide the rate by the number of nginx nodes and call it correct?Because traffic is rarely distributed evenly and the node count changes. A client pinned to one node by connection reuse gets the divided rate — far stricter than intended — while a client whose requests spread evenly gets close to the original. Division makes the limit unpredictable rather than accurate. Either accept the edge limit as approximate, or enforce the real quota behind a shared counter.
- What is the risk of keying a zone on a bearer token extracted with `map`?You are placing a live credential into a shared memory key and, if the variable appears in your log format, writing it to disk and into your log pipeline. Prefer a non-secret identifier the edge can already see — an API key ID, a tenant header — or key on a hash. The rate-limiting behaviour is identical and the disclosure risk disappears.
- How do you decide the rate number itself rather than guessing?Log the key and outcome, build the real distribution of per-key request rates over a full traffic cycle including peaks and batch windows, then set the limit at a percentile that covers legitimate use with headroom. Confirm the choice with `limit_req_dry_run on;` and `$limit_req_status`, counting who would have been rejected, before enforcing anything.
saying these in an interview costs you the question
- Applying one global rate limit to every endpoint
- Assuming a later limit_req directive overrides an earlier one
- Enforcing a new limit without measuring real traffic first
- Treating per-instance zones as a fleet-wide quota
- Leaving the 503 default so limits look like outages