skip to content

A handful of clients are exhausting an nginx server by holding many simultaneous slow downloads open. Why does `limit_conn` help here where `limit_req` does not?

level: middleimportance: must knowfreq 58%

answer

  1. arrival rate versus simultaneous occupancy
  2. one slow request never trips a rate limit
  3. only in-flight, fully-read requests are counted
  4. idle keep-alive connections are free
  5. pair concurrency caps with timeouts

basics

~20 s

limit_req meters request arrival rate, so a client that opens one request and keeps it open forever never trips it. limit_conn caps concurrent in-flight connections per key, which is the resource slow downloads and slowloris-style clients actually consume.

solid answer

~50 s

The two modules limit different resources. `limit_req` counts *arrivals* — a leaky bucket over how often a key sends a request. A client that opens fifty connections and then trickles one slow request down each of them has an arrival rate near zero, so it sails past any `limit_req` rule while occupying fifty worker slots. `limit_conn`, declared with `limit_conn_zone $binary_remote_addr zone=addr:10m;` and applied as `limit_conn addr 10;`, caps how many connections from that key are being processed at once, which is exactly the resource under pressure. Note what nginx counts: only connections where the whole request header has been read and a request is being processed. An idle keep-alive connection is not counted. In practice you pair `limit_conn` with the timeout directives — `client_header_timeout`, `client_body_timeout`, `send_timeout` — because a cap on concurrency without a cap on duration still lets a slow client sit in its slots indefinitely.

code

nginx · 19 lines
nginx
http {
    limit_conn_zone $binary_remote_addr zone=perip:10m;

    client_header_timeout 10s;
    client_body_timeout   10s;
    send_timeout          10s;

    server {
        listen 80;
        server_name files.example.com;

        location /downloads/ {
            limit_conn perip 5;
            limit_rate 512k;
            limit_rate_after 1m;
            root /srv/files;
        }
    }
}

go deeper

for a junior

Be able to name both modules and say plainly that one limits how often requests arrive and the other limits how many connections are open at the same time.

for a middle

Explain the counting rule — only connections with a fully-read request header being processed are counted — and give a concrete attack shape that a rate limit misses entirely.

for a senior

Show the layered config you would actually ship: concurrency cap plus header, body and send timeouts plus a bandwidth cap, and say what each one closes off. Mention that counters are per instance.

for a principal

Frame it as choosing the fairness unit and accepting its collateral damage: keying on client address punishes shared NAT egress, and per-node zones mean the fleet-wide ceiling is a multiple of what the config states.

## Two modules, two resources nginx ships two distinct limiting modules and candidates routinely conflate them. - `ngx_http_limit_req_module` — `limit_req_zone` / `limit_req`. Limits the **rate of request arrivals** per key. - `ngx_http_limit_conn_module` — `limit_conn_zone` / `limit_conn`. Limits the **number of concurrent connections** per key. Both use a shared memory zone keyed by an nginx variable, both are configurable in `http`, `server` and `location` contexts, and both reject with 503 by default (`limit_req_status`, `limit_conn_status`). That surface similarity is why they get muddled. ```nginx http { limit_conn_zone $binary_remote_addr zone=perip:10m; limit_conn_zone $server_name zone=persrv:10m; server { location /downloads/ { limit_conn perip 5; limit_conn persrv 200; limit_rate 512k; } } } ``` ## Why a rate limit misses this attack Rate limiting assumes the harmful thing is *how often* a client asks. Whole classes of resource exhaustion do not look like that: - **Slow reads.** A client requests a large file and then reads a few bytes per second. One request, one arrival, indefinite occupancy. - **Slow request bodies (slowloris-shaped).** The client sends headers or a body one byte at a time so the request is never complete. - **Concurrent bulk downloads.** Perfectly legitimate-looking requests, just fifty of them at once from one source. In every case the arrival rate is trivially low. `limit_req` sees nothing to reject. What is consumed is connection slots — and, behind nginx, upstream connections and backend worker threads. `limit_conn` measures exactly that. `limit_conn perip 5;` means the sixth concurrent connection from that address being processed is rejected immediately, while the first five continue. ## The counting rule that catches people out nginx counts only connections **where a request is being processed and the whole request header has already been read**. Two consequences: 1. An idle keep-alive connection does not consume a `limit_conn` slot. A browser holding six persistent connections to your host is not sitting at your limit while the user reads the page. 2. A connection that has sent partial headers is *also* not counted yet — which is precisely the slowloris shape. `limit_conn` therefore does not by itself solve the incomplete-header attack; `client_header_timeout` does, by closing connections that dawdle over their headers. So the honest answer in an interview is that `limit_conn` is one of three controls, not a complete defence: ```nginx client_header_timeout 10s; # kill slow header senders client_body_timeout 10s; # kill slow body senders send_timeout 10s; # kill clients that stop reading the response limit_conn perip 5; # cap concurrency once a request is in flight ``` ## Choosing the key The key is any variable, and the choice defines the fairness unit: - `$binary_remote_addr` — per client address. Compact, and the usual default. Punishes shared NATs and corporate egress, which is the trade you accept. - `$server_name` — a ceiling for one virtual host so a single site cannot starve its neighbours on a shared instance. - A variable you build with `map` — for example a tenant identifier extracted from a header — when the address is not the meaningful unit. One caution: the zone is shared memory inside *this* nginx instance. Behind a load balancer with several nginx nodes, each keeps its own counters, so the effective concurrency ceiling is the per-node limit multiplied by the node count. ## Pairing with bandwidth control `limit_conn` caps how many streams a client gets; `limit_rate` (and `limit_rate_after`) caps how fast each one flows. Together they bound the bandwidth one key can take: five connections at 512k each is a hard 2.5 MB/s ceiling. Setting only `limit_rate` invites the client to compensate by opening more connections; setting only `limit_conn` lets each connection run at line speed. Interviewers like this pairing because it shows you reasoned about what the client will do next. ## When to reach for which Ask what the scarce resource is. If it is *work performed per unit time* — database queries behind an API, login attempts, expensive search — use `limit_req`. If it is *simultaneous occupancy* — connections, sockets, large transfers, long-lived streams — use `limit_conn`. Real edge configs usually carry both, in separate zones, because the two abuse shapes are independent.

  • Does `limit_conn` count an idle keep-alive connection against the client's allowance?
    No. nginx counts only connections where a request is being processed and the entire request header has been read. A browser parked on six persistent connections with nothing in flight consumes no slots, which is what makes a low limit such as 5 or 10 workable for real users rather than an accidental denial of service against them.
  • If `limit_conn` does not count partially-received headers, what actually stops a slowloris client?
    `client_header_timeout`, and for bodies `client_body_timeout`. Those close a connection whose headers or body arrive too slowly, so the attacker cannot hold sockets open cheaply. limit_conn covers the phase after the request is complete; the timeouts cover the phase before. You need both, plus `send_timeout` for clients that stop reading the response.
  • How do `limit_conn` and `limit_rate` complement each other?
    `limit_conn` caps how many simultaneous streams one key gets; `limit_rate` caps the bytes per second of each response. Applying only a bandwidth cap invites the client to open more connections to compensate; applying only a concurrency cap lets each connection run at full line speed. Together they put a hard ceiling on the bandwidth a single key can consume.

saying these in an interview costs you the question

  • Claiming limit_req also caps concurrent connections
  • Assuming every open TCP connection counts toward limit_conn
  • Treating limit_conn alone as a complete slowloris defence
  • Thinking the zone is shared across separate nginx instances
  • Forgetting limit_conn rejects with 503, not 429, by default

context