An nginx error log fills with "768 worker_connections are not enough". What do the `worker_processes` and `worker_connections` directives control, how do they combine into a capacity ceiling, and why does a reverse-proxy workload hit that ceiling sooner than a static-file one?
answer
- concurrency limit, not requests per second
- multiply workers by connections
- idle keep-alive still holds a slot
- proxying spends two slots per request
- descriptors must be raised with it
basics
~20 sworker_processes sets how many worker processes nginx runs; worker_connections caps the simultaneous connections each one may hold. The ceiling is their product, and a reverse proxy spends two slots per request — one client-side, one upstream — so it reaches that ceiling at roughly half the client count.
solid answer
~50 s`worker_processes` (main context) is the number of worker processes, normally `auto`, which matches the CPU core count. `worker_connections` (events context) is the per-worker limit on simultaneous connections — not requests per second, and it counts every connection the worker holds, including idle keep-alive ones, upstream connections and internal ones. Capacity is therefore roughly `worker_processes × worker_connections`. For static file serving each client connection consumes one slot, so that product is close to the client ceiling. When nginx proxies, each request also opens or reuses a connection to the upstream, and both sit in the same per-worker budget, which halves the effective client count. The error means workers ran out of slots and nginx stopped accepting, so clients see hangs or resets while CPU looks fine. The fix is to raise `worker_connections` and, because each connection needs a file descriptor, raise `worker_rlimit_nofile` to match — but first check whether the real cause is upstreams that are slow to respond, which is what leaves so many connections open at once.
code
nginx · 15 linesworker_processes auto;
worker_rlimit_nofile 65535;
events {
worker_connections 16384;
}
http {
keepalive_timeout 30s;
upstream app {
server 10.0.0.11:8080;
keepalive 64;
}
}go deeper
Know that worker_processes counts nginx's worker processes and worker_connections limits simultaneous connections per worker, and that the two multiply into the total ceiling.
Explain what counts against the limit — idle keep-alive connections and upstream connections included — and why a proxy therefore serves roughly half as many clients as the raw product suggests.
Diagnose before tuning: relate concurrency to upstream latency and hold time, decide whether the load is legitimate, and raise the connection limit and worker_rlimit_nofile together while watching live active connections.
Own the capacity model: what connection concurrency this tier is sized for, whether long-lived protocols get their own tier, and what the proxy should do at the ceiling — shed load deliberately rather than accumulate a backlog nobody can serve.
## What the two directives actually mean nginx runs one master process and several workers. The master reads configuration, binds listening sockets and manages workers; the workers do all the connection handling, each in a single-threaded event loop. - **`worker_processes`** (main context) sets how many workers exist. `worker_processes auto;` — supported since nginx 1.2.5 — sets it to the number of CPU cores detected, which is the right default for a CPU-bound event loop. - **`worker_connections`** (events context) caps how many simultaneous connections *one worker* may hold. The nginx default is 512; several distribution packages ship 768 or 1024, which is where the number in that error message usually comes from. The key misreading is treating `worker_connections` as throughput. It is a concurrency limit — a count of open connections at one instant — and it includes everything the worker holds: active client connections, idle keep-alive client connections waiting for another request, connections to upstream servers, and connections nginx opens for its own purposes. ## The capacity arithmetic The absolute ceiling is: ``` max simultaneous connections = worker_processes × worker_connections ``` With 4 workers at 768 that is 3,072 connections in total — but connections are not clients. For a **static file server**, each client connection uses one slot, so the ceiling is roughly the client ceiling. Keep-alive already inflates it, because an idle connection a browser is holding open still occupies a slot even though no request is in flight. For a **reverse proxy**, each proxied request needs a client-side connection *and* a connection to the upstream, and both are charged to the same worker's budget. The practical client ceiling is therefore about half the product. Enabling `keepalive` in the `upstream` block reduces how often those upstream connections are created, but pooled connections still occupy slots while they are held. ## Reading the error correctly `768 worker_connections are not enough` means a worker hit its limit and could not take on more. Symptoms at the client are hangs, resets and connections that never get a response, while nginx's CPU and the application's own logs look unremarkable — nginx never got far enough to log much per request. Before raising the number, ask *why* so many connections are open at once. By Little's law, concurrency is arrival rate times how long each connection is held. A ceiling reached at modest traffic usually means connections are being held far too long, and the honest causes are: - **Slow upstreams.** If backend latency climbs from 50 ms to 5 s, concurrency rises a hundredfold at the same request rate. Raising the connection limit here just lets nginx queue more of a backlog it cannot clear. - **Long keep-alive with many idle clients.** A generous `keepalive_timeout` and a large client population means a large idle population holding slots. - **Long-lived connections by design** — WebSocket upgrades, server-sent events, streaming downloads — where high concurrency is normal and the limit genuinely needs to be sized for it. ## Raising it safely When the concurrency is legitimate, raise `worker_connections` — values in the tens of thousands are ordinary for a proxy — and raise `worker_rlimit_nofile` (main context) alongside it, because every connection and every open file consumes a file descriptor, and the per-worker descriptor limit must comfortably exceed `worker_connections` to leave room for upstream sockets, log files and cached file handles. A configuration that raises connections but not descriptors trades one error message for another. `worker_rlimit_nofile` sets the limit nginx requests for its worker processes; the value the operating system is willing to grant is a separate, host-level concern. Also reconsider `worker_processes`: on a machine where nginx is one tenant among several, pinning workers to fewer than the core count is sometimes right, and it lowers total capacity proportionally, so the two numbers must be chosen together. ```nginx worker_processes auto; worker_rlimit_nofile 65535; events { worker_connections 16384; } ``` And instrument it: the stub status module exposes the live active-connection count, so you can watch headroom against the configured ceiling instead of rediscovering it from the error log.
- Why is raising `worker_connections` sometimes the wrong response to that error?Because the connection count is a symptom of how long connections are held. If upstream latency has grown, concurrency rises at the same request rate, and a bigger limit only lets nginx accumulate a deeper backlog it still cannot serve — turning a fast rejection into a slow one for everybody. Fix the latency, or shed load deliberately, before enlarging the buffer.
- What must you change alongside `worker_connections`, and why?`worker_rlimit_nofile`, in the main context. Every connection consumes a file descriptor, and a worker also needs descriptors for upstream sockets, log files and cached file handles. It should sit comfortably above `worker_connections`, otherwise nginx trades the connection-limit error for descriptor-exhaustion errors at the same load.
- Does `worker_processes auto` always give the best throughput?It is the right default when nginx owns the machine, since each worker is a single-threaded event loop and one per core avoids context switching. On a shared host it can be wrong: workers compete with the co-located application for CPU. Whatever you choose, remember it multiplies the connection ceiling, so lowering the worker count lowers total capacity proportionally.
saying these in an interview costs you the question
- Treats worker_connections as a requests-per-second limit
- Says idle keep-alive connections do not occupy a slot
- Forgets a proxied request also holds an upstream connection
- Raises worker_connections without raising the descriptor limit
- Assumes the fix is always a bigger number rather than a slow upstream