skip to content

You added `keepalive 32;` to an nginx `upstream` block, but connections to the backend are still opened and closed for every request. What else must the proxying location set, and what does that number actually count?

level: middleimportance: should knowfreq 52%

answer

  1. one directive is never enough
  2. HTTP/1.0 cannot keep alive
  3. the default Connection header says close
  4. the number counts idle, per worker
  5. whoever times out first causes the 502

basics

~20 s

The keepalive directive alone is not enough: the location must also set proxy_http_version 1.1 and clear the Connection header with proxy_set_header Connection "". The number is the count of idle connections each worker caches, not a connection limit.

solid answer

~50 s

`keepalive` only enables a cache of idle upstream connections; two defaults still force a close on every request. Nginx proxies with HTTP/1.0 unless you set `proxy_http_version 1.1;`, and it sends `Connection: close` upstream by default, which you clear with `proxy_set_header Connection "";`. All three are required together — that is the classic incomplete config. The number itself is often misread: it is the maximum number of **idle** keepalive connections retained per worker process, not a cap on concurrent connections to the backend, so nginx can and will open more than 32 under load and simply close the least recently used idle ones above the limit. Because it is per worker, the real idle pool is roughly the number times `worker_processes`. I also check `keepalive_timeout` against the backend's own idle timeout: if the backend closes first, nginx can hand a request to a connection the peer just tore down and return a 502.

code

nginx · 16 lines
nginx
upstream backend {
    least_conn;
    server 10.0.0.11:8080;
    server 10.0.0.12:8080;
    keepalive 32;
    keepalive_timeout 60s;
}

server {
    location / {
        proxy_pass http://backend;
        proxy_http_version 1.1;
        proxy_set_header Connection "";
        proxy_set_header Host $host;
    }
}

go deeper

for a junior

Learn the three lines that must appear together: keepalive in the upstream block, proxy_http_version 1.1, and proxy_set_header Connection "". Be able to say why reuse saves a handshake per request.

for a middle

Explain why each directive is needed — HTTP/1.0 default, default Connection: close — and state precisely that the number bounds idle connections per worker rather than total connections.

for a senior

Bring the operational half: size the pool against what the backend can hold, order idle timeouts so nginx closes first, and diagnose the reuse race behind sporadic 502s using $upstream_connect_time.

for a principal

Own it as a platform default: a standard proxying snippet with the three directives, a documented rule that upstream idle timeouts exceed the proxy's, and awareness that long-lived pooled connections change how traffic redistributes after a backend is added or drained.

## Three directives, not one Upstream keepalive is the single most frequently half-configured thing in nginx. The complete form is: ```nginx upstream backend { server 10.0.0.11:8080; server 10.0.0.12:8080; keepalive 32; } server { location / { proxy_pass http://backend; proxy_http_version 1.1; proxy_set_header Connection ""; } } ``` Why each line exists: - **`keepalive 32;`** creates the idle-connection cache for the group. Without it there is nothing to reuse. - **`proxy_http_version 1.1;`** — nginx proxies using HTTP/1.0 by default, and persistent connections are not the default there. Leave this out and every response ends with a close. - **`proxy_set_header Connection "";`** — nginx's default proxied headers include `Connection: close`. Setting the header to the empty string removes it, so the backend is not asked to close. Note also that `Connection` is hop-by-hop: a client-supplied value must not be relayed, which is why nginx handles this header explicitly rather than passing it through. Miss any one and the other two do nothing observable. ## What the number means `keepalive N` is the maximum number of **idle** connections to the upstream group preserved in the cache of **one worker process**. It is not a limit on how many connections that worker may open: under a burst nginx opens as many as it needs, and when more than N of them fall idle, the least recently used are closed. Sizing it as though it were a connection cap is the standard misreading, and it leads people to set it enormous "so we don't run out". Because it is per worker, a host with `worker_processes auto;` on 8 cores and `keepalive 32;` can hold about 256 idle connections to that group. Size it against what the backend can hold open, not against your request rate. ## Two companions - **`keepalive_requests`** — how many requests may be served over one upstream connection before nginx closes it. As of nginx 1.19.10 the default is 1000; it was 100 before. It exists to stop per-connection memory allocations growing without bound and to let long-lived pools rebalance. - **`keepalive_timeout`** in the upstream context — how long an idle upstream connection is kept. One ordering rule catches people: when the group uses a balancing method other than the default round robin, the method directive must appear **before** `keepalive` in the block. ## The race that produces mystery 502s Connection reuse creates a window that does not exist without it. Nginx picks an idle connection and writes a request; at almost the same moment the backend, whose own idle timeout is shorter, closes it. Nginx sees a reset on a connection it believed usable. The symptom is a small, steady trickle of 502s under otherwise normal traffic, often clustered after quiet periods. The fix is ordering the timeouts deliberately: the backend's idle timeout should be **longer** than nginx's `keepalive_timeout` for the group, so nginx is always the side that closes. Nginx's `proxy_next_upstream` behaviour can retry an idempotent request that failed this way, but relying on retries to paper over mismatched timeouts is treating the symptom. ## What you actually gain Every avoided connection is a TCP handshake, and if the hop is encrypted, a TLS handshake too — measurable latency on every request, plus per-connection work on both machines and churn in the local port and socket-state accounting on the proxy host. On a chatty internal path the improvement is usually visible immediately in `$upstream_connect_time`, which drops to near zero for reused connections while `$upstream_response_time` stays where it was. Logging both variables is the cheapest way to prove that keepalive is actually working, and it is a much better answer than "the config looks right". ## Interview framing The question tests whether you have configured this yourself or copied it. The three-directive requirement, the per-worker idle-cache meaning of the number, and the shorter-backend-timeout race are the three things that separate the two.

  • Does keepalive 32 limit nginx to 32 connections to that upstream group?
    No. It bounds the number of idle connections cached per worker process. Nginx opens as many connections as concurrency demands, and when more than 32 go idle in a worker it closes the least recently used ones. With several workers the total idle pool is the number multiplied by worker_processes, so it is sized against what the backend can hold open.
  • You enable upstream keepalive and start seeing a small trickle of 502s. What is the likely cause?
    A race on reused connections: the backend's idle timeout is shorter than nginx's, so it closes a pooled connection just as nginx dispatches a request onto it. Order the timeouts so nginx always closes first — make the backend's idle timeout longer than the group's keepalive_timeout — rather than relying on proxy_next_upstream to retry.
  • How would you confirm from nginx's own logs that connections are actually being reused?
    Add `$upstream_connect_time` and `$upstream_response_time` to the log format. On a reused connection the connect time is effectively zero while the response time is unchanged; if connect time stays non-trivial on every request, the pool is not being used and one of the three required directives is missing.

saying these in an interview costs you the question

  • Thinks the keepalive directive alone enables connection reuse
  • Reads keepalive N as a maximum number of upstream connections
  • Forgets that nginx proxies with HTTP/1.0 by default
  • Leaves the default Connection: close header in place
  • Ignores that the number is per worker process

context