skip to content

A reverse proxy terminates the client's connection and opens its own connection to the backend. Why does the client keeping its connection alive not give you connection reuse on the backend leg, and what has to be true at the proxy for that reuse to happen?

level: middleimportance: should knowfreq 55%

answer

  1. two connections, not one pipe
  2. client leg says nothing about upstream
  3. reuse is opt-in on the backend side
  4. HTTP/1.1 plus strip the Connection header
  5. the pool is per worker process

basics

~20 s

A reverse proxy runs two independent connections per request, client-side and upstream-side, so client keep-alive says nothing about the backend leg. Upstream reuse needs an explicit pool, HTTP/1.1 on that leg, and no Connection: close forwarded through.

solid answer

~50 s

A reverse proxy is two connection endpoints glued together: it is a server to the client and a client to the backend, and those two legs have separate lifetimes, separate timeouts and separate pools. A client that holds one connection open for a hundred requests tells you nothing about how many TCP and TLS handshakes the backend paid — by default many proxies open a fresh upstream connection per request and close it after the response. To get reuse you have to enable an upstream pool explicitly, speak HTTP/1.1 (or HTTP/2) on that leg, and make sure hop-by-hop headers are stripped so the proxy is not forwarding `Connection: close` to the backend. Remember the pool is normally per worker process, so the real number of idle connections a backend sees is workers times pool size times proxy instances.

go deeper

for a junior

Know that a reverse proxy ends the client's connection and makes its own to the backend, so the two sides are counted and configured separately.

for a middle

Be ready to explain the three preconditions for upstream reuse — an explicit pool, HTTP/1.1 on that leg, hop-by-hop headers stripped — and what a missing one costs per request.

for a senior

Show that you size the pool by workers times instances against backend file descriptors, and that you keep the proxy's idle timeout shorter than the backend's so pooled connections are retired rather than found dead.

for a principal

Own the platform default: pooling that every team inherits, a documented budget for how many idle connections a backend may be asked to hold, and a way to detect tiers that quietly reopen a connection per request.

## The proxy is two endpoints, not a pipe It is tempting to picture a reverse proxy as a piece of tubing that a request slides through. It is not. A reverse proxy terminates the client's transport: it completes the TCP handshake, usually the TLS handshake too, parses the request, and only then acts as a **client** toward the backend, opening or borrowing a second connection of its own. Two sockets, two states, two sets of timers. Everything follows from that split. The client's connection can stay open for minutes across dozens of requests while the proxy opened and closed thirty separate backend connections. Or the reverse: the client can be a one-shot script that connects, sends, and disconnects, while the proxy served it entirely from a warm pooled connection to the backend. Neither leg governs the other. ## Why reuse on the upstream leg is not free On the client side, persistent connections are the norm and cost nothing to arrange. On the upstream side, several proxies default to a new connection per request, and there are reasons: pooling requires the proxy to track idle connections per backend, to detect ones the backend has closed underneath it, and to make sure a connection is never handed to a second request while the first is still using it. The cost of *not* pooling is paid on every single request: a TCP handshake (one round trip), and if the backend leg is encrypted, a TLS handshake (one to two more round trips plus asymmetric crypto on both ends). On a fast internal network the latency is small but the CPU and the socket churn are not, and the backend ends up with a large population of short-lived connections and sockets in post-close states. Under load that shows up as accept-queue pressure and CPU spent in handshakes rather than in your handler. ## What the proxy actually has to do Three things, and skipping any one of them silently disables reuse: 1. **Declare a pool.** Idle upstream connections have to be kept somewhere with a size limit. In nginx this is the `keepalive` directive inside an `upstream` block; other proxies express it as a connection-pool setting on the cluster or backend. 2. **Use a protocol version that supports persistence on that leg.** HTTP/1.0 upstream means close-per-response. Proxies that default to HTTP/1.0 upstream must be told to use 1.1. 3. **Strip hop-by-hop headers.** `Connection`, and the headers it names, apply to a single hop. A proxy that copies the client's `Connection` header onto the upstream request can hand the backend a `close` it never intended and tear down the pooled connection after every response. A typical nginx shape, for orientation: ```nginx upstream app { server 10.0.0.11:8080; keepalive 32; } location /api/ { proxy_pass http://app; proxy_http_version 1.1; proxy_set_header Connection ""; } ``` The `Connection ""` line is the one people forget, and its absence looks exactly like a working config — right up until you count connections on the backend. ## Sizing, and why the number is bigger than it looks The pool limit is almost always **per worker process**, not per proxy instance. Eight workers with a pool of 32 is 256 idle connections from one host, and ten such hosts is 2,560 before a single request arrives. Multiply before you set the number: the backend pays for each one in a file descriptor, a socket buffer and, for threaded servers, potentially a thread. Oversizing has a second effect. Idle connections are the ones a backend is most likely to close on its own idle timer, and a pooled connection that the backend closed a moment ago is exactly the connection the proxy is about to reuse. The more idle capacity you hold, the wider that window. The general defence is to keep the proxy's upstream idle timeout comfortably shorter than the backend's, so the proxy retires connections rather than discovering them dead mid-request. ## What to take away Count the two legs separately. Ask which one your metric is measuring, which one your timeout applies to, and whether the reuse you assume on one side actually exists on the other. Most surprises at this layer are someone reasoning about a single connection when there were always two.

  • How would you size the upstream keep-alive pool, and what goes wrong if you set it far too high?
    Start from concurrency, not intuition: the pool only needs to cover the in-flight requests one worker sustains. Multiply pool by workers by proxy instances and check that against the backend's file-descriptor and memory budget. Oversizing pins descriptors on the backend for connections nobody is using and widens the window in which the backend closes an idle connection just as the proxy reuses it.
  • What should a proxy do with hop-by-hop headers when it reuses an upstream connection?
    Strip them. `Connection` and any header it names apply to one hop only, so the proxy must consume them rather than copy them onto the upstream request. Forwarding a client's `Connection: close` tells the backend to end a connection the proxy wanted to keep, which quietly turns a pooled setup back into one connection per request.
  • If the backend leg is plain HTTP inside a private network, is pooling still worth it?
    Usually yes, though the payoff is smaller. You save one round trip per request rather than three, but you also stop churning sockets: fewer accepts, fewer closes, less time in connection setup on the backend's event loop or thread pool. On a high-request-rate service that churn, not the latency, is the reason to pool.

saying these in an interview costs you the question

  • Assumes client keep-alive is passed through to the backend
  • Thinks one client connection means one backend connection
  • Sets a pool size without multiplying by worker count
  • Believes upstream connection reuse is on by default everywhere
  • Forwards the client's Connection header to the upstream unchanged

context