skip to content

Why do many HTTP servers and load balancers cap how many requests they will serve on a single connection - for example advertising `Keep-Alive: timeout=5, max=100` - and what does a client observe when that cap is reached?

level: seniorimportance: should knowfreq 32%

answer

  1. max requests per connection = hygiene + rebalancing
  2. Keep-Alive: timeout=5, max=100 is an advisory hint
  3. Final response carries Connection: close, not an abort
  4. Keep-alive pins clients to old backends after scale-out
  5. HTTP/2 equivalent = GOAWAY frame

basics

~20 s

Caps reclaim per-connection memory and leaks, and force long-lived clients back through the load balancer so newly added backends get traffic. On the last permitted request the server answers normally but adds Connection: close; the client should transparently open a new connection.

solid answer

~50 s

Two reasons dominate. **Hygiene.** Every open connection holds buffers, TLS state, and whatever the server framework accumulated per connection. Recycling periodically bounds slow leaks and lets the process reclaim memory; it also gives a natural point for the server to apply new config or certificates. **Load distribution.** HTTP/1.1 keep-alive pins a client to one backend for the life of the connection. If a fleet scales out, existing long-lived connections keep hammering the old instances and the new ones stay idle. Capping requests (or connection lifetime) forces periodic re-resolution and rebalancing. Mechanically the server serves the final request normally and adds `Connection: close` to that response; some also send `Keep-Alive: timeout=5, max=100` as an informational hint. A correct client reads the response, discards the connection, and opens a new one - so this is invisible unless the cap is so low that you are paying a handshake every few requests. HTTP/2 does the same thing with a GOAWAY frame, which lets in-flight streams finish first.

code

http · 10 lines
http
GET /orders/42 HTTP/1.1
Host: api.example.com

HTTP/1.1 200 OK
Content-Type: application/json
Content-Length: 17
Keep-Alive: timeout=5, max=1
Connection: close

{"id":42,"ok":true}

go deeper

for a junior

Know that servers may limit how many requests one connection serves, and that the last response carries Connection: close so the client just opens a new connection.

for a middle

Explain both motivations - reclaiming per-connection resources and forcing periodic reconnection - and the exact wire behaviour of the final response.

for a senior

Tie it to operations: keep-alive pinning defeating autoscaling and draining, tuning keepalive_requests against handshake cost, and GOAWAY as the HTTP/2 equivalent.

for a principal

Set fleet policy: connection lifetime as a rebalancing lever versus handshake CPU, coordination with deploys, certificate rotation and load-balancer connection-duration limits, and the sharper pinning risk under HTTP/2 or HTTP/3.

## The knob Most servers expose two connection-recycling controls: an **idle timeout** (close a connection with no traffic for N seconds) and a **maximum requests per connection**. nginx calls the latter `keepalive_requests` (its default was 100 for years and was raised to 1000 in 1.19.10), Apache calls it `MaxKeepAliveRequests`, and many application servers and proxies have equivalents. Servers may advertise both as the informational `Keep-Alive: timeout=5, max=100` header, though clients are not required to honour it and mostly treat it as a hint. ## Why cap at all **Bounded resource lifetime.** A connection is a container for state: read/write buffers, TLS session material, per-connection objects in the server framework, sometimes a thread. Small leaks that would be invisible on a short-lived connection become significant on one that lives for days. Recycling turns an unbounded leak into a bounded one. **Rebalancing.** This is the operationally important reason. Under HTTP/1.1 with keep-alive, a client that opened 50 connections to a load balancer keeps using those same 50 TCP flows, and each is pinned to whichever backend it landed on. Autoscaling adds three instances and they receive **nothing**, because no new connections are being made. The fleet looks scaled out while the old nodes stay hot. Bounding the number of requests (or the absolute age) of a connection guarantees periodic re-connection, which re-runs DNS and load-balancer selection. The same mechanism helps traffic drain away from an instance being replaced. **Config and certificate rollovers.** New TLS certificates, new routing rules or new limits typically apply at connection setup. Long-lived connections keep running under old settings until they are recycled. **Defence in depth.** A cap limits how much a single connection can do, which slightly raises the cost of some slow-drip abuse patterns and keeps a per-connection accounting bug from being unbounded. ## What the client sees The server does not abort. It serves the final permitted request completely, adds `Connection: close` to that response, and closes after the body is written. A compliant client removes the connection from its pool and opens a fresh one for the next request. The only visible symptom is a periodic handshake - and the corresponding latency spike on that one request. Problems show up when the cap is misconfigured relative to traffic: - **Cap too low for a chatty client** (say 100 requests on a client doing thousands per second per connection): you pay TLS handshakes constantly, server CPU rises, and p99 gets a saw-tooth. Raising `keepalive_requests` is the standard remedy for reverse proxies fronting internal APIs. - **Cap interacting with the idle timeout race.** Recycling driven by `Connection: close` is safe because the signal rides on a response. Recycling driven purely by an idle timeout is not: the server can close while a pooled connection sits idle and a client may write into it at the same moment. That stale-connection race is a separate problem from the request cap and is why explicit `Connection: close` is the preferred signal. ## HTTP/2 and HTTP/3 Multiplexed protocols do not use `Connection: close` - it is a forbidden header. Instead the server sends a **GOAWAY** frame naming the highest stream ID it will process, letting in-flight requests finish while refusing new ones on that connection; the client opens a new connection for subsequent work. Many managed load balancers also enforce a maximum connection *duration* for the same rebalancing reasons, and GOAWAY is how they announce it. Because HTTP/2 concentrates all of a client's traffic onto one connection, pinning is worse there, which makes periodic GOAWAY-based recycling more important, not less. ## Practical guidance Set the cap high enough that handshakes are a rounding error (thousands of requests, or a duration measured in minutes), low enough that rebalancing happens on the timescale of your autoscaling. Verify clients handle `Connection: close` and GOAWAY without surfacing errors, and watch new-connection rate as a metric - a sudden spike often means someone lowered a cap.

  • You scale a backend fleet from 4 to 8 instances and the new instances receive almost no traffic. What is happening and how do you fix it?
    Long-lived keep-alive connections from the callers are pinned to the original backends, so no new connection selection occurs and the new instances stay idle. Fix it by bounding connection lifetime - a maximum requests per connection or a maximum connection age on the proxy or client - so connections are periodically recycled and re-balanced. With HTTP/2, have the proxy send periodic GOAWAY frames for the same effect.
  • How does HTTP/2 achieve the same connection recycling without Connection: close?
    The server sends a GOAWAY frame containing the highest stream ID it will still process. Streams below that ID complete normally, new requests on that connection are refused, and the client opens a fresh connection for subsequent work. It is a graceful, in-band shutdown that avoids the abrupt truncation you would get from simply closing the TCP connection.

A shop that politely says "this is the last order I'll take at this counter" and finishes serving you, so the queue can be redistributed across newly opened counters - rather than slamming the shutter mid-transaction.

saying these in an interview costs you the question

  • Thinking the server aborts the in-flight request when the cap is reached
  • Treating the advisory Keep-Alive: timeout/max header as a binding contract the client must obey
  • Assuming keep-alive automatically spreads load across backends when it actually pins traffic to them
  • Setting the cap very low 'for safety', paying constant TLS handshakes
  • Expecting Connection: close to work in HTTP/2 rather than GOAWAY

context