A service intermittently fails with 'connection reset by peer' or an empty response on roughly one request in a thousand, always just after a quiet period, while the target service is healthy. What race is happening and how do you address it?
answer
- Server FIN on idle; client pool never saw it; writes into a dead socket
- Symptoms: ECONNRESET / NoHttpResponse / EOF, clustered after quiet periods
- No server access-log entry = request never read
- Client idle timeout < server (and < LB, < NAT) idle timeout
- Validate-after-inactivity + retry once when zero bytes received
basics
~20 sThe server's keep-alive idle timeout closed a pooled connection at the same moment the client leased it and wrote a request. The client never saw the FIN, so it sends into a dying socket. Keep the client's idle timeout below the server's, validate connections before use, and retry safe requests once.
solid answer
~1 minThis is the **stale connection race**, inherent to HTTP/1.1 pooling. The server decides an idle connection has outlived its keep-alive timeout and sends a FIN. The client's pool is not actively reading idle connections, so it has not observed the close. A request arrives, the pool hands out that connection, and the client writes into a socket the server has already closed - producing an RST ("connection reset by peer") or an EOF with no response at all. There is no way to eliminate it: checking liveness and writing cannot be atomic, so the close can always land in between. You reduce and absorb it: 1. **Client idle timeout strictly below the server's** - if the load balancer idles at 60 s, expire client connections at 30-50 s so the client closes first. 2. **Validate after inactivity** - many clients re-check a connection that has been idle longer than a threshold before leasing it, which closes most of the window. 3. **Retry once** when nothing was received - safe because the server demonstrably did not respond. Restrict automatic retries to idempotent requests. 4. Have the server advertise `Keep-Alive: timeout=...` and prefer `Connection: close` on a response over a silent idle close.
code
bash · 3 linescurl -v --keepalive-time 5 https://api.example.com/a https://api.example.com/b
# then compare against the server's configured keepalive_timeout
curl -sI https://api.example.com/a | grep -i keep-alivego deeper
Know that a pooled connection can be closed by the server while the client still thinks it is usable, and that clients retry such a request on a new connection.
Describe the sequence - server FIN, unobserved by the idle pool, request written into a closing socket - and the two mitigations: lower client idle timeout and validate before reuse.
Confirm it from evidence (no server-side log entry, failures clustered after idle gaps, timing matching keep-alive) and tune the whole chain of idle timeouts including the load balancer and NAT, with a bounded retry only for safe requests.
Treat it as an invariant of HTTP/1.1 pooling: define fleet-wide idle-timeout ordering from client to edge to origin, prefer in-band close signals, and note that HTTP/2 GOAWAY largely removes the class.
## The race, step by step 1. The client's pool holds an idle, warm connection to the server. 2. The server's keep-alive idle timer expires. The server sends FIN and considers the connection finished. 3. The FIN sits in the client's socket receive queue. Nobody is reading it: pooled idle connections are not actively polled by most clients, and even those that poll cannot do so continuously. 4. A request arrives. The pool leases that connection and writes the request bytes. 5. The server, having closed, responds with RST - or the client reads EOF before any status line. The client sees `ECONNRESET`, `NoHttpResponseException`, `unexpected EOF`, `EOF occurred in violation of protocol`, or a similar name depending on the stack. The rate is proportional to how often connections sit idle near the server's timeout, which is why it clusters after quiet periods and at low traffic - and why it is often worst in staging or overnight. ## Why it cannot be fully prevented Between "the connection looks alive" and "the request has been written" there is always a window. TCP gives no way to atomically claim a connection, and the peer may close at any instant. Even a zero-byte liveness probe just before writing leaves microseconds of exposure. So the correct posture is: shrink the window, then handle the residual failure. ## Shrinking the window **Make the client close first.** The single most effective control. Every hop has an idle timeout: the application server, the reverse proxy, the cloud load balancer (60 seconds is a very common default), and any NAT gateway (which may silently drop flows after a few minutes without even sending a FIN - producing hangs rather than resets). Set the client pool's idle eviction comfortably below the smallest of these. If the client always closes first, the server never has to. **Validate after inactivity.** Clients such as Apache HttpClient expose a validate-after-inactivity setting: before leasing a connection idle longer than N milliseconds, poll the socket for readability - a readable socket with an EOF pending means the peer closed. This costs a syscall and closes most of the window, but not all of it. **Respect the server's hint.** A server sending `Keep-Alive: timeout=5` is telling clients when it will drop the connection; clients that honour it evict early. **Prefer close-on-response.** When a server wants to recycle a connection, sending `Connection: close` on a response is race-free, because it rides an in-band message. A silent idle close is not. HTTP/2's GOAWAY is the same idea, which is why multiplexed protocols suffer this far less. **Beware NAT and firewall idle drops.** Middleboxes that drop state without a FIN produce a worse variant: the client writes and simply waits for a response that can never arrive, failing only at the read timeout. TCP keep-alive probes (with a period below the middlebox's idle limit) or a low client idle timeout keep flows fresh. ## Absorbing the residual The key insight for retrying: if **no bytes of a response were received** and the failure happened on a connection reused from the pool, the request almost certainly was not processed - the server had already closed before it read anything. Most HTTP clients therefore retry such failures once automatically, on a **new** connection. Two constraints: - Restrict automatic retries to **idempotent** requests unless you know the endpoint deduplicates. A reset that arrives *after* the request was fully written but *before* the response is ambiguous only in principle here, but the discipline is worth keeping. - Retry once, not forever, and never retry a read timeout with the same reasoning - a read timeout means the server may well be processing the request. ## Confirming the diagnosis Look for: errors clustering after idle gaps and at low traffic; zero corresponding entries in the server's access log (the request was never read); a mean time-to-failure matching the server's keep-alive timeout; and disappearance of the problem when the client's idle timeout is lowered below the server's. A packet capture showing FIN from the server followed by the client's request and an RST is conclusive.
- Why is it acceptable to retry automatically in this specific failure, when retrying a timeout generally is not?Because the failure signature carries information: the connection was reused from the pool and zero bytes of a response were received, which means the server had already closed before reading the request. There is no plausible state where it processed the request and produced no output. A read timeout after the request was written is different - the server may be working on it - so retrying that risks duplicate side effects.
- The client's idle timeout is already below the server's, yet the errors persist at a low rate. What else should you check?Every hop has its own idle timeout: a cloud load balancer commonly at 60 seconds, a reverse proxy, and NAT gateways or stateful firewalls that silently drop flows without sending a FIN. The binding constraint is the smallest of them, and a NAT drop shows up as a hang until the read timeout rather than a reset. Also verify that validate-after-inactivity is enabled and that the client honours a Keep-Alive: timeout hint.
Reaching for a taxi you saw parked a minute ago: it drove off while you were not looking. You cannot check and get in at the same instant, so you shorten how long you leave it unwatched and accept that occasionally you must hail another.
saying these in an interview costs you the question
- Blaming the server or the network when the server's access log shows no request at all
- Believing the race can be eliminated by checking whether the socket is alive before writing
- Setting the client's idle timeout equal to or above the server's
- Retrying every failed request, including read timeouts, with the same non-idempotent method
- Disabling keep-alive entirely as the fix, trading a rare retry for a handshake on every request