Why does http.Client.Do sometimes fail with a bare EOF after an idle period?
answer
- it only happens after a quiet stretch
- two sides, two idle clocks, no handshake
- the message has no address in it
- you wrote into a socket already closing
- reuse of a connection the peer had dropped
basics
~20 sThe transport reused an idle keep-alive connection that the far side had already closed. The request went into a dead socket and nothing came back, so the cause under the *url.Error is io.EOF. It is a race, not a partner outage.
solid answer
~50 sThat is the stale pooled connection race. Go's transport keeps idle connections per host and reuses them; meanwhile the server, or a proxy or load balancer between you, closes connections it considers idle. If your request is written just as the close arrives, the client sends bytes into a socket that is already going away and reads nothing back, so the cause wrapped in the `*url.Error` is `io.EOF` and the message is just `Get "https://api.example.com/v1/orders": EOF` with no syscall detail. The transport heals many of these for you: it retries on a fresh connection when the connection had already been used successfully, no response bytes arrived, and the request is replayable - an idempotent method with no body, or one whose `Request.GetBody` is set. When it cannot rebuild the request, the EOF reaches your code. The fingerprint is a failure on the first call after an idle gap that succeeds instantly on retry.
code
text · 1 lineGet "https://api.example.com/v1/orders": EOFgo deeper
Be ready to say that EOF here means the connection ended before any response arrived, and that HTTP clients reuse connections rather than dialling fresh for every request.
An interviewer expects the mechanism: an idle connection taken from the pool that the peer had already closed, the resulting io.EOF under the *url.Error, and the fact that errors.Is sees through that wrapper.
Show the diagnosis. Name the fingerprint - first call after an idle gap, instant success on retry, bare EOF text, no correlation on the peer's side - and state the transport's own retry rule rather than reaching for a blind retry loop.
Own where the fix belongs: pooling configuration across every client in the fleet versus per-call-site handling, and the standard your services follow so an idle-close race never reads as a partner outage in an incident review.
## The race Go's HTTP transport maintains a pool of idle connections per host and reuses them, because a fresh TCP dial plus TLS handshake per request is expensive. The far end also has an opinion about idle connections: an origin server, a reverse proxy, or a load balancer sitting between you will close one after its own idle window, and it does so unilaterally. Those two clocks are independent, and there is no protocol handshake for 'I am about to close this'. So there is always a window in which the client picks a connection out of its pool, writes a request into it, and the peer's FIN is already on the way. The client writes successfully - the local socket accepts the bytes - then reads and gets end-of-stream. There is nothing else to report, so the wrapped cause is `io.EOF` and the message is bare: ``` Get "https://api.example.com/v1/orders": EOF ``` No `connect:`, no address, no syscall - that bareness is the fingerprint. ## Why you only sometimes see it The transport tries to hide this. Its documented rule is that it retries a request after a network error only when the connection had already been used successfully, and the request is idempotent and either has no body or has `Request.GetBody` defined. Idempotent here means `GET`, `HEAD`, `OPTIONS` or `TRACE`, or a request carrying an `Idempotency-Key` (or `X-Idempotency-Key`) header. And it only retries when no bytes of the response have arrived yet - once a response has started, retrying could double-apply work. So a plain `GET` usually heals invisibly, and you never learn the race happened. A `POST` built from an arbitrary `io.Reader` body cannot be rebuilt, so the transport gives up and hands you the EOF. That asymmetry is why the same job sees the error on some endpoints and never on others. ## Recognising it in a postmortem The shape is distinctive, and it is worth being able to state it from memory: - It hits the **first request after an idle gap** - the opening call of a nightly reconciliation run, the first request after a lull, the first call of each burst. - It **succeeds immediately** on a retry, with no backoff needed. - The error text is **just `EOF`**, with no address or syscall wording. - It does **not correlate** with anything on the peer's side: their logs show no request, their error rate is flat, their latency is flat. Contrast that with a real outage. A dead peer gives you `connect: connection refused` or `no such host`, it persists across attempts, and the peer's own metrics agree with you. A saturated peer gives you 5xx statuses with a nil error, not an EOF. Getting this distinction wrong is how a team spends a morning accusing a partner of instability. ## What to do about it Three responses, in order of how much they belong at the call site: **Recognise it.** `errors.Is(err, io.EOF)` matches straight through the `*url.Error` wrapper, because `url.Error` implements `Unwrap`. Recording that classification separately from timeouts and refused dials is what lets the next person see the pattern rather than a wall of identical log lines. **Make the request replayable.** If the body comes from a `*bytes.Reader`, `*bytes.Buffer` or `*strings.Reader`, `http.NewRequest` sets `Request.GetBody` for you, and the transport can rebuild and re-send the request itself. A body streamed from an arbitrary reader cannot be rebuilt, so the error surfaces to you instead. **Fix it in the pool, not the call site.** The durable cure is making the client's own idle-connection lifetime shorter than whatever closes connections upstream, so the client discards a connection before the peer does. That is a transport-pooling decision rather than an error-handling one, and it is where the permanent fix lives. ## The one thing not to do Do not answer this by wrapping every call in a blind retry-everything loop. It hides the diagnosis, it turns a benign race and a real outage into the same event, and on a non-idempotent request it is a correctness problem rather than a resilience one.
- How do you tell this apart from the partner actually being down?By shape. A stale-connection EOF hits the first request after an idle gap, succeeds instantly on retry, carries no address or syscall text, and does not correlate with the peer's error rate. A real outage gives connection refused, DNS failures or 5xx statuses, persists across attempts, and shows up in the peer's own metrics too.
- Why does the request's body change whether you see the EOF at all?The transport can only retry on a fresh connection if it can rebuild the request. http.NewRequest sets Request.GetBody automatically for a *bytes.Reader, *bytes.Buffer or *strings.Reader body, so those are replayable and heal invisibly. A body streamed from an arbitrary io.Reader cannot be replayed, so the transport gives up and the EOF reaches your code.
- Where does the durable fix live?In the transport's pooling configuration rather than at the call site: the client's idle-connection lifetime has to be shorter than whatever closes connections upstream, whether that is the origin server or a load balancer in between. At the call site you can only recognise the error and make requests replayable so the transport heals them for you.
The pool hands you a phone line the other side hung up on a minute ago. You speak into it perfectly well and hear nothing back - which is exactly what EOF means.
saying these in an interview costs you the question
- Blames the partner for an EOF on the first request
- Thinks EOF means the response body was empty
- Assumes Go never retries anything on its own
- Wraps every call in a blind retry-everything loop
- Reads EOF as a JSON parsing failure