Your Go client sets DialContext, TLSHandshakeTimeout and ResponseHeaderTimeout, yet one fetch hangs for an hour. Why?
answer
- the transport stops at the headers
- one row of the table has no knob
- a slow sender is not a dead socket
- total time for calls, progress for streams
- only the client field or the context covers the body
basics
~20 sThose three fields bound connecting, the TLS handshake and the wait for response headers. Nothing on http.Transport bounds reading the response body, so a server that sends headers and then trickles bytes holds you indefinitely.
solid answer
~50 sEvery timeout on `http.Transport` stops at the response headers. `net.Dialer.Timeout` behind `DialContext` bounds DNS and connect, `TLSHandshakeTimeout` bounds the handshake, and `ResponseHeaderTimeout` bounds the gap between finishing the request and the first response header arriving. Once headers are in, the transport stops policing the exchange — a server that answers `200 OK` immediately and then dribbles the body one byte a minute is holding a goroutine, a connection and a file descriptor for as long as it likes. The fix depends on the workload: for bounded responses add `Client.Timeout` or a request context deadline, both of which cover the body read; for legitimately long streams keep the phase timeouts and enforce *progress* rather than total time, by resetting a deadline on each read or having a watchdog close the body when nothing has arrived for N seconds. The phase fields are a supplement to an overall budget, not a replacement for one.
code
go · 8 linestr := &http.Transport{
DialContext: (&net.Dialer{Timeout: 2 * time.Second}).DialContext,
TLSHandshakeTimeout: 2 * time.Second,
ResponseHeaderTimeout: 3 * time.Second,
}
// Nothing here bounds reading resp.Body.
client := &http.Client{Transport: tr}go deeper
Learn the phase list — connect, handshake, headers, body — and remember that the transport's fields cover the first three only. That single fact prevents the most common hang.
Explain exactly what span each field measures, including that ResponseHeaderTimeout starts after the request is written, and why nothing at the socket layer turns a slow sender into an error.
Diagnose the hang from its symptoms — goroutines parked in Read, descriptors climbing, no errors logged — and choose between an overall budget and a progress watchdog based on whether the response is bounded or a stream.
Set the house rule that every hand-built transport must answer "what bounds the body", and decide whether streaming clients get their own construction path so a total-time budget is never applied to them by accident.
## The phases and who bounds them An outbound call in Go passes through stages, and `http.Transport` gives you a separate knob for some of them: | Stage | Bounded by | | --- | --- | | DNS resolution and TCP connect | the `Timeout` on the `net.Dialer` you install as `DialContext` | | TLS handshake | `Transport.TLSHandshakeTimeout` | | Waiting after sending `Expect: 100-continue` | `Transport.ExpectContinueTimeout` | | Writing the request | nothing on the transport | | Waiting for response headers | `Transport.ResponseHeaderTimeout` | | Reading the response body | **nothing on the transport** | That last row is the whole question. The transport's job ends when it hands you a `*http.Response`; the body is an `io.ReadCloser` streaming off a connection, and no transport field watches how long you spend on it. ## Why a hang looks like a hang and not an error A server that stops sending mid-body does not close the connection — the TCP session is alive, so your `Read` simply blocks. TCP keep-alives, if enabled, take minutes to hours to notice a dead peer, and they notice a *dead* peer, not a slow one. A deliberately slow sender is indistinguishable from a healthy one at the socket layer. So the symptom is a goroutine parked in `Read` forever, a connection and file descriptor held, and, if this is a fan-out or a loop, a steady accumulation of both until the process hits its descriptor limit or its memory ceiling. ## The fixes, in order of preference **1. An overall budget.** If the response is bounded — a JSON document, a config blob, a few megabytes of report — set `http.Client.Timeout`, or better, put a deadline on the request context. Both cover the body read: the transport watches the context for the life of the exchange, so a blocked `Read` returns an error when the deadline passes. This is the answer nine times in ten, and it is why the transport-phase fields alone are a half-configured client. **2. Phase timeouts plus progress enforcement.** If the response is legitimately long — a large download, a long-poll, a server-sent event stream — an overall budget is wrong, because it kills healthy calls at an arbitrary size. Keep the dial, handshake and header timeouts, which fail fast on a dependency that is not answering at all, and bound *progress* instead: - Wrap `resp.Body` in a reader that resets a deadline on every successful `Read`, and have the expiry close the body. Closing the body from another goroutine unblocks the read. - Or capture the `net.Conn` in your own `DialContext` and reset a read deadline per operation. The rule of thumb: **total-time budgets for request/response calls, progress budgets for streams.** Confusing the two produces either eternal hangs or downloads that fail at exactly the size where they became useful. **3. Do not reach for `ResponseHeaderTimeout` as a general-purpose timeout.** It is precise and useful — it isolates "the dependency is not answering" from "the dependency is answering slowly" — but it is set on the transport, so it applies to every request through that transport and cannot be varied per call. ## What `ResponseHeaderTimeout` actually measures It starts when the request has been fully written and ends when the response headers have been read. That makes it the closest thing Go has to a "server think time" limit, and it is the right knob when you want to abandon a dependency that has gone silent while still allowing a big body afterwards. Note that it does not cover the write of a large request body, and it does not restart across redirects — each hop gets its own header wait, which is one more reason an overall budget still belongs on top. ## Reviewing a client When you see a hand-built `http.Transport` in a diff, ask one question: *what bounds the body?* If the answer is "nothing", either the client also needs a `Timeout`, or every call site must be passing a request context with a deadline, or the code is streaming and there had better be a progress watchdog. A transport with three carefully tuned phase timeouts and no answer to that question is the configuration that produces the hour-long hang.
- Does a request context deadline actually interrupt a blocked body read, or only the connect phase?It interrupts the read. The transport watches the request's context for the whole exchange and tears the connection down when it is done, so a goroutine parked in `resp.Body.Read` returns an error rather than waiting on the socket. That is the key difference from the transport's phase fields, which stop caring once headers arrive.
- When is ResponseHeaderTimeout the right knob instead of an overall Client.Timeout?When the body is legitimately long — a large download, a long-poll, an event stream. You still want to fail fast if the dependency never starts answering, but an overall budget would kill healthy transfers at an arbitrary size. Bound the wait for headers, then enforce progress on the body separately.
- How do you bound a stream you intend to read for minutes?Bound progress, not total time. Wrap the body in a reader that resets a deadline on each successful read and closes the body when nothing arrives for N seconds; closing from another goroutine unblocks the parked `Read`. A watchdog on bytes-per-interval catches the slow-trickle attack that a total budget only catches late.
saying these in an interview costs you the question
- Assumes ResponseHeaderTimeout covers the whole response
- Thinks the dialer's timeout applies end to end
- Expects TCP keep-alives to catch a slow sender
- Sets an overall budget on a client used for large downloads
- Says a hung read will eventually return an error on its own