skip to content

A Go service's outbound requests fail with "cannot assign requested address" and the socket table is full of TIME_WAIT entries to one host. How do you diagnose it?

level: seniorimportance: should knowfreq 40%

answer

  1. The error is local, not the peer's
  2. Which side of a close holds that state?
  3. Count established against the churning state
  4. One flag per request tells you the truth
  5. Look for a Transport literal per call

basics

~20 s

That error is local ephemeral port exhaustion, so the client is not reusing connections. Look for an http.Transport built inside the request path, DisableKeepAlives, or a per-host idle pool smaller than your concurrency; httptrace's Reused flag proves it.

solid answer

~50 s

"cannot assign requested address" on a dial is the client running out of local ephemeral ports, and the `TIME_WAIT` pile-up says the client is the side closing connections — so almost every request is opening a new one. I would confirm the churn before changing anything: count `TIME_WAIT` versus `ESTABLISHED` to that address with `ss`, then instrument one call path with `net/http/httptrace` and log `GotConnInfo.Reused` and `WasIdle`, which turns the guess into a measurement. The Go-side causes, in order of likelihood: an `&http.Transport{}` or `&http.Client{...}` constructed per call so each has a private empty pool; `Transport.DisableKeepAlives` set; `MaxIdleConnsPerHost` left at its default of two under high concurrency; or a per-request client that never gets reused. The fix is one long-lived client per dependency with a sized idle pool. Raising the OS port range or enabling reuse knobs treats the symptom.

code

go · 7 lines
go
trace := &httptrace.ClientTrace{
	GotConn: func(info httptrace.GotConnInfo) {
		log.Printf("reused=%v wasIdle=%v", info.Reused, info.WasIdle)
	},
}
req = req.WithContext(httptrace.WithClientTrace(req.Context(), trace))
resp, err := client.Do(req)

go deeper

for a junior

Recognise that a dial failing with cannot assign requested address is about running out of local ports on your own machine, not about the remote service being down.

for a middle

Explain why the actively closing side holds TIME_WAIT, and name the client-side causes: a transport built per call, DisableKeepAlives, or an idle pool far smaller than the concurrency.

for a senior

Show a diagnosis sequence you would actually run under pressure: count socket states by destination, instrument one path with httptrace, identify the offending construction, then verify reuse afterwards.

for a principal

Decide what makes this class of defect impossible next time — a shared client convention, a reuse ratio exported as a metric, and a rule about who may tune kernel networking during an incident.

## Reading the symptom Two signals arrive together and they say the same thing from different ends. `connect: cannot assign requested address` (Linux `EADDRNOTAVAIL`) on an outbound dial means the kernel could not find a free local port for the four-tuple. Because a client socket's tuple is (local IP, local port, remote IP, remote port), and here the remote pair is fixed, the usable space is the ephemeral port range — commonly around 28,000 ports. `TIME_WAIT` is the state the **actively closing** side holds for roughly two maximum segment lifetimes so that late duplicate segments cannot be mistaken for a new connection. Tens of thousands of `TIME_WAIT` sockets to one address means your process is opening and closing connections to that address at a very high rate — which is the definition of not pooling. So the diagnosis is not "we need more ports". It is "why is each request getting its own connection?" ## Confirm before you change anything Count the states, so you know the ratio rather than the anecdote: ```text ss -tan state time-wait dst 10.0.4.12 | wc -l ss -tan state established dst 10.0.4.12 | wc -l ``` A healthy pooled client shows a stable, smallish number of established connections and almost no `TIME_WAIT` to that host. A non-pooling client shows the inverse, and the `TIME_WAIT` count tracks your request rate times sixty seconds. Then prove it from inside the program. `net/http/httptrace` gives a per-request hook that reports exactly what the transport did: ```go trace := &httptrace.ClientTrace{ GotConn: func(info httptrace.GotConnInfo) { log.Printf("reused=%v wasIdle=%v idle=%v", info.Reused, info.WasIdle, info.IdleTime) }, } req = req.WithContext(httptrace.WithClientTrace(req.Context(), trace)) ``` If `Reused` is false on every request, connection reuse is broken and the rest of the investigation is a code search. This is worth wiring behind a debug flag permanently; it is the difference between believing you pool and knowing it. ## The Go-side causes, in order 1. **A transport constructed per call.** `&http.Client{Transport: &http.Transport{}}` inside a helper or a handler gives every request a private, empty idle list. This is the classic cause and the one the symptom fits best, especially in an SDK where a `NewClient` helper is called per operation instead of per process. 2. **`Transport.DisableKeepAlives = true`.** Sometimes set during a past debugging session and never removed. It closes every connection after one exchange by design. 3. **`MaxIdleConnsPerHost` at its default of two** while running hundreds of concurrent requests to one host. You keep two warm and close the rest, so most requests still dial. 4. **Connections closed by the peer.** A server or intermediary with a short idle timeout, or one that sends `Connection: close`, will also produce churn — but then it is the *server* that closes and the `TIME_WAIT` sockets accumulate on its side, not yours. That asymmetry is a useful discriminator: `TIME_WAIT` on the client host points at client-side behaviour. 5. **Too many distinct destinations.** Pools are keyed per destination; if you address a downstream by rotating IPs or by unique hostnames, no pool ever warms up. ## The fix One long-lived `*http.Client` per dependency, held in package or struct state, with a cloned transport whose `MaxIdleConnsPerHost` matches sustained concurrency and whose `MaxIdleConns` is raised to match. In an SDK, that means the client is created in the constructor and reused for every call, never per operation. After the change, the same two measurements confirm it: `Reused` becomes true for the overwhelming majority of requests, `TIME_WAIT` collapses, and established connections settle at roughly your concurrency level. ## What not to reach for first Widening `net.ipv4.ip_local_port_range`, enabling `tcp_tw_reuse`, or setting `SO_REUSEADDR`-style options will buy headroom and will hide the defect. They are legitimate as a stopgap during an incident, and they are the right tool when the churn is genuinely unavoidable — but a service that burns 28,000 ports a minute against one host is nearly always a service that forgot to share its client. Fix the pooling, then decide whether the OS tuning is still needed. ## The review lesson This defect is invisible at low traffic: a per-call transport works perfectly in tests and in staging, and only fails when concurrency and duration multiply into port pressure. That makes it a code-review rule rather than a testing rule — no `http.Transport` literal below a constructor — backed by a `Reused` metric so the property is observable in production rather than assumed.

  • How would you tell client-side churn from a server that closes every connection?
    Look at which host accumulates `TIME_WAIT`. The actively closing side holds it, so a pile-up on your host points at your client opening and dropping connections, while a pile-up on the server's side suggests it is sending `Connection: close` or enforcing a short idle timeout. `httptrace`'s `Reused` flag plus a packet capture settles it.
  • Would raising the ephemeral port range fix this?
    It buys headroom and can stop an incident, but it does not fix the cause: you are still paying a TCP and TLS handshake per request, still burning CPU on crypto, and still one traffic increase away from the same failure. Treat OS tuning as a stopgap and land the shared client.
  • Why does this defect survive staging and only fail in production?
    Port exhaustion needs rate multiplied by the roughly one-minute `TIME_WAIT` window. Staging traffic rarely reaches the thousands of connections per second that exhaust a 28,000-port range, so a per-call transport behaves correctly there. That is why the guard is a review rule plus a reuse metric, not a test.

saying these in an interview costs you the question

  • Blames the downstream server for a local dial error
  • Jumps straight to widening the ephemeral port range
  • Thinks TIME_WAIT accumulates on the passively closing side
  • Adds retries, which multiply the connection churn
  • Cannot name a way to prove connections are being reused
  • Assumes the garbage collector will close abandoned connections promptly