Your Go reverse proxy hits "too many open files" an hour into peak. How do you find the cause?
answer
- two different curves, two different bugs
- does the count fall when traffic falls?
- measure reuse, not just the count
- one hook tells you if a connection came from the pool
- closed is not the same as drained
basics
~20 sChart the process's open descriptor count against request rate and against the descriptor limit it actually has. A count that rides concurrency means no ceiling on outbound connections; a count that only climbs means a response body on some path is never closed.
solid answer
~50 sFirst separate the two shapes. Plot open descriptors over the hour: a count that tracks concurrency and falls back with traffic is a **capacity** problem — the outbound transport has no `MaxConnsPerHost`, so in-flight requests each hold a socket. A count that only climbs and never returns is a **leak** — some path obtained a response and never closed its body. Then prove whether connections are being reused, which is Go-specific evidence: attach a `httptrace.ClientTrace` with a `GotConn` hook and count how many `GotConnInfo` values arrive with `Reused` true. If reuse has collapsed, look at what changed — typically an early return around a non-200 upstream response that closes the body without draining it, so the connection is dropped instead of returning to the idle pool. Fix by draining a bounded amount before `Close` and setting `MaxConnsPerHost` so the proxy queues rather than exhausting descriptors.
code
go · 12 linesvar reused, fresh atomic.Int64
trace := &httptrace.ClientTrace{
GotConn: func(info httptrace.GotConnInfo) {
if info.Reused {
reused.Add(1)
} else {
fresh.Add(1)
}
},
}
req = req.WithContext(httptrace.WithClientTrace(req.Context(), trace))go deeper
Know that an HTTP response body must be closed on every path, including error returns, and that leaving it open keeps a socket alive. Recognise the error text as the process running out of descriptors.
Explain why a connection whose body was not read to the end cannot go back into the idle connection pool, and how losing reuse turns into more simultaneous sockets under load.
Show the diagnosis in order: shape of the descriptor curve, the limit actually in force, then the reuse ratio from a client trace hook. Then fix both the drain and the missing per-host connection ceiling.
Decide what the service's descriptor budget is and make it observable, so the alert fires below the ceiling. Own the trade that a connection cap turns a total outage into per-upstream latency someone must accept.
## The scenario A reverse proxy fans each inbound request out to several upstream services. It has been fine for months. A change ships, and an hour into the daily peak it starts logging `dial tcp: socket: too many open files`; requests fail in bulk; a restart clears it and it comes back the next day. This is the classic Go descriptor incident, and it has a repeatable diagnosis. ## Step 1: confirm it is descriptors, and against what ceiling Count the process's open descriptors and compare with the `RLIMIT_NOFILE` the process actually holds. Enumerating *which* handles are open is operating-system work — on Linux the per-process descriptor directory, or standard host tooling — and it answers "sockets or files, and to whom". The Go-side question is the one worth your time: is the number bounded by design or not? One trap here. On Unix the Go runtime raises the soft `RLIMIT_NOFILE` to the hard limit at startup, so the soft limit you set in a shell profile is probably not the number in force; the hard limit, or the container's configured limit, is. Read the limit from the running process, not from the deployment manifest. ## Step 2: read the shape of the curve The curve over the hour tells you which of two different bugs you have. - **Rides concurrency, falls when traffic falls, but the peak is higher every day** — the count is not bounded. Outbound connections scale one-per-in-flight-request because `MaxConnsPerHost` is unset (zero means unlimited), and a slow upstream lengthens how long each connection lives. Nothing is leaked; the design simply has no ceiling, and peak traffic finally exceeded the limit. - **Climbs monotonically and never falls, even when traffic drops overnight** — something is held forever. In a proxy, that is virtually always an HTTP response body that some path never closes: the body keeps its connection alive, and the connection is a socket. A `restart clears it, it comes back` pattern fits both, so the curve, not the symptom, is what separates them. ## Step 3: prove whether connections are being reused This is the step people skip, and it is where Go gives you a precise instrument. `net/http/httptrace` lets you hook the client's internals per request. `ClientTrace.GotConn` fires when the transport has a connection for the request, and the `GotConnInfo` it receives carries `Reused` (did this come from the idle pool?), `WasIdle` and `IdleTime`. Counting `Reused` true against false gives you a reuse ratio, and that ratio is the single most diagnostic number in this incident: healthy fan-out traffic reuses nearly everything, and a reuse ratio near zero means every request is dialling a brand-new socket. If reuse has collapsed while the code did not change its request pattern, the cause is on the response side. ## Step 4: the two ways a body ruins you **Not closed at all.** The response body holds its connection open indefinitely. This is a true leak: the descriptor is never released, the curve only climbs, and no amount of limit-raising helps. The fix is the standard shape — `defer resp.Body.Close()` immediately after the error check, so no later branch can skip it. **Closed but not drained.** This is the subtler one, and it is what a code reviewer should be looking for on a newly added early-return path: ```go resp, err := upstream.Do(req) if err != nil { return err } defer resp.Body.Close() if resp.StatusCode != http.StatusOK { return fmt.Errorf("upstream %s: %s", req.URL.Host, resp.Status) } ``` The body is closed, so the descriptor is released — this is not a leak in the strict sense. But a connection whose body was not read to the end cannot be returned to the idle pool, because the transport does not know where the next response would begin on that stream. So it is closed instead. Reuse collapses for exactly the responses your error path handles, which under an upstream degradation is *most* of them. Now every request dials, `MaxConnsPerHost` is unset, in-flight connection count rides concurrency, and an hour into peak the process runs out of descriptors. Closing-without-draining does not consume the descriptor by itself; it removes the mechanism that was keeping the count small. ## Step 5: the fix, in two parts Drain a bounded amount before closing, so ordinary short error bodies stay poolable while a huge one does not turn into an unbounded read: ```go defer func() { _, _ = io.Copy(io.Discard, io.LimitReader(resp.Body, 64<<10)) resp.Body.Close() }() ``` And give the transport a ceiling — `MaxConnsPerHost` per upstream, plus a `MaxIdleConnsPerHost` large enough that steady-state traffic is served from the pool. The ceiling matters more than the drain: it converts a descriptor exhaustion, which takes the whole process down for every upstream at once, into queueing latency on the one upstream that is misbehaving. ## Step 6: make the next one boring Export the open-descriptor count as a metric and alert well below the limit, so the next capacity growth is a ticket rather than an outage. Record the reuse ratio too — it moves before the descriptor count does, which makes it the earlier signal.
- The descriptor count rides concurrency and falls overnight. Is that a leak?No — a leak does not give the descriptors back. That curve says the count is simply unbounded: with `MaxConnsPerHost` unset, in-flight requests each hold their own socket, so peak sockets equal peak concurrency times the number of upstreams. The bug is the missing ceiling, and the fix is to set one so requests queue instead of dialling without limit.
- You set MaxConnsPerHost and the descriptor errors stop, but p99 latency doubles. Did you fix it?You traded one failure mode for a visible one, which is progress but not a fix. Requests are now queueing for connections, which means the upstream cannot keep up with the concurrency you are sending it. Next you either raise the cap deliberately against a measured descriptor budget, shed load, or fix the upstream — but you are no longer taking the whole process down.
- Why is the reuse ratio a better early warning than the descriptor count?Because it moves first. Reuse collapses the moment a code path stops returning connections to the idle pool, whereas the descriptor count only becomes alarming once traffic is high enough for the unpooled dialling to matter. A reuse ratio that drops from near one to near zero after a deploy points straight at the change, hours before the limit is reached.
saying these in an interview costs you the question
- Raises the descriptor limit and calls the incident closed
- Cannot distinguish a climbing curve from one that tracks load
- Claims closing an undrained body leaks the descriptor forever
- Never checks whether connections are actually being reused
- Reads the limit from the manifest, not the running process