skip to content

In a Go daemon whose open socket count climbs every push cycle, how do you tell an unclosed resp.Body from an undrained one?

level: seniorimportance: should knowfreq 46%

answer

  1. same symptom, two different shapes
  2. does the count grow, or does it churn?
  3. one bug holds sockets, the other burns them
  4. ask the Transport whether it reused anything
  5. GotConnInfo.Reused false on every request

basics

~20 s

Look at the shape, not the count. A body never closed pins one live connection per request, so sockets grow without bound. A body closed before it was read releases each socket but reuses none, so you see constant re-dialling.

solid answer

~40 s

Both bugs raise the socket count, but they have different shapes. An unclosed `resp.Body` pins a live connection per request: the count grows monotonically with total requests and never recovers while the process runs. A body closed without being read releases each socket, so the count churns instead, with a fresh dial and TLS handshake per push. In code, check whether `Close` runs on every path — the early `return` on a non-2xx status is the classic miss — and whether anything ever reads to EOF. `net/http/httptrace` settles it cheaply: `GotConn` reports `Reused`, so always-false means the undrained case, while a socket count that only ever climbs means the unclosed one. The fixes differ: close on every path for the first, a bounded drain before closing for the second.

code

go · 11 lines
go
var newConns atomic.Int64

trace := &httptrace.ClientTrace{
	GotConn: func(info httptrace.GotConnInfo) {
		if !info.Reused {
			newConns.Add(1)
		}
	},
}
req = req.WithContext(httptrace.WithClientTrace(req.Context(), trace))
resp, err := client.Do(req)

go deeper

for a junior

Know the two distinct mistakes by name - a response body that is never closed, and one closed before it was read - and that the first holds connections open while the second throws them away.

for a middle

Explain the mechanism behind each: no finalizer reclaims an unclosed body, and the Transport will not pool an HTTP/1 connection whose payload was left unread. Be ready to point at the early return on a non-2xx status as the usual place the close is missed.

for a senior

Separate the two before changing anything: monotonic growth versus connection churn, confirmed with httptrace GotConn reporting Reused, or a goroutine profile full of Transport read loops. Then apply the matching fix and bound the drain rather than swallowing a large body.

for a principal

Decide how this class of bug stops recurring: one reviewed helper that every outbound call goes through, a connection-reuse metric you can alert on, and a clear position on whether raising resource ceilings is ever an acceptable response to it.

### Two bugs with one symptom A long-running daemon that pushes samples to a remote endpoint on a timer is the ideal shape for both of these mistakes, because it makes the same call thousands of times an hour and the response is something nobody looks at. Two different defects produce "our socket count is going up": **Bug one: the body is never closed.** Somewhere there is a path — usually the early return on a bad status code, sometimes a decode error, sometimes a request built in a helper that returns the response to a caller who forgets — where `resp.Body.Close()` never runs. Each of those requests permanently retains a connection. Nothing reclaims it: there is no finalizer on a response body. **Bug two: the body is closed but never read.** The daemon checks `resp.StatusCode`, closes, and moves on. Every connection is released correctly, but on HTTP/1 the Transport cannot pool a connection whose remaining payload was never consumed, so it discards it and the next push dials again. ### Telling them apart from the outside Count and shape, not count alone. * **Unclosed** looks monotonic. The number of established connections from the process to that endpoint rises roughly in step with the number of requests and never comes back down. Restarting is the only thing that clears it. Steady-state memory creeps up alongside it because each connection carries buffers and bookkeeping. * **Undrained** looks like churn. Live connection count stays low — often around one per concurrent request — but the *rate* of new connections tracks the request rate exactly, and sockets pile up in the kernel's post-close states rather than in the process. The process-level socket count fluctuates instead of climbing to the ceiling. A snapshot taken twice, ten minutes apart, distinguishes them immediately: if the number roughly doubled with load, it is the leak; if it is the same number and the *connection establishment* rate is what is high, it is the reuse failure. ### Telling them apart from the inside `net/http/httptrace` answers the reuse question directly and cheaply, in the program itself: ```go trace := &httptrace.ClientTrace{ GotConn: func(info httptrace.GotConnInfo) { if !info.Reused { newConns.Add(1) } }, } req = req.WithContext(httptrace.WithClientTrace(req.Context(), trace)) ``` `GotConnInfo.Reused` is true when the Transport took the connection out of its idle pool. If a healthy client is doing sequential pushes to one host and `Reused` is false every single time, the connection is not surviving the request, and the body handling is the first place to look. A goroutine profile is the complementary check for the unclosed case: leaked HTTP/1 connections keep their reader goroutines alive, so a growing count of goroutines parked in the Transport's read loop points at bodies that were never closed. ### The fixes are not the same For the unclosed body: make the close unconditional and structural, not something each caller remembers. Arrange the request in one place, `defer` the close immediately after the error check, and check status afterwards. If the function hands the response to a caller, document who owns the close — or better, do not hand a live body across a package boundary at all; read what you need and return that. For the undrained body: drain before closing, but bound it. ```go defer func() { _, _ = io.Copy(io.Discard, io.LimitReader(resp.Body, 64<<10)) resp.Body.Close() }() ``` The bound matters here more than usual. The endpoint in this scenario returns a large buffer nobody reads; transferring megabytes on every push to save one handshake is a worse deal than the handshake. Cap the drain, take the reuse when the leftover is small, and pay for a re-dial when it is not. ### What recent Go changes and what it does not On Go 1.27, closing an HTTP/1 response body drains a bounded amount automatically so the connection can still be pooled. That quietly fixes the undrained case for small bodies, which means an upgrade can make this class of problem disappear — and it also means the version of Go a service is built with is now part of the diagnosis. It does nothing for a body that is never closed, and nothing for a remainder larger than the bound. ### The judgment behind the fix The reason to separate the two before touching code is that they justify different work. The unclosed body is a correctness bug that will eventually take the process down and deserves an immediate, structural fix. The undrained body is a cost problem: it wastes handshakes and ephemeral ports and adds latency, and how much you should spend to fix it depends on the request rate and on how big the wasted remainder is. Someone new to the codebase who sees "sockets climbing" and reaches for a bigger descriptor allowance has treated the second bug as the first, and postponed rather than fixed either.

  • Raising the process's file-descriptor allowance makes the alert stop. Is that a fix?
    Only for the reuse case, and only by accident. If bodies are never closed the growth is unbounded, so a larger allowance just moves the failure later and makes the eventual outage happen at a worse time. Fix the ownership of the close first, then decide whether the ceiling was genuinely too low for the real connection count.
  • The daemon pushes over HTTP/2. Does the undrained case still show up?
    Not in the same way. HTTP/2 multiplexes independent streams over one connection, so closing a body early resets that stream and the connection survives. If the endpoint is HTTP/2 and connections are still not being reused, look elsewhere — at connection-level errors, or at whether the client itself is being rebuilt per push. The unclosed-body case still leaks, because the stream stays open.
  • What signal in a goroutine profile points at bodies that were never closed?
    A steadily growing population of goroutines parked in `net/http`'s connection read loop. Each live HTTP/1 connection has one, so if that count tracks the number of requests made rather than the number in flight, connections are being retained — which for a client almost always means response bodies that were never closed.

saying these in an interview costs you the question

  • Treats every rising socket count as the same bug
  • Raises the descriptor allowance and calls it fixed
  • Assumes closing the body guarantees the connection is reused
  • Drains an unbounded response just to keep one connection
  • Never checks whether the Transport reused anything
  • Looks only at the happy path and misses the early return