skip to content

Go's http.Server has no max-connection setting: how do you cap an ingest tier that OOMs holding thousands of keep-alive connections?

level: seniorimportance: should knowfreq 38%

answer

  1. prove it before you tune it
  2. the server struct has no such field
  3. the cap belongs one layer down
  4. a blocked accept is not a refusal
  5. count idle versus active connections

basics

~20 s

Prove with a live-heap profile that the memory is per-connection buffering, not a handler leak, then cap connections at the listener, because http.Server has no such field. Also lower MaxHeaderBytes and count idle connections with the ConnState hook.

solid answer

~40 s

First prove it is connections, not a handler leak: in a live-heap profile the top entries should be the `bufio` read and write buffers `net/http` allocates per accepted connection. Then cap it — and since `http.Server` has no connection-count field, the cap sits under it, at the listener. Wrap your `net.Listener` so `Accept` stops handing out connections past N (`golang.org/x/net/netutil.LimitListener` does exactly this) and pass it to `srv.Serve`. Lower `MaxHeaderBytes` from the 1 MB default so the per-connection worst case is small. Use the `ConnState` hook to see the idle-versus-active split: behind a balancer most held connections are usually idle, which points the fix at the balancer's pool size. And know what you bought — a blocking accept cap makes overload look like latency, because excess connections queue in the kernel and fail there.

code

text · 7 lines
text
$ go tool pprof -inuse_space http://localhost:6060/debug/pprof/heap
(pprof) top
Showing nodes accounting for 1.42GB, 94.1% of 1.51GB total
      flat  flat%   sum%        cum   cum%
   0.71GB 47.0% 47.0%    0.71GB 47.0%  bufio.NewReaderSize
   0.68GB 45.1% 92.1%    0.68GB 45.1%  bufio.NewWriterSize
   0.03GB  2.0% 94.1%    0.03GB  2.0%  net/http.(*conn).readRequest

go deeper

for a junior

Take away the core fact: http.Server has no setting for how many connections it will hold, so any limit has to be applied to the listener you hand it.

for a middle

Be able to enumerate what one accepted connection costs — a goroutine and stack, read and write buffering, TLS state, a descriptor — and show how that multiplies into the instance's memory.

for a senior

Lead with the diagnosis. Read a live-heap profile to separate per-connection buffering from a handler leak, apply the cap at the right layer, and state what over-capacity will look like to callers once it binds.

for a principal

Own the posture question: queueing at the kernel versus explicit shedding changes what your callers and your balancer can observe, and that choice outlives whatever number you pick today.

## The shape of the failure An ingest endpoint sits behind a load balancer that keeps a pool of upstream connections open. Request rate is unremarkable. Memory climbs anyway, plateaus high, and eventually the instance is killed. This is the failure mode that keep-alives make possible: the number of connections a Go server holds is chosen by its callers, not by its own throughput. ## Step one: prove it is connections Do not start tuning. Take a heap profile of live memory and read the top entries. ``` (pprof) top flat flat% sum% 0.71GB 47.0% 47.0% bufio.NewReaderSize 0.68GB 45.1% 92.1% bufio.NewWriterSize ``` A profile dominated by `bufio` buffers allocated on the `net/http` serving path says: each accepted connection is carrying its own read and write buffering, and there are too many connections. That is a very different diagnosis from a profile dominated by your own types, which would say a handler is retaining data. The whole point of looking first is that the two failures feel identical from the memory graph and have no overlap in their fixes. ## What a connection costs Per accepted connection, roughly: * a serving goroutine with its own growable stack; * a read buffer and a write buffer on the order of a few kilobytes each; * TLS record buffers and session state, if the listener is encrypted; * whatever the handler holds while a request is in flight; * a file descriptor, and kernel socket buffers outside your heap entirely. None of it is large. Ten thousand of them is. ## Step two: cap it at the listener `http.Server` has **no field** that limits simultaneous connections. There is no `MaxConns`. The cap must live one layer down, in the `net.Listener` you hand to `Serve`: ```go ln, err := net.Listen("tcp", ":8080") … srv := &http.Server{Handler: mux} log.Fatal(srv.Serve(limited)) // limited wraps ln and stops accepting past N ``` `golang.org/x/net/netutil.LimitListener(ln, n)` is the standard wrapper: its `Accept` blocks once `n` connections are live and resumes when one closes. A hand-rolled version is a listener holding a buffered channel as a semaphore, acquiring before `Accept` returns and releasing when the wrapped connection is closed — the release-exactly-once detail is the part people get wrong. ## Understand what the cap does to clients This matters more than the mechanism. A blocking accept cap does not refuse anyone. Connections past the limit sit in the kernel's accept queue; when that fills, the kernel starts dropping or resetting them. From a client's point of view overload therefore looks like *connect latency, then an obscure network error* — not like a server saying no. Nothing in your process logs it, because your process never saw those connections. The alternative posture is to accept and shed: let connections in and answer over-capacity requests with an explicit status so the balancer can see the instance is saturated and route elsewhere. That costs the memory of accepting, but overload becomes legible. Which of the two you want is a real decision, not an implementation detail. ## Step three: shrink the per-connection worst case Lower `MaxHeaderBytes` from the 1 MB default to what your API genuinely needs. It does not change the steady-state cost of a well-behaved connection, but it bounds what a hostile one can make you buffer, which is what turns "thousands of connections" into an arithmetic you can defend rather than an open-ended risk. ## Step four: measure the idle/active split Wire the `ConnState` hook to counters and watch how many of the held connections are idle rather than serving. On an ingest tier behind a balancer, the honest answer is usually "almost all of them". That reframes the problem: you are paying memory for a pool somebody else sized. Shrinking the balancer's upstream pool, or spreading connections across more instances, may be a cheaper fix than anything inside your process — and it is a conversation you can only have with the number in hand. ## The last resort Turning keep-alives off entirely does collapse the connection count, and it is the wrong first move: it converts a memory problem into a handshake and CPU problem for every client, and hurts tail latency across the board. It is defensible only on genuinely tiny instances, which is roughly what the `net/http` documentation says about it. ## The answer in one breath "Heap profile first to prove it is per-connection buffering and not a handler leak; cap at the listener because `http.Server` has no connection field; lower `MaxHeaderBytes`; measure the idle/active split with `ConnState`; and decide deliberately whether over-capacity should look like queueing or like an explicit refusal."

  • How does the heap profile distinguish too many connections from a handler leak?
    By what dominates live bytes. Per-connection `bufio` read and write buffers allocated on the serving path mean the count of open connections is the problem. Your own types at the top mean a handler is retaining data, and no connection cap will help.
  • What does a client see when a blocking accept cap is reached?
    Nothing from your application. The connection waits in the kernel's accept queue and, once that fills, is dropped or reset. The client experiences slow connects and then a transport error, and your process logs nothing because it never accepted the connection.
  • Why might the right fix be at the load balancer rather than in the Go process?
    Because on an ingest tier the held connections are usually a pool the balancer sized, and most of them are idle. If the idle/active split says so, shrinking that pool or spreading it over more instances removes the memory cost without capping anything you own.
  • Why is disabling keep-alives the wrong first response to this?
    It works, but it trades a memory problem for a CPU and latency one: every request then pays a TCP and TLS handshake, an accept and a new goroutine. It is justified only on very resource-constrained instances, which is what the stdlib documentation reserves it for.

saying these in an interview costs you the question

  • Reaches for a max-connections field on http.Server that does not exist
  • Tunes GOGC or GOMEMLIMIT before profiling live memory
  • Assumes a blocking accept cap returns an error to clients
  • Disables keep-alives as the first fix rather than the last
  • Treats connection count as equal to concurrent request count