skip to content

Who owns the connection cap on a memory-bound Go http.Server fleet, and how do you set the number?

level: principalimportance: should knowfreq 30%

answer

  1. it is an agreement, not a setting
  2. measure bytes per connection first
  3. connections and requests are two numbers
  4. silent queueing versus visible refusal
  5. one authoritative place, one owner

basics

~20 s

Derive the cap from measured bytes per open connection against the instance's memory budget, decide whether over-capacity queues silently or is refused visibly, and name one authoritative place it lives, with an owner and an alarm before it binds.

solid answer

~50 s

Treat the cap as a budget, not a setting. Measure bytes per open connection at a known connection count — the `ConnState` hook gives you the count, a live-heap profile the bytes — and divide the instance's memory budget by it, leaving headroom for in-flight request work and for the collector. Then make three decisions that are yours. Posture: a blocking accept cap turns overload into connect latency nobody can see, while accepting and shedding explicitly lets the balancer route away, and legibility is usually worth the memory. Location: the same cap can live in your process, the balancer's upstream pool and the platform's memory limit, and if all three exist none is authoritative — pick one. Escalation: the capacity owner defends the number, the service owner overrules it when it refuses legitimate load, and you alarm before it binds.

go deeper

for a junior

Take away that a connection limit is chosen from measured memory per connection, not guessed, and that it belongs to someone who has to answer for the instance's memory.

for a middle

Be able to explain how you would obtain the inputs: connection counts by state from the server's own hook, and per-connection bytes from a live-heap profile at a known count.

for a senior

Argue the operational posture. Show why silent queueing hides saturation from the balancer and from dashboards, and what you would alarm on so the cap is discussed before it binds.

for a principal

Own the whole agreement: where the cap is authoritative when three layers can express it, who defends the number, who may override it during an incident, and what the fleet-cost versus availability trade actually buys.

## Why this is a decision and not a config value A per-instance connection cap looks like a number in a config file. It is really an agreement between two people who are optimising different things: the person sizing the fleet wants the cap low enough that an instance cannot exhaust its memory, and the person who owns the service wants it high enough that no legitimate caller is ever made to wait. Those pull in opposite directions and neither can win outright. Making that explicit is most of the work. ## Get the number from measurement The honest derivation: 1. Instrument connection counts by state, so you know how many connections an instance holds and how many are idle rather than serving. 2. At a known count, take a live-heap profile and attribute memory to per-connection structures — the read and write buffering each accepted connection carries, serving-goroutine stacks, TLS state — as distinct from handler working set. 3. Divide the instance's memory budget by bytes-per-connection, then subtract headroom: for the peak concurrent *request* work on top of the connections, and for the collector, which needs slack to run without thrashing. That gives a cap you can explain in one sentence, which is the property that matters when somebody asks to raise it at 3am. ## Two numbers, not one Connections and concurrent requests are different quantities and reuse is what separates them. Thousands of idle upstream connections held by a balancer cost memory and almost no CPU; a few hundred active requests cost CPU and allocate. A fleet sized on connection count alone over-provisions compute; one sized on request rate alone gets surprised by memory. Carry both numbers, and know the idle/active ratio in normal traffic — it is also the fastest signal that the balancer's pool, rather than your service, is what actually needs resizing. ## Decide the posture, and say it out loud A cap enforced by refusing to accept does not tell anyone it has bound. Excess connections queue in the kernel and then fail there; clients see slow connects and transport errors, your service logs nothing, and the incident is diagnosed as "the network". Accepting connections and shedding over-capacity requests with an explicit response costs a little more memory but makes saturation visible to the balancer, to dashboards and to callers. The usual right answer for a fleet behind a balancer is: accept and shed, so the balancer can act; keep the hard accept cap as the backstop that prevents the OOM. But it is a real trade and the argument has to be made, not assumed. ## Decide where the cap is authoritative The same limit can be expressed in three places: your process's listener, the balancer's upstream connection pool, and the platform's memory limit on the instance. If all three exist and none is designated, you get the worst outcome — whichever is lowest silently governs, and nobody knows which one that is until an incident. Name one as authoritative, document the others as backstops with deliberately looser values, and make sure they are reviewed together when any of them moves. ## The compatibility lever nobody prices Lowering the per-request header ceiling from the stdlib's generous default shrinks the worst case an anonymous client can force you to buffer, so it directly improves the arithmetic behind the cap. But it is also a compatibility decision: clients with large cookies, long signed tokens, or requests that have accumulated forwarding headers through several proxies will start failing, and they will fail before any of your handler logging runs. Measure the real distribution of header sizes first, roll it out where it can be reverted, and treat it as a change to your public contract rather than a memory tweak. ## Governance: who can move it Write down three things: * **Owner** — whoever is accountable for instance memory sets the number and defends it, because raising it means either OOMs or a larger fleet, and both have a cost with a name. * **Override** — the service owner can overrule it when it demonstrably refuses legitimate load, and that override should be fast and logged rather than requiring a planning cycle. A cap that cannot be raised in an incident will be removed entirely in one. * **Alarm** — page on approaching the cap, not on hitting it. Hitting it is already an outage for somebody; approaching it is the conversation about buying instances or resizing the balancer's pool. ## What a strong answer sounds like "I derive it from measured bytes per connection against the memory budget, keep connection count and active request count as separate numbers, prefer explicit shedding over silent queueing so the balancer can see saturation, name one authoritative place the cap lives, and give it an owner, an alarm before it binds, and a fast override path for the service team."

  • Why keep concurrent connections and concurrent requests as separate capacity numbers?
    Because connection reuse decouples them. Idle pooled connections cost memory and almost no CPU; active requests cost CPU and allocate. Sizing on one number alone either over-provisions compute or gets ambushed by memory, and the ratio between them is the signal that tells you which layer to fix.
  • Why prefer explicit shedding over a hard accept cap on a fleet behind a load balancer?
    Because a blocking accept cap makes saturation invisible: connections queue in the kernel and fail there, your service logs nothing, and the incident is blamed on the network. Shedding explicitly lets the balancer route away and shows up on dashboards. Keep the accept cap as the backstop against the OOM.
  • What goes wrong when the cap exists in the process, the balancer and the platform at once?
    The lowest one silently governs and nobody knows which. Name one as authoritative, set the others deliberately looser as backstops, and review them together, otherwise a routine change to any of the three moves the effective limit without anyone deciding to.
  • What should the escalation path for raising the cap look like?
    Fast and logged. The capacity owner defends the number because raising it means OOMs or more instances; the service owner can override when it is demonstrably refusing legitimate load. A cap that takes a planning cycle to raise gets deleted during the first incident it causes.

saying these in an interview costs you the question

  • Picks a round number with no bytes-per-connection measurement
  • Sizes the fleet on request rate and ignores held connections
  • Leaves the cap enforced in three places with none authoritative
  • Accepts silent queueing without deciding it is the posture
  • Has no owner or escalation path for raising the cap