skip to content

When a Go service exhausts descriptors, how do you decide between raising RLIMIT_NOFILE and capping connections?

level: principalimportance: nice to knowfreq 25%

answer

  1. one question decides it: bounded or not
  2. who owns the number, and who owns the failure
  3. the runtime may already have raised it for you
  4. a budget you can write down, not a value you type
  5. alert below the ceiling, not at it

basics

~20 s

Ask whether the descriptor count is bounded by design. If the code has an explicit ceiling that is legitimately above the limit, raise the limit. If the count grows with load and nothing caps it, raising the limit only buys hours and moves the failure somewhere worse.

solid answer

~50 s

The argument is really about who owns the number. The platform team owns the container's `RLIMIT_NOFILE`; the service owner owns the connection posture — `MaxConnsPerHost`, per-upstream concurrency, and whether response bodies are drained. Decide with one question: is peak descriptor count bounded by something in the code? If yes, and the bound genuinely exceeds the limit, the limit is wrong — descriptors are cheap kernel objects and a proxy holding tens of thousands is legitimate. If no, raising the limit is buying time against an unbounded quantity, and the next failure will be memory, ephemeral ports, or an upstream you have now overwhelmed at higher concurrency. Note also that Go's runtime already raises the soft limit to the hard limit at startup, so "bump the ulimit" often changes nothing. The durable outcome is a written budget plus an alert below the ceiling, not a one-off number change.

go deeper

for a junior

Understand that raising a limit is not the same as fixing a leak, and that the first question is always whether the resource is being released at all.

for a middle

Be able to compute a rough descriptor budget from concurrency and per-host connection caps, and explain why an unset cap means the count follows load without any ceiling.

for a senior

Argue both sides concretely: when a higher limit is correct engineering for a bounded proxy, and when it merely relocates the failure to memory, ports or a saturated upstream.

for a principal

Own the decision and its cost. Set a written descriptor budget, make the limit a multiple of it, alert below the ceiling, and get explicit agreement that a connection cap will show up as queueing latency before it ships.

## The argument, stated fairly When a Go service dies with `too many open files`, two teams show up with two fixes. The platform team says the service is holding an absurd number of connections and should cap them. The service team says descriptors are a trivial kernel resource, the limit is set to some historical default, and raising it is a one-line change. Both are sometimes right, and the reason this is a principal-level call is that whoever wins owns the next outage. ## The question that decides it **Is peak descriptor count bounded by something in the code, or does it follow load?** That is the whole judgment, and everything else is detail. **Bounded.** A proxy with `MaxConnsPerHost` set per upstream, a listener with a known connection cap, a worker with a fixed file concurrency — these hold a number you can compute in advance: upstreams times per-host cap, plus inbound connections, plus a margin. If that computed number is above the configured limit, the *limit* is the defect. Descriptors are small kernel objects; a network proxy legitimately holding 30,000 of them is normal engineering, and refusing to configure for it is cargo-culting a default meant for a login shell. **Unbounded.** No per-host cap, connection count riding concurrency, or an outright leak. Here raising the limit is not a fix, it is a delay, and it is a delay that makes the next failure worse. At the higher limit the service will hold more simultaneous connections before it notices anything wrong, which means: more memory per connection, more ephemeral port pressure, and — the part people miss — far more concurrency aimed at the upstream that was already slow. You have turned your own outage into theirs. ## The Go-specific facts that change the argument Two details routinely embarrass one side or the other: 1. **The Go runtime already raises the process's soft `RLIMIT_NOFILE` to the hard limit at startup** on Unix, and `os/exec` restores the original soft limit for child processes. So a proposal to "raise the soft ulimit" frequently changes nothing at all — the service was already at the hard limit. The real knob is the hard limit, which in a container is set by the runtime and the orchestrator, not by the service. 2. **The zero value of `MaxConnsPerHost` means unlimited.** A team can honestly believe they configured their client carefully and still have no ceiling, because the field they did not set is the one that would have provided it. "We have a cap" should always be answered with "show me the field". ## What the durable decision looks like Not a number typed into a manifest during an incident. Three artefacts: **A stated budget.** "At p99 this instance holds N descriptors: I inbound connections, U upstreams times C per-host connections, F open files, plus overhead." That sentence is the thing that makes the limit reviewable, and it forces the code to have caps, because you cannot write the sentence otherwise. **A limit set as a multiple of the budget** — typically three to four times — so that ordinary growth does not require a platform change, and a runaway still hits a wall rather than eating the node's system-wide descriptor table. **Descriptor count as a monitored signal with an alert well below the ceiling.** The alert is the entire point: it converts "the service died at peak" into "capacity planning ticket". Pair it with connection-reuse rate, which moves earlier than the descriptor count and points at code changes rather than at traffic. ## Naming the loser The part that makes this a leadership question rather than a tuning question: say out loud what each choice costs and who is accountable. - If you raise the limit **without** a cap in the client, the service team owns the next incident, and it will be a worse one — a saturated upstream or an out-of-memory kill instead of a clean, obvious error. - If you impose a cap, requests will queue when an upstream is slow, and p99 latency will get worse in exactly the situation where it was already bad. Somebody must accept seeing that on a dashboard and not treat it as a regression. That acceptance has to be agreed before the cap ships, not argued during the next incident. - If a genuinely bounded service is denied a higher limit because of a platform-wide default, the platform team owns the availability cost of that default. "The default is 1024" is not a technical position. ## The one thing to do first, always Before either change, establish whether the count returns to baseline when traffic does. A leak is not an argument between teams; it is a bug, and neither raising the limit nor capping connections fixes it. Capping merely makes it fail earlier and more predictably, which is useful, and raising the limit makes it fail later and less predictably, which is not. Settle the leak question first, then have the ownership conversation about the bounded steady state that remains.

  • A team asks you to approve raising the container's descriptor limit from 1024 to 65536. What do you ask for?
    The budget sentence: how many descriptors does one instance hold at p99, decomposed into inbound connections, upstreams times per-host cap, and open files? If they can produce it and the number exceeds 1024, approve immediately — the default was wrong. If they cannot, the count is unbounded and the raise is a delay; the missing per-host connection cap is the actual change.
  • Why is 'descriptors are cheap, just raise the limit' both true and dangerous?
    True because a descriptor is a small kernel object and a network service legitimately holds tens of thousands. Dangerous because the descriptor limit is often the only bound the system has: remove it from an unbounded design and the next constraint reached is memory, ephemeral ports, or the upstream's own capacity — all of which fail more slowly and more confusingly than a clear open-file error.
  • How do you make the tradeoff visible rather than re-litigating it every incident?
    Write the budget into the service's operational documentation, set the limit as an explicit multiple of it, and alert on descriptor count well below the ceiling. Then make a per-upstream connection cap a review expectation for every new client, so 'we have no ceiling' shows up in a pull request rather than at peak traffic.

saying these in an interview costs you the question

  • Treats raising the limit as free with no budget behind it
  • Assumes changing the soft ulimit affects a running Go process
  • Caps connections without warning anyone that latency will rise
  • Argues the policy before establishing whether there is a leak
  • Defends a platform default with no availability reasoning