skip to content

A Go sidecar with GOMAXPROCS=4 runs 90 OS threads — what explains it, and what do you change?

level: seniorimportance: should knowfreq 40%

answer

  1. count contexts, then count threads
  2. they are not the same limit
  3. who is stuck in the kernel right now
  4. name lookups and file calls hold threads
  5. bound the blocking work, prefer the Go resolver

basics

~20 s

GOMAXPROCS bounds goroutines running Go code, not OS threads. Every goroutine parked in a blocking system call or a C call holds its own thread while the runtime spins up more to keep the scheduling contexts busy. Cap the blocking work, do not tune the scheduler.

solid answer

~50 s

That is the scheduler working as designed, not a leak. `GOMAXPROCS` limits scheduling contexts — how many goroutines execute Go code simultaneously — while every goroutine sitting in a blocking system call or a cgo call holds an OS thread of its own, and the runtime creates fresh threads so the contexts stay busy. In a sidecar doing name resolution and disk stats, the usual sources are the cgo-based system resolver and ordinary file operations, neither of which goes through the netpoller. Confirm it by reading the real thread count — `ps -o nlwp= -p <pid>` or the `Threads:` line in `/proc/<pid>/status` — and comparing it with `GOMAXPROCS` and with what the goroutines are doing. The count is a high-water mark, because a thread that finishes a call parks for reuse rather than exiting. Fixes are on the workload: bound the concurrent blocking calls, prefer the pure-Go resolver, and cache repeated stats.

code

text · 4 lines
text
$ ps -o nlwp= -p 4211
90
$ grep ^Threads: /proc/4211/status
Threads:	90

go deeper

for a junior

Take away the headline: threads and GOMAXPROCS are different limits, and a thread count above it is normal in a program that makes blocking calls.

for a middle

Be able to derive the number — contexts running Go code, plus one thread per blocking call in flight, plus parked threads kept for reuse — and name which calls block.

for a senior

Show the diagnosis: read the real thread count from the OS, correlate it with the concurrent blocking work and its latency, and fix the workload with a cap or a non-blocking path rather than a runtime knob.

for a principal

Own the standard: what thread and memory footprint a sidecar is allowed on a shared node, who enforces the concurrency caps, and whether the platform mandates the pure-Go resolver across services.

A thread count far above `GOMAXPROCS` is the single most common "is something wrong with the Go runtime?" question from engineers arriving from a one-thread-per-core mental model. Almost always nothing is wrong, but it is worth being able to explain precisely and to decide whether the number is acceptable. ## The rule that resolves it `GOMAXPROCS` bounds **scheduling contexts**, and therefore how many goroutines can be executing Go code at the same instant. It says nothing about OS threads. A thread only needs a context while it is running Go code; a thread sitting inside a system call or inside C does not hold one. So the number of threads a Go process owns is roughly: > threads ≈ (threads running Go code, capped by GOMAXPROCS) + (calls currently blocked in the kernel or in C) + (parked threads kept for reuse) With `GOMAXPROCS=4` and 80-odd blocking calls in flight, 90 threads is exactly what that formula predicts. ## Where the blocking calls come from in a sidecar Two sources dominate for an agent-style process doing name resolution and filesystem inspection: **Name resolution.** Go has two resolvers. The pure-Go one sends its queries over UDP and TCP sockets, so it rides the netpoller and costs no thread while waiting. The cgo resolver calls the system's `getaddrinfo`, which is a C call that pins its thread for the whole lookup — and lookups are slow, on the order of milliseconds to seconds against a struggling resolver. A burst of concurrent lookups through the cgo path converts directly into a burst of threads. Which resolver is used depends on the platform and configuration; `GODEBUG=netdns=go` forces the pure-Go one process-wide, and `net.Resolver{PreferGo: true}` does it for one resolver instance. **Filesystem calls.** Every `stat`, directory read, and file read on a regular file is a genuine blocking call, because readiness polling does not apply to regular files. On a network mount or a throttled volume each one can take tens of milliseconds, so a modest fan-out becomes a large thread count. A third source, on any process that links C libraries, is cgo generally: the runtime treats a call into C like a system call, and long C calls hold threads. ## Confirming it rather than guessing Read the thread count from the OS, not from the Go program's goroutine count — those are different numbers and confusing them is the root of most of these tickets. On Linux, `ps -o nlwp= -p <pid>` prints the number of threads, and `/proc/<pid>/status` has a `Threads:` line. Watch it over time: a count that rises to a plateau and stays there is the expected high-water-mark behaviour, because a thread that finishes its blocking call parks on the runtime's idle list and is reused rather than destroyed. A count that climbs without bound tracks blocking work that is itself unbounded — that is the real problem, and it is a workload problem. Then correlate. Ask what the process was doing at the peak: how many concurrent lookups, how many concurrent file operations, and how slow each was. Latency matters as much as rate, because threads are held for the duration of each call. ## What the threads actually cost Each OS thread carries kernel bookkeeping, a runtime stack for scheduler code, and a mapped stack whose size is set by the OS. The virtual footprint looks alarming; the resident cost is smaller but real, and there is a scheduling cost too — more runnable threads than cores means the kernel is doing more context switching to no benefit. For a sidecar sharing a node with the workload it monitors, that overhead is precisely what you are trying not to spend. ## The fixes, in the order to try them 1. **Bound the blocking concurrency.** A counting semaphore around lookups and file operations converts an unbounded thread count into a chosen one. This is the fix, and it is a workload change. 2. **Take the blocking path out.** Prefer the pure-Go resolver so lookups go through the poller. Cache resolution results and filesystem metadata that changes slowly. 3. **Reduce the work.** Poll less often, batch stats, or read one aggregate file instead of walking a tree. ## The fix that is not a fix Raising or lowering `GOMAXPROCS` does not change this, because the threads are not there to run Go code. Note also that since Go 1.25 the default `GOMAXPROCS` on Linux honours the container's CPU limit and is re-read periodically, so in a container the value may well be smaller than the node's core count — which makes the gap between contexts and threads look wider without being a new problem.

  • Does the thread count come back down after the burst of blocking calls finishes?
    Not usually, and that is expected. A thread that returns from its call parks on the runtime's idle list rather than exiting, so the count behaves as a high-water mark and the memory stays mapped. Judge health by whether it plateaus: a stable plateau is the design working, a count that keeps climbing means the blocking work itself is unbounded.
  • Can you read the thread count from inside the process instead of from ps?
    Yes. Go 1.26 added scheduler metrics to `runtime/metrics`, including `/sched/threads:threads`, so a service can export its own thread count alongside its other gauges. That is worth wiring up before you need it, since the interesting comparison — threads against GOMAXPROCS and against concurrent blocking work — is much easier when both numbers are already in your dashboards.
  • How do you decide whether 90 threads is actually a problem?
    Look at cost and trend, not the number. If it plateaus, the resident memory is acceptable for the instance size, and the process is not causing context-switch churn on a node it shares, it is fine. It becomes a problem when the count tracks an unbounded fan-out, when the memory competes with the workload you are monitoring, or when it is masking latency in the calls that are holding the threads.

Four tills and ninety staff is not a mistake if most of the staff are waiting on hold to a supplier. The tills bound how many customers get served at once; the phone calls bound how many people you employ.

saying these in an interview costs you the question

  • Says GOMAXPROCS should cap the OS thread count
  • Calls it a thread leak without checking whether it plateaus
  • Reads the goroutine count as if it were the thread count
  • Proposes tuning GOMAXPROCS to fix thread growth
  • Blames the garbage collector for the extra threads