skip to content

A Go transcoder's OS thread count climbs past a thousand while its goroutine count stays flat — how do you diagnose that?

level: seniorimportance: should knowfreq 45%

answer

  1. the flat line is the clue
  2. workload is not growing, yet threads are
  3. something holds a thread it cannot give back
  4. arrival rate times blocking duration
  5. cap the calls in flight, not the threads

basics

~20 s

A flat goroutine count with rising OS threads means goroutines are stuck in long blocking calls into C that hold their thread. Confirm with the threadcreate and goroutine profiles, then bound how many such calls run at once.

solid answer

~50 s

The shape of the graph is the diagnosis. Thread creation tracks *stuck* threads, not workload: when a goroutine sits inside a long C call or a long blocking system call, the runtime cannot reuse that thread, so it makes another one to keep the remaining goroutines running. A flat goroutine count rules out a goroutine leak and points straight at blocking depth. Confirm it in three readings: the `/debug/pprof/threadcreate` stacks to see which call path is producing threads, the goroutine profile to see the same goroutines parked in that call, and `GODEBUG=schedtrace=1000` or `/sched/threads:threads` for the live trend. Then look at memory — threads created for C calls get the platform's default stack reservation, so a thousand of them is gigabytes of address space. The fix is a bound: put a counting semaphore in front of the blocking call so concurrency, and therefore thread demand, is capped by design.

code

go · 12 lines
go
var codecSlots = make(chan struct{}, 16)

func transcode(ctx context.Context, blob []byte) ([]byte, error) {
	select {
	case codecSlots <- struct{}{}:
	case <-ctx.Done():
		return nil, ctx.Err()
	}
	defer func() { <-codecSlots }()

	return encodeWithCCodec(blob) // long blocking call; holds one OS thread
}

go deeper

for a junior

Know that goroutines and OS threads move independently, and that a stuck blocking call ties up a real thread while a parked goroutine does not. Being able to say which graph you would look at first is enough.

for a middle

Explain the mechanism in your own words: work is not increasing, so the extra threads come from existing ones being unavailable. Name the profiles you would pull and what each one would show.

for a senior

Do the capacity arithmetic — arrival rate times blocking duration — tie the thread count to resident memory, and land on bounding concurrency at the blocking call rather than tuning the runtime.

for a principal

Frame it as a design constraint rather than an incident: any unbounded blocking call is a latent capacity bug, and the standard is that such calls sit behind an owned, measured bound with an explicit thread ceiling behind that.

## Reading the two lines A media transcoder hands each job to a C codec and gets a large encoded blob back. Over an hour, the operating system's thread count for the process goes from 12 to 1,400. `runtime.NumGoroutine()` sits at about 60 the whole time. Resident memory climbs steadily. That combination is diagnostic on its own, so start by naming what it rules out. - **Not a goroutine leak.** A leak shows the opposite signature: goroutines climbing, threads roughly flat, because parked goroutines cost no thread. - **Not increased load.** Load would move the goroutine count too. - **Not the garbage collector.** Its worker goroutines are bounded and show in the goroutine count. What is left is **blocking depth**: a small, stable number of goroutines, each of which spends a long time somewhere the runtime cannot reclaim its thread. The runtime's response to "my threads are all stuck and I have runnable goroutines" is to make another thread. Do that a few thousand times and you have this graph. The arithmetic is worth doing out loud, because it turns the mystery into a capacity calculation: threads needed ≈ job arrival rate × average time spent inside the blocking call. Forty jobs a second at 30 seconds inside the codec is 1,200 threads, and nothing in the program says "1,200" anywhere. ## Confirming it **1. Pin down the real count and the trend.** Read `/sched/threads:threads` from `runtime/metrics`, or run with `GODEBUG=schedtrace=1000` and watch the `threads=` and `idlethreads=` fields once a second, or take the operating system's own count. You want to see the count climbing and, importantly, *not falling back* when load drops. **2. Attribute the creations.** Pull `/debug/pprof/threadcreate` and read the top stacks, not the total. A dominant stack running through the runtime's cgo entry path is a direct answer: calls into C are producing threads. Remember the profile under-reports, so use the ranking, not the number. **3. Corroborate from the goroutine side.** `/debug/pprof/goroutine?debug=2` shows every goroutine's stack. You expect roughly as many goroutines parked in the codec call as you have excess threads. If the two views disagree, one of your assumptions about where time is spent is wrong. **4. Connect it to memory.** This is the part that turns "an odd metric" into "the reason we are being killed". In a pure-Go program the runtime allocates a small system stack for each thread. Threads created for C calls are created through the platform's thread API and get the platform's *default* stack reservation — commonly 8 MB of virtual address space per thread on Linux, committed as it is touched. A thousand of those is gigabytes of address space and a resident cost that grows quietly. Overlay the thread count on resident memory and the two lines will match. ## Why it does not recover on its own The Go runtime does not aggressively destroy idle threads. Once created, an M generally sticks around, so the thread count behaves like a **high-water mark**: a five-minute burst of slow codec calls leaves a permanently fatter process, and the next burst starts from there. Any mental model that says "it will settle down when traffic drops" is wrong, and it is why the peak, not the mean, is the number your memory budget must cover. ## The fix The cause is unbounded concurrency at a blocking call, so the fix is a bound at that exact point. A counting semaphore built from a buffered channel is the idiomatic Go form: ```go var codecSlots = make(chan struct{}, 16) ``` Acquire before the call, release with `defer` after. Now at most sixteen goroutines can be inside the codec at once, thread demand is bounded by construction, and excess work waits in a queue you can see and measure instead of turning into kernel threads you cannot. Size the bound from measurement — how many concurrent codec calls the machine can actually sustain — not from the request rate. Two supporting moves, neither of which is the fix: - **An explicit thread ceiling** via `debug.SetMaxThreads`, set well above the bound and well below the 10,000 default, so a regression fails loudly and early with stacks rather than as an out-of-memory kill. - **Alerting on the shape**, not the value: thread count rising while goroutine count is flat, and thread count as a fraction of its historical peak. ## What a weak answer looks like Reaching for `GOMAXPROCS` — lowering it does not bound threads, because it limits how many goroutines execute Go code simultaneously, not how many threads exist. Adding a worker-goroutine pool without a bound on the blocking call is equally cosmetic: the goroutine count becomes tidy while thread creation continues unchanged. And raising the thread ceiling to "give it room" simply buys a larger crash.

  • Why does a flat goroutine count rule out a goroutine leak here?
    Because a leak is defined by goroutines accumulating: blocked goroutines pile up and the count only rises. Parked goroutines are cheap and hold no thread, so a leak shows as rising goroutines with roughly flat threads. This graph is the mirror image, which means the population is stable and each member is holding a thread for a long time.
  • How do you size the semaphore that fronts the blocking call?
    From measurement, not from the request rate. Benchmark how many concurrent calls the machine sustains before latency degrades — usually a small multiple of the CPU count for a compute-heavy codec — and set the bound there. Then confirm the resulting thread count and memory are within budget, and expose queue depth and wait time so saturation is visible rather than silent.
  • Would lowering GOMAXPROCS reduce the thread count?
    No. GOMAXPROCS bounds how many goroutines execute Go code simultaneously; it does not cap how many operating-system threads exist. Threads stuck in blocking calls are precisely the ones that are not executing Go code, so the runtime keeps creating more regardless of that setting. Only bounding the blocking calls bounds the threads.
  • After load drops back to normal, what do you expect the thread count to do?
    Stay near its peak. The runtime does not aggressively reap idle threads, so the count behaves like a high-water mark and the process keeps the memory those thread stacks reserved. That is why a short incident leaves a permanently fatter process, and why restarting is often the only way back to the original footprint.

Every job that goes into the codec takes a checkout lane with it and does not give it back until it is done. The shop does not have more customers than yesterday; it just keeps opening lanes because none of the open ones ever frees up.

saying these in an interview costs you the question

  • Calls rising threads with flat goroutines a goroutine leak
  • Expects the thread count to fall back after the burst
  • Suggests lowering GOMAXPROCS to bound threads
  • Adds a worker pool without bounding the blocking call
  • Raises the thread ceiling and calls it fixed
  • Ignores the memory cost of a thousand thread stacks