Under load a handler's cgo call blocks in C for 300 ms and OS-thread count climbs past 200. Why, and how do you bound it?
answer
- C code cannot be preempted
- one blocked call, one real thread
- GOMAXPROCS does not bound threads
- cancellation is cooperative, C ignores it
- cap admission with a buffered channel
basics
~20 sA goroutine executing C holds a real OS thread for the whole call and cannot be preempted. The runtime detaches its execution slot and starts another thread so Go work continues, so peak concurrent C calls sets thread count. Bound it with a semaphore at the boundary.
solid answer
~50 sWhile a goroutine is inside C, the runtime cannot preempt it or scan its frames, so it marks the goroutine as having left Go code; shortly afterwards its execution slot is handed to another OS thread so the rest of the program keeps running. The blocked thread does not count against `GOMAXPROCS`, so the number of threads tracks the number of concurrent C calls, not your core count — 200 in-flight 300 ms calls means about 200 threads, each with its own stack, and the runtime does not shrink that count back afterwards. Cancelling the request's `context.Context` does not help: nothing in Go can interrupt a running C function. The fix is admission control — a buffered channel used as a semaphore in front of the call, sized from cores or an engine licence — so excess load queues in Go, where a deadline can shed it. Confirm the diagnosis with the threadcreate profile and the process's thread count read next to request latency.
code
go · 12 lines// At most 8 goroutines may be inside the C engine at once.
var slots = make(chan struct{}, 8)
func Score(ctx context.Context, v []float64) (float64, error) {
select {
case slots <- struct{}{}:
case <-ctx.Done():
return 0, ctx.Err() // shed at the gate, not inside C
}
defer func() { <-slots }()
return callEngine(v) // blocks in C; ctx cannot interrupt this
}go deeper
Know that a goroutine inside C occupies a real operating-system thread and cannot be paused, so a slow C call is nothing like a slow Go call. That single fact explains most of the surprise here.
Be ready to explain that the runtime hands the goroutine's execution slot to another thread so Go work continues, that GOMAXPROCS bounds Go execution rather than threads, and that concurrent C calls therefore set the thread count.
Demonstrate the full loop: read the threadcreate profile and thread count against latency, prove the growth tracks in-flight C calls, then put a sized semaphore in front of the boundary and shed load at the gate with the request deadline.
Decide and publish the ceiling. Frame it as capacity: threads are memory and scheduling cost, the cap converts overload into visible queueing, and the number should come from hardware or licence limits with an alert when gate wait time grows.
## Why a C call costs a whole OS thread Goroutines are cheap because the runtime can move them: at a safe point it can stop one, park it, and run another on the same OS thread. A goroutine inside a C function is not movable. The runtime does not know where the C code is, has no stack maps for its frames, and will not interrupt foreign code with a preemption signal. So for as long as the C function runs, one operating-system thread is committed to it and does nothing else. That is fine for a 200 ns call. It is a capacity question for a 300 ms call. ## What the runtime does about it Entering C, the runtime records that the goroutine has stopped executing Go code — the same treatment a blocking system call gets. Shortly after, a background monitor notices the goroutine has been away for a while and hands its execution slot to another OS thread, which is created or taken from the idle list. That is the behaviour you want: it means one slow C call does not stop the other requests, and Go code keeps running at full width. The price is arithmetic. The number of threads the process needs is roughly the peak number of concurrent blocking C calls, plus the threads doing Go work. `GOMAXPROCS` bounds how much Go code runs simultaneously; it does not bound threads, and a thread parked in C is not counted against it. With 300 ms calls and 700 requests per second arriving, Little's law says about 200 calls are in flight at any moment, so about 200 threads. Each carries an OS stack and runtime bookkeeping, and Go does not destroy threads once created — the count is a high-water mark for the life of the process. The symptom set is characteristic: thread count climbing and never falling, memory that grows in steps unrelated to the Go heap, rising scheduling latency and context switching, and p99 latency degrading across *every* endpoint, not just the one calling the engine. ## Why the request's context does not save you Cancellation in Go is cooperative: a `context.Context` deadline is a signal that Go code chooses to observe. A C function observes nothing. When the deadline fires you can stop waiting and return an error to the caller, but the C call runs to completion and its thread stays committed. Worse, if the caller retries on that timeout you now have two calls in flight for one request, and the thread count climbs faster. A timeout at the C boundary is honest only if the C API itself accepts a deadline or a cancel handle. ## Bounding it The lever is admission control at the boundary. A buffered channel used as a semaphore caps the number of goroutines allowed inside C at once; that cap is the maximum number of threads the engine can cost you. Size it from something real — the machine's cores, the engine's own internal parallelism, a licence limit — not from request concurrency. Once the gate exists, overload turns into queueing in Go rather than thread growth, and queueing is something you can measure and shed: select on the semaphore and on `ctx.Done()`, so a request that would wait past its deadline gives up before it ever enters C. Time spent waiting at the gate becomes a metric that tells you the engine is undersized, which is exactly the signal a bare thread count fails to give you. Secondary moves, in rough order of preference: make the C call shorter or batch several requests into one crossing; move the work off the request path entirely to a bounded set of long-lived engine goroutines fed over a channel; or run the engine out of process and talk to it over a socket, trading serialisation cost for isolation. ## Confirming it rather than assuming it Three readings settle it. The **threadcreate profile** (`/debug/pprof/threadcreate`, or `go tool pprof` against it) attributes thread creation to the stacks that caused it. The process's own thread count — read from the operating system, or from the `/sched/threads:threads` metric in `runtime/metrics` on recent Go — plotted next to request latency shows whether the climb tracks load. And a goroutine dump shows the offending goroutines sitting in the runtime's cgo entry path, marked as being in a system call, which distinguishes engine calls from ordinary Go blocking. If thread count tracks in-flight C calls, the diagnosis is closed.
- Does cancelling the request's context stop the running C call?No. Cancellation is cooperative and only Go code observes it. The handler can stop waiting and return an error, but the C function runs to completion and keeps its OS thread the whole time — and if the caller retries, you now have two calls in flight for one request. Only a C API that takes its own timeout or cancel handle gives you a real deadline.
- Why does the runtime not simply preempt a goroutine sitting in C?Preemption works by interrupting Go code at points where the runtime has stack maps and can safely park the goroutine. Inside a C function there are no Go frames to inspect, no safe points, and no way to resume elsewhere, so the runtime leaves it alone and instead detaches its execution slot so other threads can carry on.
- How do you tell engine-blocked threads apart from ordinary Go blocking in a dump?A goroutine blocked on a channel or mutex shows that in its stack and is not holding a thread. A goroutine inside C shows the runtime's cgo entry path and is marked as in a system call. Cross-check with the threadcreate profile: if the stacks creating threads are the engine's callers, the growth is the C boundary.
- Does thread count come back down once the burst is over?No. The runtime parks idle threads for reuse but does not destroy them, so the count is a high-water mark that persists for the life of the process. That is why a single bad burst leaves a permanently fatter process, and why the ceiling has to be enforced on entry rather than cleaned up afterwards.
saying these in an interview costs you the question
- Thinks GOMAXPROCS caps the number of OS threads
- Expects the scheduler to preempt a running C call
- Believes a context deadline aborts the C function
- Assumes thread count shrinks after the burst passes
- Adds goroutines to make a serial C engine faster