skip to content

Overhead and Thread Pinning

A C call costs far more than a Go call and takes the goroutine off the scheduler onto an OS thread it cannot preempt, so a chatty or blocking boundary shows up as thread growth and latency.

part ofGo (Golang)overview, primer and where to startread it →
on this pageshow

questions

5

Why is calling a C function from Go through cgo far more expensive than a plain Go function call?

level: middleimportance: must knowfreq 42%

answer

  1. not a jump, a transition
  2. the goroutine changes stacks
  3. the runtime is told you left Go
  4. nothing optimises across the boundary
  5. fixed per crossing, so batch

basics

~20 s

A cgo call is a runtime transition, not a jump. The goroutine moves onto the thread's system stack, the runtime records that it has left Go code, and both are undone on return. That costs tens of nanoseconds; a Go call costs a couple.

solid answer

~50 s

A Go-to-Go call is a jump the compiler fully understands, so it can inline it, keep arguments in registers and prove they never escape. A cgo call cannot be any of those. The goroutine must leave its small growable Go stack and run C on the thread's system stack, and the runtime must mark the goroutine as no longer executing Go code so its execution slot can be handed to another thread if the call is slow; on return the reverse happens. On top of the transition, the compiler cannot inline across the boundary, cannot optimise through it, and escape analysis gives up, so anything you point C at is usually heap-allocated. The practical consequence is that the boundary cost is fixed per call: a C function that takes 5 ns is dominated by it. The fix is batching — pass the whole float64 feature vector in one call instead of one call per feature — and measuring the gap with a benchmark rather than guessing it.

code

go · 13 lines
go
var sink float64

func BenchmarkCgoCall(b *testing.B) {
	for i := 0; i < b.N; i++ {
		sink = float64(C.score_one(C.double(1.5)))
	}
}

func BenchmarkGoCall(b *testing.B) {
	for i := 0; i < b.N; i++ {
		sink = scoreOne(1.5)
	}
}

go deeper

for a junior

Remember the headline: calling C from Go is not a normal function call, it goes through the runtime, and it costs roughly a hundred times a Go call. Be able to say that fewer, bigger calls beat many small ones.

for a middle

Be ready to name the three costs — the switch onto the system stack, telling the runtime the goroutine has left Go code, and the optimisations lost at an opaque boundary — and to explain why arguments pointed at by C tend to be heap-allocated.

for a senior

Show that you measure rather than quote. Expect to describe benchmarking the boundary on your own hardware, reading a CPU profile to see whether crossings or C work dominate, and redesigning the C API so one call does the whole vector.

for a principal

Own the API shape. A fine-grained C interface imported into a hot Go path is a cost you inherit for years, so argue for a batched entry point at integration time and for a benchmark in CI that fails when someone reintroduces a per-element crossing.

## What a normal Go call costs When one Go function calls another, the compiler knows both sides. It can inline the callee, keep arguments in registers, and run escape analysis across the call to decide that a value can stay on the caller's stack. Nothing has to be told to the runtime: the goroutine keeps running on the same stack, remains preemptible, and the garbage collector can still find every pointer because the compiler emitted stack maps for both frames. The cost is a call instruction — order of a nanosecond, often zero because the call disappeared into an inline. ## What a cgo call has to do instead A call written as `C.score_one(x)` is not that. cgo generates a Go wrapper that routes through the runtime, and three separate things happen. **A stack switch.** A goroutine runs on a small, growable, movable Go stack. C code assumes a large, fixed, never-moving stack and knows nothing about Go's stack-growth check that begins every Go function. So the runtime switches the thread onto its system stack before entering C, and switches back on return. **A scheduler notification.** While the thread is executing C, the runtime cannot see what it is doing, cannot preempt it, and cannot scan its frames for pointers. It therefore records that the goroutine has stopped executing Go code, much as it does for a system call, so that if the call turns out to be slow the goroutine's execution slot can be given to another OS thread and the rest of the program keeps running. That bookkeeping is cheap but not free, and it is paid twice per call. **Lost compiler work.** The boundary is opaque. The C function cannot be inlined into Go, no optimisation crosses it, and escape analysis stops there: because the compiler cannot prove what C does with a pointer you hand it, the pointed-to value escapes and is heap-allocated. `go build -gcflags=-m` will show the escape, and `go test -bench . -benchmem` will show the resulting allocation. ## The number, and why the number is not the point Order of magnitude: a cgo call is tens of nanoseconds against roughly one for a Go call — call it one to two orders of magnitude. The exact figure moves with hardware and with the Go release (Go 1.26 cut cgo call overhead by roughly 30%), which is precisely why you should measure it on your own machine rather than quoting a blog post. Two benchmarks, one over the C function and one over an equivalent Go function, give you the gap directly. What matters more than the number is that the cost is **fixed per crossing and independent of how much work the C function does**. If the C side runs for a millisecond, the boundary is noise and you should stop thinking about it. If the C side runs for 20 ns and you call it once per element of a large vector, the boundary is most of your runtime. ## The engineering response The response is almost always **batching**: change the C API so one crossing does N units of work. An RPC handler that scores a request should hand the entire feature vector to the engine in a single call, not call a per-feature entry point in a Go loop. This is also why a naive port of a C library that exposes a fine-grained API often benchmarks worse in Go than in C — the API shape, not the language, is the problem. Secondary responses: hoist repeated conversions out of loops, avoid handing C pointers to values that could otherwise stay on the stack, and reuse a single long-lived buffer across calls rather than allocating per call. ## What this is not It is not the C compiler being slow — that cost was paid at build time. It is not garbage collection — the collector is not stopped for the duration of your call. And it is not something you can annotate away: there is no directive that inlines C into Go. The boundary is a real runtime transition and the only lever you have is how many times you cross it.

  • Why does passing a Go pointer to a C function usually cause a heap allocation?
    Escape analysis has to prove a value's lifetime ends with the frame. It cannot see what the C function does with a pointer, so the value escapes and moves to the heap. `go build -gcflags=-m` reports it, and `-benchmem` shows the allocation count. Batching helps here too: one pointer per vector instead of one per element.
  • Does the boundary cost change if the C function is trivial?
    No, and that is the trap. The transition is a fixed cost paid on every crossing, so a C function that runs in 5 ns is almost entirely boundary. Cheap, frequently called C functions are the worst possible cgo shape; expensive, rarely called ones make the boundary irrelevant.
  • How would you decide whether cgo overhead is actually your problem?
    Benchmark the C path against an equivalent Go path with `go test -bench . -benchmem`, and take a CPU profile of the real service. If the profile shows time concentrated in the cgo entry path rather than inside the C work, you are paying for crossings and should batch. If the C work dominates, the boundary is noise.

It is the difference between stepping into the next room and clearing airport security. The walk is trivial; the checkpoint costs the same whether you are flying an hour or a day.

saying these in an interview costs you the question

  • Says a cgo call compiles down to an ordinary call instruction
  • Blames the overhead on the C compiler being slow
  • Thinks a compiler directive can inline C into Go
  • Claims argument type conversion is the whole cost
  • Assumes cgo is free because C itself is fast
open as a page

What does runtime.LockOSThread do, and when does a C library force you to use it?

level: middleimportance: should knowfreq 34%

basics

~20 s

runtime.LockOSThread wires the calling goroutine to the OS thread it is running on: that goroutine runs nowhere else and no other goroutine runs there until UnlockOSThread. You need it when a C library keeps per-thread state and must be driven from one fixed thread.

open as a page

Under load a handler's cgo call blocks in C for 300 ms and OS-thread count climbs past 200. Why, and how do you bound it?

level: seniorimportance: should knowfreq 38%

basics

~20 s

A goroutine executing C holds a real OS thread for the whole call and cannot be preempted. The runtime detaches its execution slot and starts another thread so Go work continues, so peak concurrent C calls sets thread count. Bound it with a semaphore at the boundary.

open as a page

Why can't a Go panic unwind out of a C frame, and what must an exported Go callback do about it?

level: seniorimportance: should knowfreq 27%

basics

~20 s

Panicking walks Go frames and runs their deferred calls; C frames carry none of that information, so the runtime cannot travel through them and aborts the whole process instead. Every Go function exported to C must recover in its own deferred call and return a status code.

open as a page

A cgo call into a C scoring engine takes 50-400 ms; how do you decide whether it stays on the RPC request path?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Turn it into a capacity and blast-radius decision. Concurrency times call duration is the OS-thread budget, so publish a cap and enforce it on entry; then judge whether a library that can kill the whole process, and that ignores your deadlines, belongs in the request path at all.

open as a page