skip to content

Why is calling a C function from Go through cgo far more expensive than a plain Go function call?

level: middleimportance: must knowfreq 42%

answer

  1. not a jump, a transition
  2. the goroutine changes stacks
  3. the runtime is told you left Go
  4. nothing optimises across the boundary
  5. fixed per crossing, so batch

basics

~20 s

A cgo call is a runtime transition, not a jump. The goroutine moves onto the thread's system stack, the runtime records that it has left Go code, and both are undone on return. That costs tens of nanoseconds; a Go call costs a couple.

solid answer

~50 s

A Go-to-Go call is a jump the compiler fully understands, so it can inline it, keep arguments in registers and prove they never escape. A cgo call cannot be any of those. The goroutine must leave its small growable Go stack and run C on the thread's system stack, and the runtime must mark the goroutine as no longer executing Go code so its execution slot can be handed to another thread if the call is slow; on return the reverse happens. On top of the transition, the compiler cannot inline across the boundary, cannot optimise through it, and escape analysis gives up, so anything you point C at is usually heap-allocated. The practical consequence is that the boundary cost is fixed per call: a C function that takes 5 ns is dominated by it. The fix is batching — pass the whole float64 feature vector in one call instead of one call per feature — and measuring the gap with a benchmark rather than guessing it.

code

go · 13 lines
go
var sink float64

func BenchmarkCgoCall(b *testing.B) {
	for i := 0; i < b.N; i++ {
		sink = float64(C.score_one(C.double(1.5)))
	}
}

func BenchmarkGoCall(b *testing.B) {
	for i := 0; i < b.N; i++ {
		sink = scoreOne(1.5)
	}
}

go deeper

for a junior

Remember the headline: calling C from Go is not a normal function call, it goes through the runtime, and it costs roughly a hundred times a Go call. Be able to say that fewer, bigger calls beat many small ones.

for a middle

Be ready to name the three costs — the switch onto the system stack, telling the runtime the goroutine has left Go code, and the optimisations lost at an opaque boundary — and to explain why arguments pointed at by C tend to be heap-allocated.

for a senior

Show that you measure rather than quote. Expect to describe benchmarking the boundary on your own hardware, reading a CPU profile to see whether crossings or C work dominate, and redesigning the C API so one call does the whole vector.

for a principal

Own the API shape. A fine-grained C interface imported into a hot Go path is a cost you inherit for years, so argue for a batched entry point at integration time and for a benchmark in CI that fails when someone reintroduces a per-element crossing.

## What a normal Go call costs When one Go function calls another, the compiler knows both sides. It can inline the callee, keep arguments in registers, and run escape analysis across the call to decide that a value can stay on the caller's stack. Nothing has to be told to the runtime: the goroutine keeps running on the same stack, remains preemptible, and the garbage collector can still find every pointer because the compiler emitted stack maps for both frames. The cost is a call instruction — order of a nanosecond, often zero because the call disappeared into an inline. ## What a cgo call has to do instead A call written as `C.score_one(x)` is not that. cgo generates a Go wrapper that routes through the runtime, and three separate things happen. **A stack switch.** A goroutine runs on a small, growable, movable Go stack. C code assumes a large, fixed, never-moving stack and knows nothing about Go's stack-growth check that begins every Go function. So the runtime switches the thread onto its system stack before entering C, and switches back on return. **A scheduler notification.** While the thread is executing C, the runtime cannot see what it is doing, cannot preempt it, and cannot scan its frames for pointers. It therefore records that the goroutine has stopped executing Go code, much as it does for a system call, so that if the call turns out to be slow the goroutine's execution slot can be given to another OS thread and the rest of the program keeps running. That bookkeeping is cheap but not free, and it is paid twice per call. **Lost compiler work.** The boundary is opaque. The C function cannot be inlined into Go, no optimisation crosses it, and escape analysis stops there: because the compiler cannot prove what C does with a pointer you hand it, the pointed-to value escapes and is heap-allocated. `go build -gcflags=-m` will show the escape, and `go test -bench . -benchmem` will show the resulting allocation. ## The number, and why the number is not the point Order of magnitude: a cgo call is tens of nanoseconds against roughly one for a Go call — call it one to two orders of magnitude. The exact figure moves with hardware and with the Go release (Go 1.26 cut cgo call overhead by roughly 30%), which is precisely why you should measure it on your own machine rather than quoting a blog post. Two benchmarks, one over the C function and one over an equivalent Go function, give you the gap directly. What matters more than the number is that the cost is **fixed per crossing and independent of how much work the C function does**. If the C side runs for a millisecond, the boundary is noise and you should stop thinking about it. If the C side runs for 20 ns and you call it once per element of a large vector, the boundary is most of your runtime. ## The engineering response The response is almost always **batching**: change the C API so one crossing does N units of work. An RPC handler that scores a request should hand the entire feature vector to the engine in a single call, not call a per-feature entry point in a Go loop. This is also why a naive port of a C library that exposes a fine-grained API often benchmarks worse in Go than in C — the API shape, not the language, is the problem. Secondary responses: hoist repeated conversions out of loops, avoid handing C pointers to values that could otherwise stay on the stack, and reuse a single long-lived buffer across calls rather than allocating per call. ## What this is not It is not the C compiler being slow — that cost was paid at build time. It is not garbage collection — the collector is not stopped for the duration of your call. And it is not something you can annotate away: there is no directive that inlines C into Go. The boundary is a real runtime transition and the only lever you have is how many times you cross it.

  • Why does passing a Go pointer to a C function usually cause a heap allocation?
    Escape analysis has to prove a value's lifetime ends with the frame. It cannot see what the C function does with a pointer, so the value escapes and moves to the heap. `go build -gcflags=-m` reports it, and `-benchmem` shows the allocation count. Batching helps here too: one pointer per vector instead of one per element.
  • Does the boundary cost change if the C function is trivial?
    No, and that is the trap. The transition is a fixed cost paid on every crossing, so a C function that runs in 5 ns is almost entirely boundary. Cheap, frequently called C functions are the worst possible cgo shape; expensive, rarely called ones make the boundary irrelevant.
  • How would you decide whether cgo overhead is actually your problem?
    Benchmark the C path against an equivalent Go path with `go test -bench . -benchmem`, and take a CPU profile of the real service. If the profile shows time concentrated in the cgo entry path rather than inside the C work, you are paying for crossings and should batch. If the C work dominates, the boundary is noise.

It is the difference between stepping into the next room and clearing airport security. The walk is trivial; the checkpoint costs the same whether you are flying an hour or a day.

saying these in an interview costs you the question

  • Says a cgo call compiles down to an ordinary call instruction
  • Blames the overhead on the C compiler being slow
  • Thinks a compiler directive can inline C into Go
  • Claims argument type conversion is the whole cost
  • Assumes cgo is free because C itself is fast