skip to content

A Go CLI wrapping a C library grows in RSS for hours while Go's heap profile stays flat. What do you suspect?

level: seniorimportance: should knowfreq 18%

answer

  1. two allocators, one process
  2. the profiler is telling the truth
  3. RSS counts what the runtime never allocated
  4. correlate growth with calls through the wrapper
  5. the evidence has to come from the C side

basics

~20 s

Memory allocated on the C side. Go's heap profile and MemStats only account for the Go heap, so malloc'd buffers are invisible to them. The usual cause is a C.CString or a library-returned pointer that is never passed to C.free.

solid answer

~50 s

The two numbers disagree because they measure different heaps. Go's pprof heap profile and `runtime.MemStats` describe only memory the Go runtime allocated; anything obtained through `malloc` on the C side — every `C.CString`, plus any buffer the C library returns — is outside their accounting but firmly inside the process's RSS. So a flat Go profile alongside climbing RSS is the signature of a leak on the **C** side, not a Go one. I would confirm the shape first: does growth track call volume through the wrapper, and does RSS minus `MemStats.HeapSys` account for the gap? Then I would look at the wrapper's allocation sites, checking that every `C.CString` has a matching `C.free(unsafe.Pointer(...))` on a path that always runs, and that each pointer the library returns is freed or copied according to its documented ownership. Confirming it from the C side — malloc/free counters, or running the binary under a leak checker — turns the hypothesis into evidence.

go deeper

for a junior

Take away one fact: Go's heap profile only reports memory the Go runtime allocated, so memory obtained by C code does not appear in it even though it counts toward the process's RSS.

for a middle

Be able to explain the mechanism — two allocators in one process — and name the concrete causes, chiefly a C.CString on a path that skips its free and a deferred free stuck inside a long-running loop.

for a senior

Show a method rather than a guess: rule out the Go side first, correlate growth with calls through the wrapper, compare RSS against MemStats, then produce C-side evidence with allocation counters or a leak checker before changing code.

for a principal

Argue for the structural fix. Confine all C. traffic to one wrapper package with paired allocate-and-release, give long-lived C handles an explicit Close, and decide whether keeping allocation counters behind a debug flag is worth the permanent diagnosability.

## Why the two numbers disagree This is one of the few production symptoms in Go where the standard tooling tells you almost nothing, and understanding *why* is the whole answer. Go's heap profile is produced by the runtime's own allocator: it samples allocations that the Go runtime performed, and reports them by call site. `runtime.ReadMemStats` fills a `runtime.MemStats` with the same universe of information — `HeapAlloc`, `HeapSys`, `HeapInuse` all describe the Go heap. Neither mechanism knows anything about `malloc`. When a C library allocates a buffer, the operating system maps pages into the process and RSS rises, but from the Go runtime's point of view nothing was allocated at all. So the disagreement is not a bug in the profiler. It is the profiler correctly reporting that the growth is not Go's. ## Narrowing it down Before reaching for heavy tooling, three cheap observations discriminate between candidate causes: 1. **Does the growth track call volume?** Instrument the wrapper with a counter of calls made and sample RSS beside it. A leak proportional to calls points straight at the per-call allocation path. Growth that continues while the tool is idle points elsewhere. 2. **How big is the gap?** Read `runtime.MemStats` periodically and compare `HeapSys` (plus the other `Sys` fields) against RSS. Some gap is always expected — the runtime's own bookkeeping, stacks, the binary itself — but a gap that widens monotonically is the leak, and its slope tells you how much per call. 3. **Is anything else growing?** Check the goroutine count too, from the goroutine profile. If goroutines are flat and only RSS climbs, that rules out the far more common Go-side explanation and keeps you honest. Ruling out the Go side first matters, because "it must be cgo" is a comfortable story that will waste a day if the actual cause was an unbounded Go-side cache. ## The usual causes Once the C heap is implicated, there is a short list, in rough order of frequency: - **A `C.CString` that is never freed.** The wrapper allocates a C string for every call, and one code path — an early `return` on a validation error, a path added months after the happy path was written — skips the release. This is why the allocate-then-`defer`-free-on-the-very-next-line habit exists: it makes the leak impossible to introduce by adding a return statement later. - **A `defer C.free(...)` inside a loop.** Deferred calls run at function return, so a long-running loop accumulates every buffer until the function finally exits. In a batch CLI processing a large file tree, that function may not return for hours. - **A returned pointer whose ownership was misread.** The C function `malloc`'d its result and documented that the caller frees it; the wrapper calls `C.GoString` to copy the text out and then drops the pointer. The Go string is correct, the memory is gone. The mirror mistake — freeing a pointer into the library's own static buffer — does not leak; it corrupts the C heap and crashes, usually somewhere unrelated. - **A C-side handle that has an explicit destructor.** Many C libraries pair a `*_new`/`*_open` with a `*_free`/`*_close`. Every one of those constructors is a resource the Go wrapper must release, and the Go type wrapping it needs a `Close` method that callers are expected to call. ## Confirming it on the C side Hypothesis becomes evidence only with C-side accounting: - **Malloc counters.** The cheapest instrumentation is often a pair of tiny `static` wrapper functions in the preamble that increment counters around allocation and release, exposed to Go so the CLI can print them. Allocations minus frees, per call, is a number that ends the argument. - **A leak checker.** Running the binary under a memory-error detector such as valgrind reports unreachable `malloc`'d blocks at exit with the C stack that allocated them, which names the leaking call site directly. It is slow, and a cgo binary produces some baseline noise from the runtime, so run it over a short deterministic workload rather than the production one. - **Allocator statistics.** On glibc, `malloc_stats` or `mallinfo` can be called from a preamble helper to show the C allocator's arena size growing. ## The fix, and the shape that prevents recurrence The individual fix is small: pair each allocation with its release on the adjacent line, move loop bodies into their own functions so deferred frees fire per iteration, and read the library's documentation for every pointer it hands back. The durable fix is structural. Keep all `C.` traffic inside one wrapper package, allocate and free within the same function wherever possible, and where a C handle must outlive a call, wrap it in a Go struct with an explicit `Close` method so the lifetime is visible to callers as ordinary Go. Where the leak counters proved useful during the incident, keeping them behind a debug flag costs nothing and makes the next occurrence a five-minute diagnosis. ## What an interviewer is listening for The key sentence is that Go's heap profile only covers the Go heap. After that they want a method: rule out the Go side, correlate growth with call volume, then get C-side evidence rather than guessing. A candidate who proposes tuning `GOGC` or calling `runtime.GC()` to fix it has missed that the collector cannot see this memory at all.

  • Would the race detector or go vet have caught this?
    No. The race detector instruments Go memory accesses to find concurrent unsynchronised access, and it finds only races on paths actually executed; a leak is not a race. `go vet` inspects Go source for suspicious constructs and has no model of C ownership. Nothing in the Go toolchain accounts for malloc'd memory, which is why C-side instrumentation is required.
  • Would tuning GOGC or setting GOMEMLIMIT help here?
    No. Both govern the Go heap: GOGC sets the growth target that triggers a collection, and GOMEMLIMIT is a soft limit on memory the Go runtime manages. C memory obtained through malloc is invisible to the collector, so no pacing knob can reclaim it. Worse, tightening them adds CPU cost while the RSS curve keeps climbing.
  • How would you instrument the C side cheaply, without a full leak checker?
    Add two small static helpers to the preamble that wrap allocation and release and bump counters, then expose the counts to Go so the tool can report allocations minus frees per run. It costs a few lines, runs at production speed, and turns "probably a leak" into a number tied to call volume — leaving the slow leak checker for when you need the allocating stack.
  • The C function returns a char*. How do you decide whether your wrapper should free it?
    The library's documented ownership contract decides. If it returns malloc'd memory, copy it out with `C.GoString` and then free the pointer; if it returns a pointer into static or internal storage, freeing corrupts the heap and the contents may be overwritten by the next call, so copy immediately and free nothing. When the documentation is silent, treat it as borrowed.

It is like auditing one company's books to explain a shared building's rising rent. The ledger is accurate and complete, and the spending simply happened in the other tenant's name.

saying these in an interview costs you the question

  • Blames the Go garbage collector for not reclaiming C memory
  • Proposes tuning GOGC or GOMEMLIMIT to fix it
  • Insists the heap profile would show any real leak
  • Never checks whether growth tracks call volume
  • Frees a library-owned pointer to stop the growth
  • Concludes it is cgo without ruling out a Go-side cache