skip to content

Profiles and pprof

Which profile answers which symptom - CPU, heap and allocs, block and mutex - and the workflow around one: exposing it from a live process, labelling it by request, and reading it.

part ofGo (Golang)overview, primer and where to startread it →
on this pageshow

explore

questions

25

How do you capture a CPU profile from a running Go program with runtime/pprof?

level: juniorimportance: must knowfreq 60%

answer

  1. two calls bracket the work
  2. the first one takes a writer
  3. deferred calls run last-in-first-out
  4. the second call is what flushes
  5. os.Exit skips defers, so the file is empty

basics

~10 s

Open a file, call runtime/pprof.StartCPUProfile(f), run the workload, then call StopCPUProfile before exiting. Stop flushes the buffered samples; without it the file is unusable. Inside tests, go test -cpuprofile does the same thing.

solid answer

~50 s

You bracket the work with two calls. `pprof.StartCPUProfile(w io.Writer)` starts sampling and returns an error if a CPU profile is already running — only one can be active per process. `pprof.StopCPUProfile()` takes no arguments and returns only after every sample has been written, so it is the call that actually makes the file readable. The usual shape is `f, _ := os.Create("cpu.out")`, then `pprof.StartCPUProfile(f)`, then `defer pprof.StopCPUProfile()`. Two things bite people: exiting through `os.Exit` or `log.Fatal` skips deferred calls, so you get a truncated or empty profile; and the profile covers wall-clock time between the two calls, so if you start it at process start you are profiling initialisation too. For a hot path like a search-ranking scorer, start the profile immediately before the batch of queries you care about and stop it right after.

code

go · 12 lines
go
f, err := os.Create("cpu.out")
if err != nil {
	log.Fatal(err)
}
defer f.Close()

if err := pprof.StartCPUProfile(f); err != nil {
	log.Fatal(err)
}
defer pprof.StopCPUProfile()

scoreQueries(queries) // only this is profiled

go deeper

for a junior

Be ready to write the four lines from memory: create a file, StartCPUProfile with it, defer StopCPUProfile, run the work. Say out loud that Stop is what makes the file readable.

for a middle

Explain why the pair is required, why defer order matters when the file is also deferred closed, and why only one CPU profile can be active in a process at a time.

for a senior

Show judgment about the window: profile the scoring loop, not process start-up, and choose a duration that yields enough samples for the traffic you actually have.

for a principal

Be able to say why this instrument is cheap enough to leave available in production — fixed cost per second rather than cost per call — and what a team should standardise on for capturing profiles.

## What a CPU profile is A Go CPU profile is a statistical record of where the program's threads were executing while the profiler was on. Roughly 100 times a second, per thread that is burning CPU, the runtime interrupts execution, walks the stack of the goroutine running on that thread, and records it. At the end you have a set of stacks with counts; multiplied by the sample period each count becomes an amount of CPU time. It is a sample, not a log: nothing counts calls, and nothing is instrumented into your functions. ## The two calls The whole API for an in-process capture is two functions in `runtime/pprof`: - `func StartCPUProfile(w io.Writer) error` — turns sampling on and streams the encoded profile to `w`. It returns a non-nil error if a CPU profile is already in progress; a process can only have one at a time. - `func StopCPUProfile()` — turns sampling off. It takes nothing and returns nothing, but it does not return until all pending writes to `w` have completed. That second point is the one juniors trip over. Samples are buffered and drained by a background goroutine; the file on disk is not a valid profile until `StopCPUProfile` has run. If the program exits without it — `os.Exit`, `log.Fatal`, a panic that is not recovered, a `kill -9` — you are left with a truncated file that `go tool pprof` refuses or reads as almost empty. ## The idiomatic shape ```go f, err := os.Create("cpu.out") if err != nil { log.Fatal(err) } defer f.Close() if err := pprof.StartCPUProfile(f); err != nil { log.Fatal(err) } defer pprof.StopCPUProfile() ``` Note the defer order. Deferred calls run last-in-first-out, so `StopCPUProfile` runs before `f.Close()` — which is what you want, because Stop still needs to write into the file. Reversing the two lines gives you writes to a closed file. ## Scoping the window The profile covers exactly the interval between the two calls, and it covers the whole process — every goroutine on every thread, not just the one that called Start. If your goal is to cut the cost of a search-ranking scorer that runs once per query inside a larger service, wrapping `main` gives you a profile dominated by configuration loading, index warm-up and connection setup. Start the profile immediately before the scoring loop instead, or drive the scorer from a benchmark and use `go test -cpuprofile`, which brackets the run for you. A second consequence of the interval being wall-clock: you need enough of it. At roughly 100 samples per second per busy core, a two-second profile of a mostly idle process yields a handful of samples and tells you nothing. Ten to thirty seconds of genuinely busy work — or a benchmark run with a longer `-benchtime` — is the normal target. ## Cost and safety Sampling is cheap and bounded: it is a fixed number of stack walks per second, independent of how many function calls the program makes. That is why the overhead stays in the low single-digit percent and does not scale with call volume, and why taking a CPU profile of a live service is a routine operation rather than an emergency measure. Long-running services usually expose the capture through an HTTP endpoint rather than calling Start and Stop by hand, but the underlying mechanism is exactly these two functions. ## Common mistakes - Calling Start and never calling Stop, then wondering why the file is empty. - Calling `os.Exit` inside the profiled region, which skips every deferred call. - Calling Start twice — the second call errors and, if you ignore the error, you believe you are profiling when you are not. - Profiling for two seconds and treating a function with three samples as a hotspot. - Expecting the profile to explain latency. It records CPU time. A goroutine parked on a channel, a lock or a network read burns no CPU and therefore leaves no samples. ## Where the data goes The output is a gzipped protocol-buffer profile. You do not read it by hand; you feed the file to `go tool pprof` together with the binary that produced it so symbols resolve. Keeping the binary that produced the profile matters — a profile from one build inspected against a different build produces nonsense addresses.

  • What happens if you call runtime/pprof.StartCPUProfile while a CPU profile is already running?
    It returns a non-nil error and does nothing else — the first profile keeps running and the new writer receives nothing. A process supports one CPU profile at a time. Code that ignores the error believes it is profiling when it is not, which is why the error return should always be checked.
  • Why does a profile come out empty when the program ends with log.Fatal?
    log.Fatal calls os.Exit, and os.Exit does not run deferred functions. StopCPUProfile never executes, so the buffered samples are never flushed and the file on disk is truncated. If a code path can exit that way, call StopCPUProfile explicitly before exiting rather than relying on defer.
  • How long should you leave a CPU profile running?
    Long enough to accumulate thousands of samples. At roughly 100 samples per second per busy core, ten seconds of one saturated core is about a thousand samples — enough for a function with a few percent share to be visible. For a low-traffic service, profile longer or drive the code from a benchmark instead.

saying these in an interview costs you the question

  • Thinks StartCPUProfile alone writes a complete file
  • Says the profiler records every function call
  • Calls os.Exit inside the profiled region
  • Expects two concurrent CPU profiles to both work
  • Profiles the whole of main and blames initialisation code
open as a page

What does a blank import of net/http/pprof add to a Go program, and how do you reach it?

level: juniorimportance: must knowfreq 52%

basics

~20 s

A blank import of net/http/pprof runs the package's init, which registers /debug/pprof handlers on http.DefaultServeMux. It starts no listener of its own, so the routes are reachable only if some server actually serves that mux.

open as a page

In `go tool pprof`, what does the `top` command list, and what do its flat and cum columns mean?

level: juniorimportance: must knowfreq 76%

basics

~20 s

top ranks functions by their share of a profile's samples. flat counts samples taken inside that function's own code; cum counts those plus everything it called. A wrapper has a tiny flat and a huge cum.

open as a page

How does Go's mutex profile differ from its block profile in what each records?

level: middleimportance: must knowfreq 50%

basics

~20 s

Go's block profile records the waiting goroutine's stack for any synchronisation wait — channels, select, mutex acquisition, WaitGroup. The mutex profile records contention on sync.Mutex and sync.RWMutex from the holder's side, attributing the waiters' delay to the unlock that released the lock.

open as a page

In a Go heap profile, what is the difference between inuse_space and alloc_space?

level: middleimportance: must knowfreq 72%

basics

~20 s

inuse_space is memory still live at the moment of the snapshot; alloc_space is every byte allocated since the process started, freed or not. One profile file carries both, and go tool pprof -sample_index chooses which you look at.

open as a page

In a Go service, throughput plateaus while CPU sits near 30% — how do block and mutex profiles find the lock?

level: seniorimportance: must knowfreq 45%

basics

~20 s

Low CPU with flat throughput means goroutines are parked, not computing. Enable both profilers coarsely on one instance and diff two captures: the mutex profile's top delay names the critical section, and the block profile separates lock waits from channel waits.

open as a page

Why is Go's block profile empty until you call runtime.SetBlockProfileRate?

level: juniorimportance: should knowfreq 30%

basics

~20 s

Go's block and mutex profilers are switched off by default because recording costs throughput. Until you call runtime.SetBlockProfileRate for blocking events, or runtime.SetMutexProfileFraction for lock contention, the runtime samples nothing and the profile comes back with no samples.

open as a page

How do you capture a Go heap profile with runtime/pprof, and why call runtime.GC() first?

level: juniorimportance: should knowfreq 58%

basics

~20 s

Open a file and call runtime/pprof.WriteHeapProfile on it, usually right after runtime.GC(), because the in-use numbers come from the most recently completed collection. In tests, go test -memprofile mem.out writes the same profile for you.

open as a page

What does runtime/pprof.Do add to a CPU profile, and how do you build its labels?

level: juniorimportance: should knowfreq 30%

basics

~20 s

runtime/pprof.Do runs a function with key/value labels attached to the current goroutine, so CPU profile samples taken during it carry those tags. Build the label set with pprof.Labels, which takes alternating key and value strings.

open as a page

How does Go's CPU profiler sample execution, and why can a short profile miss a real hotspot?

level: middleimportance: should knowfreq 45%

basics

~20 s

Go's CPU profiler is statistical: about 100 times a second per CPU-burning thread the runtime interrupts execution and records the running goroutine's stack. A function only shows up if its total accumulated CPU time is large enough to catch samples, so short profiles hide small shares.

open as a page

Why does /debug/pprof/profile?seconds=30 block for thirty seconds while /debug/pprof/heap answers at once?

level: middleimportance: should knowfreq 42%

basics

~20 s

The /debug/pprof/profile endpoint turns CPU profiling on, waits out the requested window (30 seconds by default) and then streams the samples, so the request lasts the whole window. The /debug/pprof/heap endpoint needs no window and dumps current allocation records immediately.

open as a page

Which goroutines inherit runtime/pprof labels, and when do those labels go away?

level: middleimportance: should knowfreq 42%

basics

~20 s

A goroutine inherits the labels its creator carried when the go statement ran, and keeps them for life unless it sets its own. Already-running goroutines get nothing, and pprof.Do restores labels only on its own goroutine.

open as a page

What does `go tool pprof -http=:8080 cpu.pprof` open, and how do you read its flame graph?

level: middleimportance: should knowfreq 47%

basics

~20 s

It starts a local web server and opens pprof's browser UI over that profile: Top, Graph, Flame Graph and Peek views. In the flame graph each box is a stack frame, and its width is that call path's share of samples.

open as a page

In `go tool pprof`, what do the `list` and `peek` commands show that `top` does not?

level: middleimportance: should knowfreq 54%

basics

~20 s

list <regexp> prints a matching function's source annotated with per-line flat and cum weight, so you see which statement costs. peek <regexp> prints that function's callers and callees with the share each edge carries. top only ranks whole functions.

open as a page

A Go service's p99 doubled, yet a 30-second CPU profile shows no new hot function. Why might the cost be invisible?

level: seniorimportance: should knowfreq 52%

basics

~20 s

A CPU profile samples only threads that are executing, so time a goroutine spends parked on a channel, a mutex, a syscall or a network read produces no samples at all. Latency spent waiting, or lost to CPU throttling, is invisible to that instrument by construction.

open as a page

A Go pipeline stage decoding JSON into map[string]any keeps the collector busy while its live heap stays flat. How do you use the heap and allocs profiles to find and cut the allocation volume?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Open the profile at alloc_space and alloc_objects, not the live view: a flat live heap is what churn looks like. Then cut objects per event - a concrete struct instead of map[string]any, json.RawMessage for unread fields - and re-measure.

open as a page

A Go service's internet-facing listener answers /debug/pprof/heap although no code registers that route — how, and how do you take it off?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Something in the build blank-imports net/http/pprof, whose init registers the handlers on http.DefaultServeMux, and the public server was started with a nil handler — which means exactly that mux. Give the public listener its own http.NewServeMux and serve pprof elsewhere.

open as a page

How do you split a Go CPU profile's hottest function by tenant using pprof labels?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Label each unit of work with the tenant using runtime/pprof.Do where it executes, collect a CPU profile, then in go tool pprof list the tags and use -tagfocus to keep only one tenant's samples, comparing totals across tenants.

open as a page

Would you leave Go's block and mutex profilers enabled permanently in production, and at what rates?

level: principalimportance: should knowfreq 28%

basics

~20 s

Usually yes for the mutex profiler at a coarse fraction, and off or very coarse for the block profiler. Cost scales with how often goroutines block, so measure it per service and make both rates adjustable at runtime.

open as a page

How do you get a CPU profile of one Go benchmark using go test -cpuprofile?

level: middleimportance: nice to knowfreq 38%

basics

~20 s

Run go test with -cpuprofile plus -bench naming the benchmark and -run=^$ so no ordinary tests execute, for example: go test -run=^$ -bench=BenchmarkScoreQuery -benchtime=10s -cpuprofile=cpu.out ./ranking. The flag also leaves the compiled test binary behind for symbol resolution.

open as a page

What do the arguments to runtime.SetBlockProfileRate and SetMutexProfileFraction mean?

level: middleimportance: nice to knowfreq 22%

basics

~20 s

They are not the same unit. runtime.SetBlockProfileRate takes nanoseconds: it aims for one sampled event per that many nanoseconds spent blocked, and 1 samples everything. runtime.SetMutexProfileFraction takes a fraction: on average one contention event in N is reported.

open as a page

What does runtime.MemProfileRate control in Go, and what is its default value?

level: middleimportance: nice to knowfreq 34%

basics

~10 s

It sets Go's memory-profile sampling rate: about one allocation is recorded per MemProfileRate bytes allocated, default 512 KB. Totals are scaled estimates. A rate of 1 records everything; 0 turns profiling off.

open as a page

Which Go runtime profiles carry pprof labels, and which ignore them entirely?

level: middleimportance: nice to knowfreq 24%

basics

~20 s

Only the CPU profile and the goroutine profile carry runtime/pprof labels. The heap and allocation profiles, the block profile, the mutex profile and threadcreate ignore them, so labels cannot split memory or contention by tenant.

open as a page

How do you use `go tool pprof -diff_base` on two saved CPU profiles to find what got slower between builds?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Run go tool pprof -diff_base=old.pprof new.pprof. pprof subtracts the base sample by sample, so rows are deltas: positive means the new profile spends more there, negative means less. It is only meaningful if both profiles cover comparable work.

open as a page

How do you decide whether net/http/pprof ships in a production build, and on which listener?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Treat it as a policy with three settings: always on a loopback or authenticated operator listener, never on a public one, and compiled out only where a debug surface is genuinely unacceptable. Decide once per class of service, not per incident.

open as a page