skip to content

How does Go's CPU profiler sample execution, and why can a short profile miss a real hotspot?

level: middleimportance: should knowfreq 45%

answer

  1. interrupts, not counters
  2. about a hundred times a second
  3. each hit stands for ten milliseconds
  4. share of CPU decides visibility, not call length
  5. five hundred samples cannot resolve half a percent

basics

~20 s

Go's CPU profiler is statistical: about 100 times a second per CPU-burning thread the runtime interrupts execution and records the running goroutine's stack. A function only shows up if its total accumulated CPU time is large enough to catch samples, so short profiles hide small shares.

solid answer

~50 s

It samples rather than instruments. On Unix the runtime arms per-thread CPU timers that deliver SIGPROF at roughly 100 Hz; the signal handler walks the stack of the goroutine running on the interrupted thread and pushes it into a buffer that a background goroutine encodes. Each sample therefore stands for about 10 ms of CPU time. With four busy cores you collect around 400 samples a second, and the numbers pprof reports are counts multiplied by the period, not measured durations. The consequence is statistical: a function needs enough total CPU share to be hit repeatedly. Over a five-second profile of one busy core — about 500 samples — a function using 0.5% of CPU picks up two or three samples, which is noise. Per-call duration is irrelevant: a 200 microsecond function called 50,000 times is very visible, while a rarely called one-second function may not be.

go deeper

for a junior

Know that Go's CPU profiler takes periodic snapshots rather than timing every call, and that the default is roughly 100 snapshots a second.

for a middle

Explain the mechanism end to end: CPU-time timers, a signal handler that walks the running goroutine's stack, counts converted to time, and why the overhead is fixed per second.

for a senior

Do the sample arithmetic before drawing conclusions, and say what window and load you would need to trust a claim about a function with a small percentage share.

for a principal

Frame the tradeoff for the team: higher rates and longer windows buy resolution at a cost, and no rate makes off-CPU time appear, so decide which instrument answers which question.

## Sampling, not instrumentation Go does not insert counters into your functions. Instead the runtime asks the operating system to interrupt execution at a fixed frequency and records what was running at that instant. On Unix systems that is done with CPU-time timers that deliver SIGPROF; on Windows a dedicated thread performs the same job. The mechanism differs, the model does not: periodically, whatever is on the CPU gets its stack recorded. The default rate is 100 Hz. When the signal arrives, the handler runs on the interrupted thread, unwinds the stack of the goroutine executing there, and appends the stack to a lock-free buffer. A separate goroutine drains that buffer and writes the encoded profile to the writer you supplied. Nothing in this path is proportional to the number of calls your program makes, which is why the overhead is roughly constant — a fixed number of stack walks per second — instead of scaling with call volume. ## What a sample is worth Because the timers are CPU-time timers, one sample means about 10 ms of CPU was spent in that stack. The times pprof reports are reconstructions: sample count times the sampling period. A function credited with 1.2 seconds was hit about 120 times; it was never timed. Sampling is per running thread, so the sample budget scales with how much CPU the process actually uses. One saturated core yields ~100 samples/second. Four saturated cores yield ~400. A process that is 5% busy on one core yields about 5 samples a second, no matter how long the wall-clock window is. ## Why a hotspot can disappear This is a statistics problem, and the arithmetic is worth doing out loud in an interview. A five-second profile of one busy core is about 500 samples. A function responsible for 0.5% of CPU expects 2.5 samples. The observed count could easily be 0, 1 or 5 — a factor of several, and completely indistinguishable from a function that does nothing. At 10% share, the same profile expects 50 samples and the estimate is solid. The rule of thumb that follows: to trust a function's share you want it to have collected tens of samples at minimum. To find a 1% cost in a search-ranking scorer — exactly the kind of thing you hunt when you have been told to cut cost per query by twenty percent — you need a much longer window, a much busier process, or both. ## The misconception to kill: fast functions are not invisible Candidates often say the profiler cannot see functions that return in less than 10 ms. That is wrong, and the reason matters. The profiler does not need to catch a single call; it needs the function to be on the CPU when a random interrupt lands. A normalisation routine over UTF-8 query text that takes 200 microseconds but runs 50,000 times during the profile accumulates ten seconds of CPU and will be hit constantly. Conversely a function that takes a full second but is called twice in a five-second window accumulates two seconds and might show up as 200 samples — also fine. What is invisible is small *total* share, not small per-call duration. One genuine subtlety: sampling is not perfectly uniform. Samples land where the interrupt happens to fall, and correlations between the timer and the workload's own periodicity can bias attribution, which is another reason to prefer long profiles over short ones. ## Getting more resolution The honest ways to get more signal, in order of preference: 1. **Profile for longer.** Doubling the window doubles the samples. 2. **Make the process busier.** Drive the code under test from a benchmark so the CPU is saturated by the code you care about instead of by idle time. 3. **Raise the rate.** `runtime.SetCPUProfileRate(hz)` exists, but it must be called while no profile is running, higher rates cost more, and very high rates can lose samples. It is a last resort, not a default. ## What sampling cannot fix More samples never reveal time that used no CPU. A goroutine blocked on a channel, a mutex, a network read or a file read is not running on any thread, so no timer fires for it and no sample is attributed to it. Increasing the rate from 100 Hz to 1000 Hz multiplies your resolution on CPU-bound code and changes nothing at all about waiting. That is a limit of the instrument, not of the settings. ## Inlining and stacks Inlined functions are still reported: the compiler records inline information and pprof expands it, so a small helper that was inlined into its caller still appears as its own frame. Very deep stacks are truncated at a fixed depth, so extremely recursive code can lose the top of its call chain, but ordinary application stacks are recorded whole.

  • If a function takes only 200 microseconds per call, can Go's CPU profiler see it at all?
    Yes, easily, provided it is called often. The profiler catches whatever is running when an interrupt lands, so what matters is the function's total accumulated CPU time, not its per-call duration. A 200 microsecond function called 50,000 times owns ten seconds of CPU and will collect a great many samples.
  • Why does raising the sampling rate not help when the code is waiting rather than computing?
    The timers driving the profiler are CPU-time timers: they only fire for threads that are actually executing. A goroutine parked on a channel or a syscall occupies no thread, so no interrupt can land on it. Multiplying the rate multiplies resolution on CPU work and leaves off-CPU time exactly as invisible as before.
  • Roughly how many samples does a 30-second profile of a process using two fully busy cores contain?
    About 6,000 — roughly 100 samples per second per busy core, times two cores, times thirty seconds. That is plenty to resolve anything with a few percent share, and marginal for anything below half a percent. Doing this arithmetic before trusting a profile is the habit interviewers are looking for.

saying these in an interview costs you the question

  • Says the profiler cannot see functions faster than 10 ms
  • Believes pprof measures each call's duration
  • Treats a function with three samples as a hotspot
  • Thinks overhead grows with the number of function calls
  • Suggests raising the rate to expose blocking time