skip to content

How do you get a CPU profile of one Go benchmark using go test -cpuprofile?

level: middleimportance: nice to knowfreq 38%

answer

  1. one flag, but the filters do the work
  2. keep ordinary tests out of the run
  3. give it longer than the default second
  4. setup is in the profile even after ResetTimer
  5. it leaves a .test binary behind on purpose

basics

~20 s

Run go test with -cpuprofile plus -bench naming the benchmark and -run=^$ so no ordinary tests execute, for example: go test -run=^$ -bench=BenchmarkScoreQuery -benchtime=10s -cpuprofile=cpu.out ./ranking. The flag also leaves the compiled test binary behind for symbol resolution.

solid answer

~50 s

`go test -cpuprofile=cpu.out` profiles the entire test binary run, so the filters matter more than the flag. `-run=^$` matches no test function, which keeps ordinary tests out of the samples; `-bench=BenchmarkScoreQuery` selects the one benchmark; `-benchtime=10s` extends the run so you collect thousands of samples instead of a few hundred. Two details people miss: the profile covers everything the binary did, including package-level initialisation, benchmark setup and any other benchmark the `-bench` pattern happened to match, and `b.ResetTimer()` does not exclude setup from it — it only adjusts the reported ns/op. Passing a profiling flag also makes `go test` retain the compiled test binary in the current directory, as `-c` would, because pprof needs that binary to resolve symbols. Profile one package at a time; the profiling flags are not usable across a multi-package pattern.

code

text · 1 line
text
go test -run=^$ -bench=BenchmarkScoreQuery -benchtime=10s -cpuprofile=cpu.out ./ranking

go deeper

for a junior

Know that go test accepts -cpuprofile and that you normally combine it with -bench to profile a benchmark rather than a test.

for a middle

Explain what each flag in the command contributes, and why the profile includes setup and every benchmark the pattern matched.

for a senior

Demonstrate that you sanity-check the run before trusting it: enough samples, the right benchmarks selected, setup not dominating, inputs representative of production.

for a principal

Have a position on where benchmark profiles fit against profiles from real traffic, and how much of a team's optimisation budget should be steered by fixtures you wrote yourself.

## Why profile through a benchmark at all A benchmark is the cleanest way to get a CPU profile of a specific piece of code. It saturates the CPU with the code you care about, it runs for as long as you ask, and it needs no changes to production code. If you have been told to cut the cost per query of a search-ranking scorer, a benchmark that scores a representative corpus of query strings will give you a far denser profile than an equivalent window captured from the live service, where most of the samples belong to everything else the process does. ## The command ``` go test -run=^$ -bench=BenchmarkScoreQuery -benchtime=10s -cpuprofile=cpu.out ./ranking ``` Each part is load-bearing: - `-cpuprofile=cpu.out` — write a CPU profile to that file before exiting. - `-bench=BenchmarkScoreQuery` — the pattern selecting which benchmarks run. It is an unanchored regular expression, so a loose pattern such as `-bench=Score` may run several benchmarks and blend all of them into one profile. - `-run=^$` — a regular expression matching no test name. Without it, every ordinary test in the package still runs and contributes its CPU to the same profile. - `-benchtime=10s` — run the benchmark for ten seconds instead of the default (about one second). More time means more samples, which is usually the difference between a readable profile and noise. - `./ranking` — one package. Profiling flags apply to a single package's test binary; point them at one package rather than a recursive pattern. ## The profile covers the whole binary run This is the detail most candidates get wrong. `-cpuprofile` is not scoped to the benchmark loop. Sampling starts when the test binary starts profiling and stops as it exits, so package-level `init` work, `TestMain` setup, fixture construction and every selected benchmark all land in the same profile. In particular, `b.ResetTimer()` has nothing to do with profiling. It resets the timer used to compute the reported ns/op and B/op figures. Expensive setup before `ResetTimer` is excluded from the *reported numbers* and fully included in the *profile*. If a benchmark builds a large index of UTF-8 candidate strings before scoring, that index build can dominate the profile while the benchmark output looks clean. The fixes are to move heavy setup out of the measured function, share it across runs, or lengthen `-benchtime` so the setup's fixed cost is diluted by the repeated work. ## The retained binary When any profiling flag is passed, `go test` writes the compiled test binary into the current directory, exactly as `-c` does — for the package `ranking`, that is `ranking.test`. This is not clutter, it is a requirement: a profile stores addresses and needs the matching binary to turn them into function names and line numbers, and the tool will not symbolise correctly against a different build. Keep the pair together, and expect to see the `.test` file in `git status` after profiling. ## Related flags `go test` exposes the same shape for other data: `-memprofile`, `-blockprofile`, `-mutexprofile` and `-trace` each write their own artefact from the same run, and `-o` controls where the retained binary lands. Combining `-cpuprofile` with `-benchmem` is common — `-benchmem` adds allocation columns to the printed results while the CPU profile keeps recording time. ## Sanity checks before you believe the numbers - Did the benchmark actually run long enough? Ten seconds of a saturated core is about a thousand samples. - Did the pattern select only what you meant? Re-read the `-bench` regular expression. - Is the compiler optimising your work away? A benchmark whose result is never used can be eliminated; assign to a package-level sink so the work survives. - Is the profile dominated by setup? Compare the benchmark's own reported time against the total CPU in the profile. ## When a benchmark is the wrong vehicle A benchmark profiles the code you wrote a benchmark for, with the inputs you chose. If real queries have a different length distribution or a different mix of scripts than your fixture, the profile is honest about the benchmark and misleading about production. Treat it as a fast, precise instrument for a hypothesis you already have, and cross-check the shape against a profile taken from a real workload before spending a week on what it points at.

  • Does b.ResetTimer() keep benchmark setup out of the CPU profile?
    No. ResetTimer only affects the timing and allocation figures the benchmark reports. The CPU profile covers the entire test binary run, so setup before ResetTimer is still sampled. If setup is expensive, move it out of the benchmark, share it across iterations, or raise -benchtime so the repeated work dominates.
  • Why does go test leave a .test binary in the directory when you pass -cpuprofile?
    Because pprof needs it. A profile records addresses, and turning them into function names and line numbers requires the exact binary that produced them. Passing a profiling flag makes go test retain the compiled test binary as -c would, so the profile can be symbolised against the right build.
  • You pass -bench=Score and the profile is dominated by a function you did not expect. What is the first thing to check?
    Which benchmarks actually ran. The -bench value is an unanchored regular expression, so Score matches BenchmarkScoreQuery, BenchmarkScoreCorpus and anything else containing it, and all of them share one profile. Tighten the pattern, and check that ordinary tests were excluded with -run=^$.

saying these in an interview costs you the question

  • Believes -cpuprofile only covers the benchmark loop
  • Thinks b.ResetTimer excludes setup from the profile
  • Runs benchmarks for the default second and trusts the result
  • Deletes the retained .test binary before symbolising
  • Points profiling flags at a recursive multi-package pattern