skip to content

Optimising with Evidence

The loop around a change: repeat a benchmark until the delta is real, read what the compiler inlined or moved to the heap, and feed a production profile back to it.

part ofGo (Golang)overview, primer and where to startread it →
on this pageshow

explore

questions

12

How do you use benchstat to compare Go benchmark results from two revisions of a package?

level: juniorimportance: must knowfreq 45%

answer

  1. two files, not two numbers
  2. several samples per revision
  3. -count is what makes samples
  4. rows pair by benchmark name
  5. installs from golang.org/x/perf

basics

~20 s

Run go test -bench with -count several times on each revision, saving the output to old.txt and new.txt, then run benchstat old.txt new.txt. benchstat pairs benchmarks by name and prints the percent change with a p-value.

solid answer

~40 s

benchstat compares distributions, not single numbers, so the first job is producing several samples per revision. On the baseline I run something like `go test -run=^$ -bench=Encode -count=10 -benchmem > old.txt`, then apply the change and run the identical command into `new.txt`. `benchstat old.txt new.txt` reads both files, pairs rows by benchmark name (including the `-8` GOMAXPROCS suffix), and for each metric — the time column, plus `B/op` and `allocs/op` when `-benchmem` was used — prints the two medians, the percent change, and `p=… n=…`. A `~` in the delta column means the difference was not statistically significant. A `geomean` row summarises the whole table. benchstat is not part of the toolchain; install it with `go install golang.org/x/perf/cmd/benchstat@latest`.

code

text · 5 lines
text
git stash -q
go test -run=^$ -bench=Encode -count=10 -benchmem > old.txt
git stash pop -q
go test -run=^$ -bench=Encode -count=10 -benchmem > new.txt
benchstat old.txt new.txt

go deeper

for a junior

Be ready to name the two steps out loud: several -count runs per revision saved to files, then benchstat over both files. Know that benchstat is installed from golang.org/x/perf, not bundled with the go command.

for a middle

Explain what a sample is, why -run=^$ and -benchmem belong on the command, and how rows pair by benchmark name including the GOMAXPROCS suffix. Say why both files must come from one machine and one toolchain.

for a senior

Treat the comparison as a designed experiment: fixed environment, enough samples, identical commands, and the raw files kept so a reviewer can re-run it instead of trusting a pasted number.

for a principal

Own what counts as proof of a speedup on your team — sample files and a benchstat table rather than a screenshot — and decide where those runs execute so numbers from different pull requests are comparable at all.

## Why two numbers are not a comparison A Go benchmark prints a time per operation for one binary, on one machine, at one moment. Run the identical binary again and the number moves: CPU frequency scaling, co-tenant processes, memory layout, and garbage-collection timing all shift it by a few percent and sometimes much more. So comparing one number from before a change against one number from after tells you almost nothing — the difference you see is *the change plus the noise*, and a subtraction cannot separate them. `benchstat`, from the `golang.org/x/perf` repository, does that separation. It reads *samples* — repeated measurements of the same benchmark — from two or more files, and reports per benchmark how far the metric moved and how confident it is that the move is real. ## Step one: produce the samples One repetition of a benchmark produces one output line, which is one sample. `-count=N` repeats every selected benchmark N times, so `-count=10` gives ten samples: ``` go test -run=^$ -bench=Encode -count=10 -benchmem > old.txt ``` - `-run=^$` is a regular expression that matches no test name, so only benchmarks run. Without it the package's unit tests execute first and add noise and time. - `-bench=Encode` selects benchmarks whose name matches; `-bench=.` selects all of them. - `-count=10` produces the samples. Ten is the usual starting point; noisy machines need more. - `-benchmem` adds the bytes-per-operation and allocations-per-operation columns, which are worth having on every comparison because they are far steadier than time. The file is just the raw text output. benchstat parses the `goos:`, `goarch:` and `pkg:` header lines and the benchmark result lines, and ignores the rest. Then check out or apply the change and run **the same command** into `new.txt`. ## Step two: compare ``` benchstat old.txt new.txt ``` You get one row per benchmark, a column per input file, a `vs base` percentage delta, and a `p=… n=…` annotation. `~` in the delta column means the test could not distinguish the two sample sets. The bottom `geomean` row combines the table into a single overall figure, which is useful when a change moves several benchmarks in different directions. benchstat accepts one file (it just summarises it) or more than two (each extra file becomes another column compared against the first). ## What keeps the two files comparable - **The same machine**, ideally an idle one. Two files produced on different hardware are not a comparison of your change. - **The same toolchain version**, since compiler changes move benchmark numbers on their own. - **The same flags**, especially `-benchtime` and `-benchmem`. - **The same benchmark names.** benchstat pairs rows by name; if you rename `BenchmarkEncode` while you are at it, you get two unpaired rows and no delta. - **The same GOMAXPROCS.** The `-8` suffix on `BenchmarkEncode-8` is the GOMAXPROCS value the benchmark ran with, and it is part of the row's identity — samples taken on a 8-core and a 16-core machine will not pair. ## Installing it benchstat does not ship with the `go` command: ``` go install golang.org/x/perf/cmd/benchstat@latest ``` It then lives in `$(go env GOPATH)/bin`. ## What benchstat will not do for you It reports *whether* something moved, never *why*. It cannot tell you whether the benchmark resembles production traffic, and it cannot rescue an experiment that had too few samples or ran on a machine that was doing other work. Those judgments stay with you; benchstat only makes sure you are not reading noise as a result.

  • Why does benchstat want `-count=10` rather than one run per revision?
    With one measurement per side there is no distribution to test: benchstat cannot estimate how much the machine wobbles, so it cannot say whether the gap between the two numbers is your change or the afternoon. Ten samples give the comparison enough evidence to distinguish a real shift from ordinary jitter, and they spread the measurement over time so one unlucky moment does not decide the result.
  • What happens if you rename a benchmark function between the two runs?
    benchstat pairs rows by benchmark name, so a rename produces two unpaired rows — the old name present only in the first file, the new one only in the second — and no delta or p-value for either. Keep benchmark names stable across the revisions you are comparing, and rename in a separate commit.
  • What does the `-8` suffix in `BenchmarkEncode-8` mean for a comparison?
    It is the GOMAXPROCS value the benchmark ran with, and benchstat treats it as part of the row name. Files produced on machines with different core counts, or with different `-cpu` values, therefore produce unpaired rows instead of a comparison. Fix GOMAXPROCS or run both sides on the same machine.

saying these in an interview costs you the question

  • Compares a single ns/op number against another
  • Produces the two files on different machines
  • Believes benchstat ships with the go toolchain
  • Uses -count=1 and reports the percentage difference
  • Renames benchmarks between runs and loses the pairing
open as a page

What does the Go compiler do differently at a call site that a PGO profile marks hot?

level: middleimportance: must knowfreq 44%

basics

~20 s

Two things: it inlines hot callees it would otherwise judge too expensive, and it devirtualizes hot indirect calls by emitting a direct call to the dominant concrete type behind a type check, with the original indirect call as the fallback. Cold code is left alone.

open as a page

How do you make `go build` print the Go compiler's inlining and escape-analysis decisions?

level: juniorimportance: should knowfreq 50%

basics

~10 s

Run go build -gcflags=-m ./... or go test -gcflags=-m. The compiler prints its inlining and escape decisions to stderr: can inline, inlining call to, cannot inline, moved to heap. Use -m=2 for more detail.

open as a page

Where do you put a Go PGO profile so that `go build` uses it without extra flags?

level: juniorimportance: should knowfreq 24%

basics

~20 s

Save a pprof CPU profile as a file named default.pgo in the main package's directory. The go command defaults to -pgo=auto, so build, run, install and test pick that file up with no flag. Use -pgo=off to opt out.

open as a page

In benchstat's comparison output, what do the p-value, the n count, and a ~ in the delta column tell you?

level: middleimportance: should knowfreq 38%

basics

~20 s

The p-value is the chance of seeing a difference this large if both revisions performed the same, and n is the samples per side. A ~ means the difference was not significant at 95% confidence, not that the two are equal.

open as a page

What makes the Go compiler refuse to inline a function, and how does the `cannot inline` line say so?

level: middleimportance: should knowfreq 45%

basics

~20 s

Go's inliner scores a function body against a fixed cost budget of 80 and refuses anything over it, printing cannot inline f: function too complex. A //go:noinline directive, an assembly body, or a construct it cannot handle also blocks inlining.

open as a page

A pull request claims a 20% speedup from benchstat output produced on a shared CI runner. How do you check it?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Reproduce it before believing it. Run the unchanged commit into both files to learn the harness's noise floor, then interleave old and new runs on one quiet machine with a high -count, and check whether the allocation columns moved too.

open as a page

`go build -gcflags=-m` shows a scratch buffer in your encoder escaping to the heap. How do you find out why?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Read the escape line together with the inlining lines around it: look for cannot inline on the helper the buffer is passed to, and a leaking param note on its parameter. Then split that helper and re-measure allocs/op.

open as a page

The default.pgo committed beside your Go service is three releases old. Is the build still correct, and does the profile still help?

level: seniorimportance: should knowfreq 34%

basics

~20 s

The build stays correct: a stale profile can only lose value, never produce wrong code. It degrades gradually as functions are renamed, split or deleted and their samples stop matching. Confirm the remaining benefit by building the same commit with -pgo=off and -pgo=auto and comparing under the same traffic.

open as a page

What does the `//go:noinline` directive do, and when would you deliberately add one?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

//go:noinline is a compiler directive on the line immediately above a function declaration, no space after the slashes. It forbids inlining that function anywhere. Its honest use is experimental: turning inlining off to see what it was worth.

open as a page

What evidence should a Go benchmark speedup claim carry before you merge it, and should regressions fail CI?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Require the exact command, the sample files or a benchstat table showing n and p, both sides measured on one machine, and a delta larger than that machine's measured noise. Gate CI on deterministic allocation metrics; track timing as a reported trend.

open as a page

Who owns refreshing a Go service's committed default.pgo, and when would you ban the committed profile outright?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Treat default.pgo as a build input with a named owner, a documented capture source and a refresh cadence tied to releases. Ban it when nobody will own the refresh, when its provenance cannot be audited, or when the build-time and reproducibility cost outweighs a few percent of CPU.

open as a page