How do you use benchstat to compare Go benchmark results from two revisions of a package?
answer
- two files, not two numbers
- several samples per revision
- -count is what makes samples
- rows pair by benchmark name
- installs from golang.org/x/perf
basics
~20 sRun go test -bench with -count several times on each revision, saving the output to old.txt and new.txt, then run benchstat old.txt new.txt. benchstat pairs benchmarks by name and prints the percent change with a p-value.
solid answer
~40 sbenchstat compares distributions, not single numbers, so the first job is producing several samples per revision. On the baseline I run something like `go test -run=^$ -bench=Encode -count=10 -benchmem > old.txt`, then apply the change and run the identical command into `new.txt`. `benchstat old.txt new.txt` reads both files, pairs rows by benchmark name (including the `-8` GOMAXPROCS suffix), and for each metric — the time column, plus `B/op` and `allocs/op` when `-benchmem` was used — prints the two medians, the percent change, and `p=… n=…`. A `~` in the delta column means the difference was not statistically significant. A `geomean` row summarises the whole table. benchstat is not part of the toolchain; install it with `go install golang.org/x/perf/cmd/benchstat@latest`.
code
text · 5 linesgit stash -q
go test -run=^$ -bench=Encode -count=10 -benchmem > old.txt
git stash pop -q
go test -run=^$ -bench=Encode -count=10 -benchmem > new.txt
benchstat old.txt new.txtgo deeper
Be ready to name the two steps out loud: several -count runs per revision saved to files, then benchstat over both files. Know that benchstat is installed from golang.org/x/perf, not bundled with the go command.
Explain what a sample is, why -run=^$ and -benchmem belong on the command, and how rows pair by benchmark name including the GOMAXPROCS suffix. Say why both files must come from one machine and one toolchain.
Treat the comparison as a designed experiment: fixed environment, enough samples, identical commands, and the raw files kept so a reviewer can re-run it instead of trusting a pasted number.
Own what counts as proof of a speedup on your team — sample files and a benchstat table rather than a screenshot — and decide where those runs execute so numbers from different pull requests are comparable at all.
## Why two numbers are not a comparison A Go benchmark prints a time per operation for one binary, on one machine, at one moment. Run the identical binary again and the number moves: CPU frequency scaling, co-tenant processes, memory layout, and garbage-collection timing all shift it by a few percent and sometimes much more. So comparing one number from before a change against one number from after tells you almost nothing — the difference you see is *the change plus the noise*, and a subtraction cannot separate them. `benchstat`, from the `golang.org/x/perf` repository, does that separation. It reads *samples* — repeated measurements of the same benchmark — from two or more files, and reports per benchmark how far the metric moved and how confident it is that the move is real. ## Step one: produce the samples One repetition of a benchmark produces one output line, which is one sample. `-count=N` repeats every selected benchmark N times, so `-count=10` gives ten samples: ``` go test -run=^$ -bench=Encode -count=10 -benchmem > old.txt ``` - `-run=^$` is a regular expression that matches no test name, so only benchmarks run. Without it the package's unit tests execute first and add noise and time. - `-bench=Encode` selects benchmarks whose name matches; `-bench=.` selects all of them. - `-count=10` produces the samples. Ten is the usual starting point; noisy machines need more. - `-benchmem` adds the bytes-per-operation and allocations-per-operation columns, which are worth having on every comparison because they are far steadier than time. The file is just the raw text output. benchstat parses the `goos:`, `goarch:` and `pkg:` header lines and the benchmark result lines, and ignores the rest. Then check out or apply the change and run **the same command** into `new.txt`. ## Step two: compare ``` benchstat old.txt new.txt ``` You get one row per benchmark, a column per input file, a `vs base` percentage delta, and a `p=… n=…` annotation. `~` in the delta column means the test could not distinguish the two sample sets. The bottom `geomean` row combines the table into a single overall figure, which is useful when a change moves several benchmarks in different directions. benchstat accepts one file (it just summarises it) or more than two (each extra file becomes another column compared against the first). ## What keeps the two files comparable - **The same machine**, ideally an idle one. Two files produced on different hardware are not a comparison of your change. - **The same toolchain version**, since compiler changes move benchmark numbers on their own. - **The same flags**, especially `-benchtime` and `-benchmem`. - **The same benchmark names.** benchstat pairs rows by name; if you rename `BenchmarkEncode` while you are at it, you get two unpaired rows and no delta. - **The same GOMAXPROCS.** The `-8` suffix on `BenchmarkEncode-8` is the GOMAXPROCS value the benchmark ran with, and it is part of the row's identity — samples taken on a 8-core and a 16-core machine will not pair. ## Installing it benchstat does not ship with the `go` command: ``` go install golang.org/x/perf/cmd/benchstat@latest ``` It then lives in `$(go env GOPATH)/bin`. ## What benchstat will not do for you It reports *whether* something moved, never *why*. It cannot tell you whether the benchmark resembles production traffic, and it cannot rescue an experiment that had too few samples or ran on a machine that was doing other work. Those judgments stay with you; benchstat only makes sure you are not reading noise as a result.
- Why does benchstat want `-count=10` rather than one run per revision?With one measurement per side there is no distribution to test: benchstat cannot estimate how much the machine wobbles, so it cannot say whether the gap between the two numbers is your change or the afternoon. Ten samples give the comparison enough evidence to distinguish a real shift from ordinary jitter, and they spread the measurement over time so one unlucky moment does not decide the result.
- What happens if you rename a benchmark function between the two runs?benchstat pairs rows by benchmark name, so a rename produces two unpaired rows — the old name present only in the first file, the new one only in the second — and no delta or p-value for either. Keep benchmark names stable across the revisions you are comparing, and rename in a separate commit.
- What does the `-8` suffix in `BenchmarkEncode-8` mean for a comparison?It is the GOMAXPROCS value the benchmark ran with, and benchstat treats it as part of the row name. Files produced on machines with different core counts, or with different `-cpu` values, therefore produce unpaired rows instead of a comparison. Fix GOMAXPROCS or run both sides on the same machine.
saying these in an interview costs you the question
- Compares a single ns/op number against another
- Produces the two files on different machines
- Believes benchstat ships with the go toolchain
- Uses -count=1 and reports the percentage difference
- Renames benchmarks between runs and loses the pairing