A pull request claims a 20% speedup from benchstat output produced on a shared CI runner. How do you check it?
answer
- the machine is the instrument
- compare the commit against itself first
- noise floor before any claim
- alternate the runs, do not block them
- allocation columns barely feel the noise
basics
~20 sReproduce it before believing it. Run the unchanged commit into both files to learn the harness's noise floor, then interleave old and new runs on one quiet machine with a high -count, and check whether the allocation columns moved too.
solid answer
~50 sA shared runner is the wrong instrument: the two files were produced minutes apart on a machine whose speed depends on co-tenants, frequency scaling and thermal state, so a 20% delta with a small p-value can be entirely the machine. First I ask how it was measured — one machine or two, back to back or interleaved, what `-count` and `-benchtime`. Then I calibrate: run the *same* unchanged commit into `old.txt` and `new.txt` and look at what benchstat reports for a change of nothing. If that A/A run already produces significant-looking deltas, the harness cannot support the claim. Then I re-measure properly: alternate old and new runs, appending to the two files, on an idle machine with `-count=10` or more. I also check `B/op` and `allocs/op` — allocation counts are deterministic, so a real algorithmic win usually shows there as well, and a timing win with no allocation change on a noisy box deserves suspicion.
code
text · 5 linesfor i in 1 2 3 4 5; do
git checkout -q main && go test -run=^$ -bench=Encode -benchtime=2s -count=2 >> old.txt
git checkout -q perf && go test -run=^$ -bench=Encode -benchtime=2s -count=2 >> new.txt
done
benchstat old.txt new.txtgo deeper
Remember the one rule that survives everything: both numbers must come from the same idle machine, and one run per side is never enough to claim anything.
Explain why a shared runner breaks the statistics rather than just adding scatter — the noise is correlated with time, so block-then-block measurement turns machine drift into a fake, significant delta.
Demonstrate the A/A calibration run and interleaving as standard practice, and use the deterministic allocation columns as corroboration for a timing claim you cannot fully trust.
Decide where benchmarks are allowed to run at all, and whether the cost of a dedicated machine is worth it, given that measurements taken anywhere else cannot settle the arguments your team keeps having.
## Why the runner is the problem Shared CI machines are built for throughput, not for measurement. Several jobs share cores, the kernel migrates goroutines between them, turbo and thermal throttling change the effective clock during the run, and virtualised instances may be credit-limited so that a burst of speed is followed by a slow patch. None of that is random in the way a statistical test assumes: it is *correlated with time*. If `old.txt` was collected while a neighbouring build was linking and `new.txt` after it finished, benchstat sees a large, consistent, highly significant improvement — and it is entirely the neighbour. This is the trap the p-value cannot catch. Statistics defend you against random jitter; they cannot defend you against a systematically different environment on each side. ## Step 1: ask how it was measured Before doing any work, ask three questions. Were both files produced on the same machine and the same toolchain? Were the runs back to back — the whole baseline, then the whole new revision — or interleaved? What were `-count` and `-benchtime`? A claim built from `-count=1` on each side has no statistics behind it at all, and a claim built from two consecutive blocks has time confounded with the change. ## Step 2: measure the noise floor with an A/A run The single most useful move: compare a commit with **itself**. Run the unchanged code into `old.txt`, run the unchanged code again into `new.txt`, and hand both to benchstat. Everything it reports is noise by construction. If that A/A comparison comes back with tildes and deltas under 1%, the harness can resolve small effects and a 20% claim is plausible. If it comes back with a significant 12% "improvement" for no code change at all, you have learned that this environment cannot support any claim smaller than that, and the pull request's evidence is worthless regardless of its p-value. Run the A/A on the runner in question — the point is to characterise *that* machine, not an ideal one. ## Step 3: re-measure properly Two controls fix most of the damage. **Interleave.** Instead of ten baseline runs followed by ten new runs, alternate them and append to the two files. Any slow drift in the machine then lands on both sides equally instead of being attributed to the change. ``` for i in 1 2 3 4 5; do git checkout -q main && go test -run=^$ -bench=Encode -benchtime=2s -count=2 >> old.txt git checkout -q perf && go test -run=^$ -bench=Encode -benchtime=2s -count=2 >> new.txt done benchstat old.txt new.txt ``` **Quieten the machine.** An idle host, no other jobs, and where you control it, frequency scaling pinned. This is why teams that care keep one unshared machine for benchmarks rather than trying to statistically correct a busy one. Raise `-count` to give the test more independent samples, and raise `-benchtime` so each sample runs long enough that scheduler quanta and a GC cycle or two average out inside it. They fix different things: `-benchtime` makes each sample steadier; only `-count` adds samples for the significance test to consume. ## Step 4: cross-check with something that does not depend on the clock Allocation counts are deterministic for a given input: the same code path allocates the same number of times regardless of what else the machine is doing. So `allocs/op` and `B/op` from `-benchmem` are nearly noise-free, and they are excellent corroboration. A change that claims to be 20% faster because it stopped allocating a slice per call should show that drop in `allocs/op` on every run. A 20% timing win with allocation columns showing `~` is not disproven, but it now needs an explanation. ## Step 5: sanity-check the benchmark itself Two more failure modes worth a glance. Did the benchmark change in the same pull request? Then the two files measure different work and the comparison is meaningless. And is the result of the work still used — if the change let the compiler eliminate the computation entirely, a spectacular delta is measuring nothing. A win far larger than the mechanism can explain deserves a look at the benchmark body before it deserves congratulations. ## What you say in the review Not "I do not believe you". Something reproducible: *this was measured on a shared runner, here is the A/A run showing that environment reports plus or minus 9% for no change, please re-measure interleaved on the dedicated box and repost the table with n and p.* That converts an argument about numbers into a procedure anyone can repeat.
- Why does interleaving the runs matter if you already have twenty samples per side?Because twenty samples taken in one block share whatever the machine was doing during that block. If a neighbouring job started halfway through, the baseline block and the new block sit in different worlds and the test reports that difference as your change. Alternating spreads any drift across both files, so time is no longer confounded with the revision under test.
- To make a comparison more convincing, do you raise -count or -benchtime?-count, if the goal is statistical evidence: each repetition adds an independent sample, and the significance test consumes samples. -benchtime lengthens each individual measurement, which smooths jitter inside a sample and helps when a benchmark is so short that timer resolution or a single GC cycle dominates. In practice you raise -benchtime until each sample is stable, then raise -count until the comparison is conclusive.
- The timing improved by 20% but allocs/op is unchanged. Does that sink the claim?Not by itself, but it shifts the burden. Many real wins do not touch allocation — better branch behaviour, fewer bounds checks, a cheaper hash. What it removes is the noise-free corroboration, so the timing evidence has to stand alone and therefore has to be strong: quiet machine, interleaved, high -count, and a mechanism the author can explain.
saying these in an interview costs you the question
- Accepts a p-value as proof the environment was sound
- Runs all baseline samples then all new samples
- Never checks what the harness reports for no change
- Compares files produced on two different machines
- Ignores that the benchmark itself changed in the same PR