In benchstat's comparison output, what do the p-value, the n count, and a ~ in the delta column tell you?
answer
- a statistical test, not a subtraction
- assume both revisions equal, then measure surprise
- tilde means not detected, not equal
- n is the honesty column
- significant and large are different questions
basics
~20 sThe p-value is the chance of seeing a difference this large if both revisions performed the same, and n is the samples per side. A ~ means the difference was not significant at 95% confidence, not that the two are equal.
solid answer
~50 sbenchstat does not subtract two averages — it runs a statistical test, by default the non-parametric Mann-Whitney U test, on the two sets of samples. The p-value is the probability of observing a difference at least this large if the two revisions came from the same distribution; benchstat's default confidence is 95%, so p below 0.05 is reported as a percentage delta and anything above it prints as `~`. `n` says how many samples from each file went into that row, which is your reality check on whether the test had any power. The `±` figure shows how spread the samples were: a large `±` means a noisy machine and a weak conclusion. Crucially, significance is not magnitude — `-0.4% (p=0.000)` is a real and useless result, while a 20% win at `n=3` may be real and simply unproven.
code
text · 3 linesEncode-8 1.203µ ± 2% 0.972µ ± 3% -19.20% (p=0.000 n=10)
Decode-8 842.1n ± 1% 838.4n ± 2% ~ (p=0.353 n=10)
geomean 1.006µ 903.0n -10.24%go deeper
Learn the three symbols before the theory: the percentage is how much it moved, p says whether to believe it, and a tilde means no difference was detected. Do not report a tilde row as an improvement.
Explain the hypothesis being tested, why the default test is rank-based rather than mean-based given the shape of timing data, and why n and the ± spread decide how much the p-value is worth.
Show the judgment that separates significance from importance: refuse to ship complexity for a well-measured 0.4%, and refuse to abandon a promising change on an underpowered tilde.
Set what your team is allowed to claim from a table like this, so pull request descriptions stop turning tilde rows into speedups and stop turning tiny significant deltas into headline numbers.
## benchstat is running a hypothesis test The mental model that makes the output readable: benchstat holds two *sets* of numbers for each benchmark — say ten samples from `old.txt` and ten from `new.txt` — and asks a single question. *If these two sets had actually been drawn from the same underlying distribution, how surprising would a gap this big be?* The answer to that question is the p-value. The default test is the Mann-Whitney U test, a rank-based, non-parametric test. Non-parametric matters here: benchmark timings are not normally distributed — they have a floor and a long right tail, because a run can be slowed by an interrupt but never sped up below the work's cost. A test that assumed normality would be wrong about exactly the data you are feeding it. ## Reading each piece **The delta**, e.g. `-19.20%`, is how far the new column's central value sits from the base column's. It is the size of the effect, and it is the only part of the row that tells you whether the change is worth having. **`p=0.000`** is the significance. Small p means "a gap this large would be surprising if nothing changed". benchstat's default confidence level is 95%, so the cut-off is 0.05. It is *not* the probability that your change is good, and it is *not* the size of the improvement — the single most common misreading. **`n=10`** is how many samples per side went into the row. This is the honesty column. A stunning delta at `n=3` is an anecdote; a modest delta at `n=20` on a quiet machine is a result. **`~`** is printed instead of a delta when p exceeds the threshold. It means *we failed to detect a difference*, which has two very different causes: there really is no meaningful change, or your experiment was too noisy or too small to see one. A `~` is never proof of equivalence. **`± 3%`** describes the spread of the samples behind a column. Read it before you read the delta: if the base column's samples wobble by ±8%, a 5% delta was never going to be visible, whatever the p-value says. **`geomean`** is the bottom row: the benchmarks in the table combined with a geometric mean, so a change that helps some benchmarks and hurts others gets one overall number. A geometric mean is the right average for ratios — it does not let one benchmark that got 10x faster drown out five that got slightly slower the way an arithmetic mean would. It is only meaningful if the benchmarks in the table are all ones you care about; a suite padded with irrelevant micro-benchmarks produces a flattering geomean. ## The two failure modes **Chasing significance.** `-0.35% (p=0.001 n=20)` is a genuine, reproducible, statistically solid improvement that no user will ever perceive, and defending a complicated change with it is a bad trade. Significance answers "is it real"; only the delta answers "is it worth it". **Dismissing on a tilde.** A promising change that comes back `~` has not been refuted. Look at `n` and `±` first: raise `-count`, move to a quieter machine, and re-run before concluding the idea was worthless. ## A worked reading ``` Encode-8 1.203µ ± 2% 0.972µ ± 3% -19.20% (p=0.000 n=10) Decode-8 842.1n ± 1% 838.4n ± 2% ~ (p=0.353 n=10) ``` Encode is 19% faster, the samples were tight, and the result is very unlikely to be chance — believe it. Decode moved by a fraction of a percent and the test declined to call it; treat Decode as unchanged, and specifically do not report "Decode also improved slightly" in the pull request. ## Where the p-value stops helping The test only knows about the numbers it was handed. If both files were produced on the same busy shared machine, but the base ran while a build was hogging the cores and the new one ran after it finished, the test will happily report a large effect with a tiny p-value. Statistics protect you from random noise, not from a biased experiment.
- benchstat reports -0.4% with p=0.000. Do you call that a speedup in your pull request?I would report it accurately and not sell it. The result is real — a tiny effect measured well enough to be certain of — but four tenths of a percent buys nothing and does not justify extra complexity. Significance says the measurement is trustworthy; the delta says whether anyone benefits. If the change also makes the code simpler I keep it on those grounds, not on the benchmark.
- Your change comes back as ~. Is that evidence the change made no difference?No. A tilde says the test could not distinguish the two sample sets, which happens both when there is no effect and when the experiment lacked power. Check n and the ± spread first: with few samples or a noisy machine, even a genuine 10% win can hide. Raise -count, run on a quiet machine, re-measure, and only then conclude the idea did nothing.
- Why does benchstat use a non-parametric test rather than a t-test by default?Benchmark timings are not normally distributed. They have a hard floor at the real cost of the work and a long right tail from interrupts, scheduling and GC, so the mean is dragged around by outliers. The Mann-Whitney U test compares ranks rather than means, which makes it robust to that tail and to the occasional wildly slow sample.
saying these in an interview costs you the question
- Reads the p-value as the size of the improvement
- Treats ~ as proof the two revisions are identical
- Calls any p below 0.05 a meaningful win
- Ignores n and the ± spread entirely
- Thinks benchstat just averages and subtracts