skip to content

What evidence should a Go benchmark speedup claim carry before you merge it, and should regressions fail CI?

level: principalimportance: nice to knowfreq 26%

answer

  1. two decisions, not one
  2. review bar versus machine gate
  3. a red build that lies gets muted
  4. gate what is deterministic
  5. publish the machine's noise floor once

basics

~20 s

Require the exact command, the sample files or a benchstat table showing n and p, both sides measured on one machine, and a delta larger than that machine's measured noise. Gate CI on deterministic allocation metrics; track timing as a reported trend.

solid answer

~50 s

The bar has to be cheap for the pull requests that make no performance claim, and firm for the ones that do. If a description says "20% faster", it must carry the command used, both sides measured on the same machine, and a benchstat table with `n` and `p` rather than two pasted numbers, plus a delta bigger than that machine's known noise floor. On gating: a red build that lies gets muted, and a muted gate is worse than none because it launders false confidence. So I gate on what is deterministic — `allocs/op` and `B/op` on a short list of hot-path benchmarks — and publish timing as a tracked trend, promoting it to a blocking check only if we own a dedicated, quiet machine. If the lead still finds it too flaky, I would rather delete the gate and keep the trend.

go deeper

for a junior

When you claim a speedup, include the command you ran and the benchstat table rather than two timings, and say which machine produced both sides.

for a middle

Be able to argue why allocation counts can be gated but timings usually cannot, and what makes a failing check actionable enough that people fix it instead of retrying it.

for a senior

Show that you weigh the cost of a false red against the cost of a missed regression, and that you would rather narrow or downgrade a gate than keep one the team routinely overrides.

for a principal

Own the policy end to end: what evidence a claim must carry, which benchmarks are gated, on what hardware, who pays for it, and what you do when the lead is right that the gate is costing more than it saves.

## Two separate decisions This question hides two calls that get muddled: what a *person* must show to have a performance claim accepted in review, and what a *machine* is allowed to block a merge for. They have different economics and should be answered separately. ## The review bar The cost of the bar falls on every author, so it must be proportionate. Most changes claim nothing about performance and need nothing. The bar applies when a description asserts a number, or when a change is *justified* by performance — complexity accepted, an abstraction removed, a cache introduced. What such a claim should carry: - **The command.** Which benchmarks, what `-count`, what `-benchtime`, whether `-benchmem` was on. Without it the result is not reproducible and the reviewer cannot judge whether it had any power. - **A benchstat table, not two numbers.** With `n` and `p` visible. Two pasted per-operation figures are the artefact this whole discipline exists to reject. - **One machine, ideally interleaved.** Stated explicitly, because it is the assumption most often violated and never mentioned. - **A delta above the known noise floor of that machine.** Which presumes somebody has measured that floor by comparing a commit against itself. Publishing that number once, per benchmark machine, is one of the highest-leverage things a team can do here — it converts "is 6% real?" from an argument into a lookup. - **The benchmark itself, committed.** A win measured with a throwaway benchmark cannot be re-checked next quarter and cannot regress detectably. And what you deliberately do *not* require: a formal statistical case for every micro-refactor. If the bar makes ordinary work expensive, authors stop measuring rather than start. ## The gating decision Now the harder call. A blocking benchmark gate has a failure mode that most quality gates do not: it can be *wrong at random*. A flaky red build teaches the team that red means nothing, and that lesson transfers to the gates that were telling the truth. So the question is not "would a gate be nice" but "can this gate be right often enough to keep its authority". That depends almost entirely on the hardware you can give it. **On shared runners**, timing gates cannot be right often enough. The machine's own variation is frequently larger than the regressions worth catching. But not everything a benchmark reports is noisy: allocation counts and bytes per operation are deterministic for a given input. A gate that fails when `allocs/op` on a listed hot-path benchmark rises above a recorded value is reproducible, explains itself precisely, and catches a large fraction of real regressions — accidentally boxing a value, an added copy, a logging call in a loop. That is the gate I would build first. **On dedicated, quiet hardware**, a timing gate becomes defensible: run it nightly rather than per pull request, alert rather than block for the first quarter, set the threshold well above the measured A/A noise floor, and only then let it block. Bisect-friendliness matters more than immediacy here — knowing which of yesterday's twenty merges cost you 8% is usually enough. **Whatever the gate, it must show its work.** A failure that prints the benchmark name and the benchstat table is actionable. A failure that says "performance check failed" gets a retry, then an exemption, then deletion. ## Owning the tradeoff, and being overruled This is where the call stops being technical. A tech lead who says "the benchmark gate has cost us four hours of retries this month, turn it off" is making a legitimate argument about throughput, and the honest answer is usually to concede the gate and keep the signal: move to alert-only, narrow it to the two benchmarks that guard the actual hot path, or spend the money on a machine that makes the gate trustworthy. Insisting on a gate that the team has already learned to ignore is worse than losing the argument, because it leaves everyone believing they are protected. The other cost worth naming out loud: benchmark suites accumulate. Every benchmark added to a gate is a permanent tax on every merge, and most of them do not guard anything anyone notices. Keeping the gated list short, tied to code paths with a stated performance requirement, is what keeps the whole arrangement affordable. ## The last question Before any of this, ask whether the benchmark represents something users experience. A suite that gates on micro-benchmarks nobody's traffic resembles will happily block a change that would have made production faster. The evidence bar and the gate are both instruments for a purpose, and the purpose is the service's behaviour, not the suite's numbers.

  • Your lead says the benchmark gate is too flaky and wants it removed. What do you propose?
    I concede the blocking behaviour and try to keep the signal. Move it to alert-only or nightly, narrow it to the two or three benchmarks that guard a path with a stated requirement, and switch the blocking check to the deterministic allocation metrics that do not flake. If we want a real timing gate back, the ask is a dedicated machine, priced honestly against what a regression costs us.
  • How would you set the threshold for a regression gate?
    From the measured noise floor of the machine it runs on, not from a round number. Compare a commit against itself repeatedly, take the largest delta that appears with no code change, and set the threshold comfortably above it. If that leaves the threshold so wide it catches nothing worth catching, the honest conclusion is that this machine cannot host a timing gate at all.
  • A change is 15% faster on the benchmark but production latency does not move. What now?
    Then the benchmark is not measuring what the service is spending its time on, and that is the finding worth acting on. I would still take the change if it costs nothing, but I would treat the mismatch as a signal that the suite's coverage is pointed at the wrong code, and reprioritise measuring the real path before anyone optimises further against a benchmark that does not predict anything.

saying these in an interview costs you the question

  • Wants a strict timing gate on shared CI runners
  • Accepts two pasted numbers as a performance claim
  • Sets a regression threshold as a round number
  • Keeps a flaky gate everyone has learned to retry
  • Applies the full evidence bar to every pull request