skip to content

How do you prove a Go benchmark's fast ns/op comes from the compiler deleting the loop body rather than from real speed?

level: seniorimportance: should knowfreq 30%

answer

  1. cheapest evidence first, then decisive
  2. does the number scale with input size?
  3. convert it into bytes per second
  4. -m explains inlining, not deletion
  5. disassemble the compiled test binary

basics

~10 s

Scale the input and see whether the timing moves, check the implied throughput against physical limits, and disassemble the built test binary. Then re-run with the result anchored: an orders-of-magnitude jump proves deletion.

solid answer

~50 s

I work from cheap evidence to decisive evidence. First arithmetic: convert ns/op into bytes per second and ask whether a machine can do that — 1 MiB in 0.3 ns is petabytes per second, which settles it. Then scale the input: double the buffer and re-run; genuine work roughly doubles, a deleted body does not move. Then `go test -gcflags=-m` tells me whether the function was inlined, which is the precondition for elision, though `-m` never says "I removed your loop". The decisive step is `go test -c` followed by `go tool objdump -s BenchmarkHashBlock` — if the loop holds only a counter increment and a compare, the argument is over. Finally, add a sink or switch to `for b.Loop()` and compare. Inlining that real callers also get is honest speed; deleting work nobody asked for is not.

code

text · 5 lines
text
go test -c -o pkg.test ./checksum
go tool objdump -s 'BenchmarkHashBlock' pkg.test

# and, separately, what the compiler decided about the call:
go test -gcflags=-m -run='^$' -bench=HashBlock ./checksum

go deeper

for a junior

Learn the two cheap checks you can run yourself: work out the implied bytes per second, and re-run with a bigger input to see whether the timing moves with it.

for a middle

Be able to explain what inlining has to do with elision, and to use the compiler's inlining output correctly — as the mechanism that made deletion possible, not as evidence that it happened.

for a senior

Show a ladder of evidence ending in something decisive: disassembling the built test binary, or a controlled before-and-after with the result anchored. Separate elision from inlining and from work that genuinely is small.

for a principal

Set the bar for what a performance claim must carry before it can steer a decision, so credibility comes from how benchmarks are written rather than from whoever reviews them most carefully.

## The situation A pull request adds a benchmark for a checksum routine on a hot path, and the number in the description is 0.31 ns/op for a 1 MiB block. The author is pleased. Your job as the reviewer is to decide whether that number can be quoted in a design discussion, and "it looks too fast" is not a review comment. Here is the ladder of evidence, cheapest first. ## 1. Arithmetic against physics Divide the input size by the time. 1,048,576 bytes in 0.31 ns is about 3.4 PB/s. Server memory bandwidth is in the tens of GB/s and L1 bandwidth in the hundreds. Five orders of magnitude of gap is not a fast implementation. A useful habit that makes this automatic is calling `b.SetBytes(int64(len(buf)))` in the benchmark, so the harness reports throughput next to the timing and an impossible number is visible without doing the division by hand. This step costs nothing and settles most cases. It is also the argument that convinces a defensive author, because it does not depend on any claim about the compiler. ## 2. Scale the input Run the same benchmark with a 2 MiB and an 8 MiB buffer. Real work scales roughly with the input; a loop whose body was deleted reports the same sub-nanosecond figure at every size, because what it measures — the counter increment and the bounds compare — never depended on the buffer at all. A flat line across a 8x change in input is strong evidence on its own. ## 3. Ask the compiler what it did to the call `go test -gcflags=-m -bench=HashBlock -run='^$' ./...` prints the compiler's inlining and escape-analysis decisions for the package including its test files. What you are looking for is `can inline hashBlock` and `inlining call to hashBlock` at the benchmark's line. That is the precondition: the compiler had to see through the call before it could decide the work was dead. Be honest about the limit of this tool. `-m` reports inlining and escape decisions; it does not print "I eliminated this loop body". It gives you the mechanism, not the conclusion. Presenting `-m` output as proof of elision overstates it, and a careful author will say so. ## 4. Look at the machine code This is the step that ends the argument. Build the test binary rather than running it: ``` go test -c -o pkg.test ./checksum go tool objdump -s 'BenchmarkHashBlock' pkg.test ``` Read the loop. If the body between the backward branch and its target is an increment, a compare and a jump, nothing was hashed. If you see the load-and-mix sequence you expect from the hash, the work is there and the number came from somewhere else — a wrongly sized input, an inlined constant, or a genuinely fast routine. `go build -gcflags=-S` gives the same information as compiler output if you prefer reading the assembler listing to disassembling. ## 5. The controlled experiment Change one thing: anchor the result in a package-level sink, or rewrite the loop as `for b.Loop()`. Re-run. If ns/op moves from 0.31 to hundreds of microseconds, you have a before-and-after that needs no interpretation. This is the fix and the proof in the same edit, which is why it belongs in the PR rather than in a comment thread. ## Distinguishing elision from honest speed The reviewer's real skill is not assuming every fast number is a lie. Three legitimate causes look similar at first glance. **Inlining.** If the function under test is inlined into the benchmark, call overhead vanishes from the measurement. That is not cheating when production callers also inline it — it is what they will experience. It becomes misleading only when the benchmark is used to argue about a call that in practice is not inlined, for example because it is reached through an interface. If you want the call included, `//go:noinline` on the function makes the choice explicit and should be commented. **Constant folding.** A compile-time-constant input lets the compiler compute the answer at build time; the sink then holds a constant and the loop is empty again. The cure is to keep the input in a package-level variable or build it at run time. **The work really is small.** Sometimes the routine is genuinely a few nanoseconds and the reviewer's instinct is simply wrong. That is what step 2 is for: if the number scales with the input, it is measuring the input. ## What to write in the review Ask for two things. First, the anchored benchmark — sink or `b.Loop` — so the number is credible by construction rather than by argument. Second, `b.SetBytes` where the benchmark has a natural unit of input, because a throughput figure makes the next impossible number obvious to whoever reads it after you.

  • Why is go build -gcflags=-m evidence but not proof of elision?
    `-m` reports inlining and escape-analysis decisions. Inlining is the precondition for the body being removed — the compiler must see through the call before it can judge the work dead — but `-m` never states that a loop body was deleted. It tells you the mechanism was available, not that it fired. Disassembling the loop, or the before-and-after with a sink in place, is what actually settles it.
  • The author argues the benchmark is fast because the function was inlined. When is that a fair defence?
    When real callers get the same inlining. Measuring an inlined form is honest if that is how the function is reached in production. It stops being fair when the production path defeats inlining — a call through an interface, a function value, or a package boundary the compiler will not cross — because the benchmark then omits overhead every real caller pays. Say which path the number is meant to represent.
  • What would you ask the author to add so the next reader does not have to repeat this analysis?
    An anchored loop — a package-level sink or `for b.Loop()` — so the result cannot be deleted, and `b.SetBytes` with the input length so the harness reports throughput alongside the timing. A throughput figure turns "is this plausible?" into a one-glance check, which is the only version of this review that scales.
  • The number is flat as you scale the input, but the disassembly clearly contains the hash. What now?
    Look at what the loop is actually fed. A common cause is that the benchmark hashes a fixed small prefix, or a compile-time constant, while the buffer you scaled is never reached. Another is that setup, not the hash, dominates and is being timed. Print or assert the input length inside the benchmark once, and re-check where the buffer comes from.

saying these in an interview costs you the question

  • Accepts a number because the code looks correct
  • Treats -gcflags=-m output as proof the body was deleted
  • Never converts the timing into implied throughput
  • Assumes every fast benchmark result is elision
  • Calls inlining cheating even when real callers inline too
  • Reruns with more iterations and calls the number confirmed