A Go benchmark that hashes a 1 MiB buffer reports 0.3 ns/op — what most likely happened?
answer
- the number is physically impossible
- who ever reads that hash result?
- no side effects and no observer
- dead code once the call is inlined
- store it into a package-level variable
basics
~20 sThe compiler deleted the work. A function with no side effects whose result is never used becomes dead code once it is inlined, so the loop measures only its own counter. Keep the result alive by storing it in a package-level variable.
solid answer
~40 s0.3 ns/op is about one iteration of an empty loop on a modern core, not a megabyte of hashing — that figure implies petabytes per second, which no machine does. The benchmark almost certainly called the hash function and threw the result away. The compiler inlines the call, sees that nothing ever observes the returned value and that the function has no side effects, and removes it; what remains is a counting loop. The fix is to make the result observable: accumulate it into a local and store that local into a package-level sink variable, or write the loop as `for b.Loop()`, which keeps call arguments and results alive. `b.ResetTimer` does not help — the timer was honest, there was simply no work inside it.
code
go · 7 linesvar buf = make([]byte, 1<<20)
func BenchmarkHashBlock(b *testing.B) {
for i := 0; i < b.N; i++ {
hashBlock(buf) // result discarded
}
}go deeper
Be ready to say out loud that a sub-nanosecond result for a megabyte of input is impossible, and that an unused result lets the compiler delete the work. Know the one-line fix: store the result in a package-level variable.
Explain the two preconditions for elision — the call gets inlined, and the function has no side effects — and why assigning to the blank identifier is not an observation. Show the sink pattern and say what it costs per iteration.
Demonstrate the reflex of sanity-checking a benchmark number against physical throughput before trusting it, and show how you would prove elision by scaling the input and re-running with a sink in place.
Own the habit at team scale: benchmarks that silently measure nothing are worse than no benchmarks, because they get quoted in design decisions. Decide what convention the codebase uses so a fast number is credible by construction.
## The symptom You write your first benchmark for a checksum package: ```go var buf = make([]byte, 1<<20) func BenchmarkHashBlock(b *testing.B) { for i := 0; i < b.N; i++ { hashBlock(buf) } } ``` and `go test -bench=HashBlock` prints something like `1000000000 0.31 ns/op`. Two things in that line should stop you: the iteration count has hit the ceiling the benchmark runner will go to, and the per-operation time is under a nanosecond for a megabyte of input. ## Why the number cannot be real Do the arithmetic before you do anything else. One MiB is 1,048,576 bytes. At 0.31 ns per operation that is roughly 3.4 petabytes per second of throughput. Main memory on a large server delivers tens of gigabytes per second; even reading the buffer once out of L1 cache is orders of magnitude away from that. A number that is physically impossible is not a fast implementation, it is a measurement of nothing. 0.31 ns is, however, exactly what one iteration of an empty loop costs on a ~3 GHz core: one increment and one comparison, about a cycle. That is the tell — the loop ran, the body did not. ## What the compiler did Go's compiler inlines small functions. Once `hashBlock(buf)` is inlined into the benchmark, the compiler is looking at plain arithmetic over a slice whose computed value is assigned to nothing. It then applies ordinary dead-code elimination: a value that no reachable code ever reads, produced by operations with no side effects (no stores to memory that outlives the function, no channel operations, no calls it cannot see through), can be deleted. Delete the value and the operations feeding it become dead too, so the entire body disappears and the loop is left with its counter. Nothing here is a bug or an aggressive flag. It is the same optimisation that makes real Go code fast; the benchmark simply asked the compiler to compute something nobody wanted. Two preconditions matter. The call has to be inlined or otherwise transparent — if the compiler cannot see the body it must assume the call has effects and must emit it. And the function has to be effect-free — a hash that writes into a package-level table, prints, or takes a lock cannot be removed. ## Making the result observable The idiom is a package-level sink: ```go var Sink uint64 func BenchmarkHashBlock(b *testing.B) { var h uint64 for i := 0; i < b.N; i++ { h ^= hashBlock(buf) } Sink = h } ``` The store into `Sink` is observable outside the function, so the compiler must perform it; that keeps `h` live, which keeps every `hashBlock` call live. The `^=` costs one cheap ALU operation per iteration, which is noise next to hashing a megabyte. Since Go 1.24 the shorter route is `for b.Loop() { hashBlock(buf) }`. `testing.B.Loop` was designed for this: arguments and results of calls inside the loop are kept alive, so the call cannot be deleted outright, and the loop manages the timer for you. ## What does not fix it `_ = hashBlock(buf)` does not fix it — assignment to the blank identifier is defined to discard, and the compiler treats it as no observation at all. `b.ResetTimer()` before the loop does not fix it either; it resets the clock, and the clock was never the problem. Raising `-benchtime` or `-count` just repeats the empty measurement more times. ## Confirming it in thirty seconds Double the buffer to 2 MiB and re-run. Real hashing roughly doubles ns/op; an elided loop does not move, because its cost never depended on the input. Then add the sink and watch the number jump by four or five orders of magnitude. That before-and-after is the evidence you bring to a review.
- Does calling b.ResetTimer before the loop rescue a benchmark whose body was optimised away?No. `b.ResetTimer` only zeroes the elapsed time and allocation counters so setup done before it is not charged to the measurement. It has no influence on what the compiler emits. If the loop body was deleted, resetting the timer just measures the empty loop more carefully. The fix has to make the computed value observable — a package-level sink, or a `for b.Loop()` loop.
- How would you confirm the elision in under a minute, without reading assembly?Change the input size and re-run. Hash 2 MiB instead of 1 MiB: genuine work roughly doubles ns/op, while a deleted body reports the same sub-nanosecond figure because its cost never depended on the buffer. As a second check, add the sink and compare — a jump from 0.3 ns/op to hundreds of microseconds tells you exactly what the earlier number was measuring.
- Why does the elided benchmark still report about 0.3 ns rather than 0?The loop itself survives. Each iteration still increments `i` and compares it against `b.N`, which is roughly one cycle on a modern core — about 0.3 ns at 3 GHz. The benchmark harness also divides total elapsed time by the iteration count, so you are seeing the real cost of an empty counting loop, not a rounding artefact.
saying these in an interview costs you the question
- Says the machine is simply very fast
- Blames CPU cache warming for the tiny number
- Thinks b.ResetTimer makes any number trustworthy
- Assumes the compiler never deletes function calls
- Adds an assignment to the blank identifier and expects it to help
- Never sanity-checks implied throughput against the input size