skip to content

A Go benchmark reports different ns/op at -benchtime=1s and -benchtime=10s. What causes that?

level: seniorimportance: nice to knowfreq 32%

answer

  1. the flag is a knob on b.N
  2. the division should cancel out
  3. amortised setup gets cheaper the longer you run
  4. leftovers make later iterations worse
  5. repeat the measurement before blaming the code

basics

~20 s

A longer benchmark duration means a larger b.N, and ns/op is the whole timed run divided by it. A figure that moves means iterations are not independent: one-off setup amortised away, or state accumulating across iterations.

solid answer

~40 s

`-benchtime` sets how long a single measurement must last, so raising it raises `b.N`, and `ns/op` is the whole timed region divided by `b.N`. For a single repeatable operation that division cancels out and the figure stays flat. When it does not, the iterations are not independent. The two usual causes pull in opposite directions: a fixed cost inside the timed region — fixture construction above the loop with no `b.ResetTimer` — is divided by a bigger `b.N`, so `ns/op` **falls** with a longer run; and state that grows per iteration, such as appending to a package-level slice, makes later iterations genuinely more expensive through collector work and cache misses, so `ns/op` **rises**. Machine noise is the third possibility: use `-count` to repeat the measurement and see whether the spread already explains the gap.

code

text · 4 lines
text
$ go test -bench=EncodeChunk -run=^$ -benchtime=1s
BenchmarkEncodeChunk-8   	 2421000	       412.6 ns/op
$ go test -bench=EncodeChunk -run=^$ -benchtime=10s
BenchmarkEncodeChunk-8   	15840000	       631.2 ns/op

go deeper

for a junior

Know that -benchtime sets how long one measurement runs, default one second, and that raising it raises the iteration count the harness picks.

for a middle

Explain why the iteration count should cancel out of ns/op, and name the two shapes that break it: a fixed cost inside the timed region, and state carried from one iteration to the next.

for a senior

Walk the diagnosis end to end: sweep the duration, repeat measurements to see the spread, watch whether allocations per operation move too, then find what an iteration leaves behind. Say plainly that a duration-dependent benchmark cannot support a decision.

for a principal

Own the standard of evidence: what a team is allowed to conclude from a micro-benchmark, what hardware and repetition discipline the numbers need before they enter a review, and when the honest answer is to measure the real workload instead.

## What -benchtime actually changes `-benchtime` sets the minimum duration of one measurement; the default is `1s`. The harness keeps raising `b.N` until a run reaches that duration, then reports elapsed-time-over-`b.N`. It also accepts an `x` suffix for an exact iteration count: `-benchtime=100x` pins `b.N` to 100 and skips the ramp entirely. So `-benchtime` is really a knob on `b.N`. If your benchmark measures one repeatable operation, the value of `b.N` cancels in the division and the reported cost should not care what you set. When it does care, the benchmark is telling you something about itself. ```text $ go test -bench=EncodeChunk -benchtime=1s BenchmarkEncodeChunk-8 2421000 412.6 ns/op $ go test -bench=EncodeChunk -benchtime=10s BenchmarkEncodeChunk-8 15840000 631.2 ns/op ``` ## Cause 1 — a fixed cost inside the timed region (figure falls) This is the classic. Fixture construction above the loop, or a one-time cache warm-up on the first iteration, is charged to the whole run. Divided by 2.4 million iterations it adds a lot per operation; divided by 15.8 million it adds much less. The symptom is a number that gets *better* the longer you run, and it is the reason `b.ResetTimer()` exists. A subtler version: the first iteration populates a cache, opens a connection, or triggers a lazily-initialised package variable. The benchmark then reports the amortised cost of a warm path plus a fraction of a cold one, and the fraction depends on `b.N`. ## Cause 2 — iterations that are not independent (figure rises) If each iteration leaves something behind, later iterations run in a different, worse environment than earlier ones: * Appending to a package-level slice, or inserting into a map that is never cleared, grows the live heap. A larger live heap means the collector runs more often and scans more, and that cost lands inside the timed region. * Retaining buffers or byte slices per iteration turns a garbage-free benchmark into one that allocates monotonically. * Growing data structures destroy locality: at ten thousand entries everything is in cache, at ten million nothing is. The symptom is a number that gets *worse* with a longer run, and the fix is to make each iteration leave no trace, or to reset the accumulated state under `b.StopTimer()` at a fixed interval. ## Cause 3 — the machine, not the code A laptop that boosts its clock for the first second and then throttles will report a slower `ns/op` for a ten-second run of *any* benchmark. Background processes, another build, a virtual machine's noisy neighbours and CPU frequency governors all move the number by several percent. This is what `-count` distinguishes. `go test -bench=EncodeChunk -benchtime=10s -count=8` repeats the whole measurement eight times and prints eight lines. If the eight lines at 10s span 600-660 ns/op and the 1s runs span 400-660 ns/op, the difference between durations is inside the spread and there is nothing to explain. If the eight 10s lines cluster tightly at 631 and the 1s lines cluster tightly at 412, the benchmark itself is the story. ## How to run the diagnosis 1. **Make the difference reproducible.** Several runs at each duration, `-count` for repetitions, with `-benchmem` on so you can see whether allocations per operation move too. If `allocs/op` also changes with the duration, the iterations are definitely not independent — allocation counts are deterministic in a clean benchmark. 2. **Sweep the duration.** `1s`, `3s`, `10s`. Falling monotonically points at amortised fixed cost; rising monotonically points at accumulating state; noise looks like neither. 3. **Pin the iteration count.** `-benchtime=1000x` and `-benchtime=100000x` remove the harness's ramp from the picture and let you compare like with like. 4. **Read the loop with fresh eyes.** What survives an iteration? A package-level variable, a channel nobody drains, an unclosed resource, a buffer being appended to, a lazily-built index. Anything that persists is a candidate. 5. **Fix the benchmark before you trust either number.** A benchmark whose result depends on how long you ran it cannot support a decision about which implementation is faster; both figures are measuring the benchmark rather than the code. ## The judgment to bring The hardest part is resisting the instinct to pick the number you like. Engineers routinely quote the shorter run because it is prettier, or the longer one because it looks more rigorous. Neither is defensible while the figure is duration-dependent: the honest position is that the benchmark is broken, and the first work item is making it report the same cost at 1s, 10s and a pinned iteration count.

  • Which direction does the figure move when fixture setup sits inside the timed region?
    It falls as the duration rises. The fixed cost is divided by a larger `b.N`, so each additional second of running makes the operation look cheaper. A monotonically improving `ns/op` across `1s`, `3s` and `10s` is close to a signature of an un-reset timer or a one-off warm-up on the first iteration.
  • How does -count help separate a broken benchmark from a noisy machine?
    `-count=N` repeats the whole measurement N times and prints N lines, showing the spread at a single duration. If the gap between two durations sits inside that spread, the machine explains it. If each duration clusters tightly on its own value, the benchmark's iterations are not independent and the code needs fixing.
  • Why is a changing allocs/op across durations especially damning?
    Allocation counts are deterministic for a clean benchmark: the same operation repeated should allocate the same amount every time, so the average is flat regardless of `b.N`. If it climbs with the duration, iterations are leaving allocations behind — accumulating state — and the timing difference has a concrete, findable cause.
  • When is -benchtime with an x suffix the right tool?
    `-benchtime=1000x` pins `b.N` to exactly 1000, removing the harness's ramp-up from the comparison. It is useful for reproducing a duration-dependent result deterministically, and for very slow operations where a time-based run would take minutes. For routine measurement a time-based run gives a better-averaged number.

saying these in an interview costs you the question

  • Picks whichever duration gives the nicer number
  • Assumes a longer run is automatically more accurate
  • Never suspects state shared between iterations
  • Thinks -benchtime changes what one iteration does
  • Confuses -benchtime with repeating the whole measurement