In go test -benchmem output, what do B/op and allocs/op measure, and what does b.ReportAllocs do?
answer
- two extra columns on the benchmark line
- bytes requested, not bytes still live
- the stack never shows up here
- counts of events survive a noisy machine
- one benchmark can opt in without the flag
basics
~20 sB/op is the average bytes of heap memory allocated per iteration and allocs/op the average number of heap allocations. Pass -benchmem to go test for every benchmark, or call b.ReportAllocs inside one benchmark to always report those two columns.
solid answer
~50 s`go test -bench=. -benchmem` adds two columns to every benchmark line: `B/op`, the average number of bytes allocated **on the heap** per iteration, and `allocs/op`, the average number of distinct heap allocations per iteration. Both are totals accumulated over the timed region and divided by `b.N`, taken from the runtime's allocation counters — so they count bytes *allocated*, not bytes still live at the end, and they exclude anything the compiler kept on the stack via escape analysis. `b.ReportAllocs()` does the same thing for a single benchmark without the flag, which is handy when allocation count is the point of that particular benchmark. In practice `allocs/op` is the most valuable column: it is far more stable across machines and runs than `ns/op`, so a jump from 4 to 9 allocations is a credible regression signal even on a noisy laptop.
code
text · 5 lines$ go test -bench=EncodeChunk -run=^$ -benchmem
goos: linux
goarch: amd64
BenchmarkEncodeChunk-8 2874316 412.6 ns/op 272 B/op 4 allocs/op
PASSgo deeper
Know the flag and the two columns by name: go test -bench=. -benchmem prints B/op and allocs/op, which are bytes and allocation counts averaged per iteration.
Explain that both come from the runtime's heap counters over the timed region divided by b.N, that they measure bytes allocated rather than bytes retained, and that b.ReportAllocs opts one benchmark in without the flag.
Argue for allocation counts as the stable signal on noisy hardware, read the shape of a cost from four allocations totalling 272 bytes, and recognise the truncation case where B/op is non-zero while allocs/op prints 0.
Decide what a team actually gates on: which benchmarks are load-bearing, whether an allocation-count budget belongs in review, and how you keep such a budget from turning into micro-optimisation that nobody can tie to production behaviour.
## Turning the columns on Two equivalent switches: ```text $ go test -bench=EncodeChunk -benchmem BenchmarkEncodeChunk-8 2874316 412.6 ns/op 272 B/op 4 allocs/op ``` ```go func BenchmarkEncodeChunk(b *testing.B) { b.ReportAllocs() // ... } ``` `-benchmem` is a flag on the `go test` command line and applies to every benchmark in the run. `b.ReportAllocs()` is a method on `*testing.B` that opts one benchmark in permanently, so the columns appear even when a colleague forgets the flag. Using both is harmless. ## What the numbers actually are The testing package snapshots the runtime's memory statistics at the start and end of the timed region and reports the deltas divided by `b.N`: * **`B/op`** — total bytes of heap memory *requested* during the run, divided by `b.N`. It is cumulative allocation, not peak usage and not resident memory. A benchmark that allocates a 1 KB buffer and immediately drops it every iteration reports about 1024 B/op even though the program never holds more than a kilobyte. * **`allocs/op`** — the number of distinct heap allocation events, divided by `b.N`. One `make([]byte, 1<<20)` is a single allocation of a megabyte; a thousand small `append` growth steps are a thousand allocations. Both are integer-divided, so a cost paid on only some iterations rounds down. A benchmark that allocates once every other iteration prints `0 allocs/op` while `B/op` still shows roughly half the object's size — a genuinely confusing output the first time you meet it. ## What they do not include * **Stack allocations.** If the compiler proves a value does not outlive the frame, it lives on the stack and never appears in these columns. A change that drops `allocs/op` from 3 to 0 usually means values stopped escaping to the heap, not that the work disappeared. * **Work excluded from the timed region.** Anything between `b.StopTimer()` and `b.StartTimer()`, or before `b.ResetTimer()`, is left out of the allocation counters just as it is left out of the clock. That is deliberate and it is why fixture construction does not pollute the figures. * **Memory the operating system gave the process.** These are Go heap counters, not RSS. ## Why allocs/op is the column to watch `ns/op` is a physical measurement and inherits every source of noise on the machine: other processes, CPU frequency scaling, thermal throttling, cache state, memory layout. Two runs of an unchanged benchmark can differ by several percent. `allocs/op` is a *count of events in your program's execution*. Barring randomised inputs, it is usually identical run to run and identical across machines. That makes it the sharpest tool in the set: * A pull request that leaves `ns/op` inside the noise but takes `allocs/op` from 4 to 11 has almost certainly introduced a regression that will show up under production load, where allocation drives collector work. * Conversely, an optimisation that halves `allocs/op` is real even if the micro-benchmark's `ns/op` barely moves, because the saved collector pressure is paid elsewhere in the process. This is why teams that gate on benchmark output usually gate on the allocation columns first: they are the part of the output you can believe from a single run on ordinary hardware. ## Reading a line end to end ```text BenchmarkEncodeChunk-8 2874316 412.6 ns/op 272 B/op 4 allocs/op ``` * `EncodeChunk` — the benchmark's name minus the `Benchmark` prefix. * `-8` — the `GOMAXPROCS` value the run used. * `2874316` — the final `b.N`, the iteration count the reported numbers are divided by. * `412.6 ns/op` — elapsed time per iteration. * `272 B/op` — heap bytes allocated per iteration. * `4 allocs/op` — heap allocation events per iteration. Four allocations totalling 272 bytes tells you the shape of the cost too: several small objects rather than one large buffer, which points at things like a per-record header slice, a string conversion, or an interface boxing a value — each a candidate for removal. ## A related column you can add yourself When a benchmark processes a payload, `b.SetBytes(n)` declares how many bytes one iteration handles and makes the harness print a throughput column in MB/s alongside `ns/op`. It is a reporting aid only: it changes no measurement, and `n` must be the bytes handled per single iteration, not per run.
- Why is allocs/op often a more trustworthy signal than ns/op?`allocs/op` counts events inside your program, so it is essentially deterministic and stable across machines and runs. `ns/op` is a physical measurement that absorbs CPU frequency scaling, cache state and neighbouring processes. A jump in allocations is believable from one run; a small change in time usually is not.
- A benchmark reports 512 B/op and 0 allocs/op. How is that possible?Both columns are integer divisions of a total by `b.N`. If the code allocates a 1 KB object on roughly every second iteration, the byte total halves to about 512 per operation while the allocation count of about 0.5 truncates to 0. It signals an allocation that is amortised, not absent.
- Do these columns show stack allocations?No. They come from the runtime's heap counters, so anything the compiler proved could stay in the function's frame is invisible here. That is why an optimisation which stops a value escaping shows up as `allocs/op` dropping to 0 rather than as a smaller number: the work moved off the heap entirely.
saying these in an interview costs you the question
- Reads B/op as peak memory or resident set size
- Believes stack allocations appear in allocs/op
- Thinks allocs/op counts garbage collection cycles
- Assumes 0 allocs/op proves nothing was ever allocated
- Treats a 2% ns/op change as more meaningful than an allocation jump