A CPU profile blames a generic helper's per-element `func(T) U` callback in a batch job. Why is it slower than the loop it replaced, and how do you fix it without changing behaviour?
answer
- the type parameters are not the cost
- one shape, one instantiation, plus a dictionary
- something is called once per element
- the old loop inlined its body
- gcflags=-m tells you what inlined
basics
~20 sThe generics are rarely the cost. The callback is an indirect call the compiler usually cannot inline, once per element, where the old loop inlined its body. Benchmark it, read go build -gcflags=-m, then specialise the hot path.
solid answer
~50 sGo compiles generic code by GC-shape stenciling with dictionaries: one instantiation per shape, with a hidden dictionary argument supplying per-instantiation type information. There is no boxing and no reflection, so the type parameters themselves cost close to nothing. What costs is the `func(T) U` value: it is an indirect call executed once per element, and the compiler usually cannot inline through it, whereas the hand-written loop had its arithmetic inlined into the loop body. On a hundred-million-element pass that difference is visible in a CPU profile. Fix it by measuring first — a benchmark with `-benchmem`, then `go build -gcflags=-m` to see what did and did not inline — and then specialising: write the concrete loop for the one hot type and keep the generic helper for the cold callers. Preallocating the output and hoisting work out of the callback usually matter more than the generics.
code
go · 16 linesfunc Reduce[T, A any](s []T, init A, f func(A, T) A) A {
acc := init
for _, v := range s {
acc = f(acc, v) // indirect call, usually not inlined
}
return acc
}
// specialised for the one type the profile blamed
func total(ds []time.Duration) time.Duration {
var sum time.Duration
for _, d := range ds {
sum += d
}
return sum
}go deeper
Know that passing a function value means a real call for every element, while a loop body written inline can be optimised into the loop. That difference only matters over very large inputs.
Explain GC-shape stenciling with dictionaries: types sharing a shape share one instantiation, nothing is boxed, and nothing is resolved by reflection. Contrast it with monomorphisation and with type erasure.
Demonstrate the loop of benchmark, gcflags=-m, targeted change, re-measure. Show that you specialise only the path the profile blamed and can justify the duplicated code by the number it moved.
Decide how much duplication a measured win is worth and who maintains it. A specialised copy is a permanent fork of a helper's behaviour, so it needs a benchmark in the tree and an owner, or it silently diverges from the generic version it was cloned from.
## What the compiler actually generates Go does not implement generics by full monomorphisation and does not erase them either. It uses **GC-shape stenciling with dictionaries**. A GC shape is roughly "how the garbage collector sees this type": all pointer-shaped type arguments — `*Foo`, `*Bar`, `*bytes.Buffer` — share one shape and therefore share a single compiled instantiation, while distinct non-pointer shapes such as `int64` and `string` get their own. Alongside the ordinary arguments, an instantiated call passes a hidden **dictionary**: a static table carrying the type descriptors, method entries and other per-instantiation data the shared code needs. Three consequences follow: - No value is boxed into an interface, and nothing is decided by reflection. A generic `Map` over `[]int64` really does move `int64`s. - Code size is bounded: a helper instantiated with twenty pointer types produces one body, not twenty. - Calling a *method* on a value of a type parameter goes through the dictionary, which is an indirect call rather than a direct one. If your helper's hot inner operation is a constraint method, that is a real per-element cost. ## Where the time in this profile is actually going For a transform helper the dominant cost is usually not the dictionary at all — it is the callback. `f` is a function *value*. Calling it is an indirect call through a pointer, and the compiler generally cannot inline through an indirect call, so every element pays call overhead and loses the optimisations that inlining unlocks: constant folding, keeping an accumulator in a register, eliminating a bounds check. The loop the helper replaced had its two lines of arithmetic sitting directly in the loop body, where all of that applied. Two further costs often ride along: - **Allocation.** The helper returns a new slice. The loop it replaced may have accumulated into a single variable or written into a buffer the caller already owned. On a batch pass over a latency histogram, one allocation per call and the resulting GC pressure can outweigh the call overhead. - **Escapes.** A closure that captures variables may put them on the heap. `go build -gcflags=-m` prints both the escape decisions and the inlining decisions, and it is the fastest way to see which of these is happening. ## The order of work The chair here is someone making a batch job faster without changing what it computes, so the discipline matters more than the trick. 1. **Reproduce it in a benchmark.** A `testing.B` benchmark over realistic input with `-benchmem` turns "the profile blames this" into a number you can move. Go 1.24 added `for b.Loop()`, which keeps the timed body from being optimised away. 2. **Read the compiler's decisions.** `go build -gcflags=-m` (twice, `-m -m`, for more detail) says what inlined and what escaped. If the callback did inline, generics are not your problem and the profile is pointing at the work inside it. 3. **Attack the biggest term first.** Preallocating the output slice, hoisting a repeated computation out of the callback, or avoiding a per-element allocation usually beat anything structural. 4. **Specialise only the hot path.** Write the concrete loop for the one type that dominates the profile, keep the generic helper for the dozen cold call sites, and put a comment on the specialised copy explaining which benchmark justifies its existence. Duplicating a helper for every type "because generics are slow" is the failure mode. 5. **Re-measure.** The same benchmark, the same input. A change that does not move the number gets reverted, not kept for its plausibility. Profile-guided optimisation is worth knowing about here: building with a representative CPU profile lets the compiler inline more aggressively along hot paths and devirtualise some indirect calls, and it costs no source change at all. It became generally available in Go 1.21. ## What not to conclude The wrong lesson is "generics are slow, go back to `interface{}`". An `any`-based helper is worse on both counts: it boxes values and adds a type assertion per element. The right lesson is that a *callback per element* is the expensive part of a functional-style helper in any implementation, and that in Go, where a plain `for` loop is idiomatic and inlines, the cheapest fix at the hot spot is often to write the loop out again — locally, deliberately, with a benchmark next to it — and leave the generic helper serving everywhere the cost does not matter.
- Would rewriting the helper to take `any` instead of a type parameter be faster?No, it is worse on both counts. Passing values as `any` boxes non-pointer values, which allocates, and every element then needs a type assertion. Generics avoid both: the instantiated code moves concrete values with no boxing. The indirect callback cost stays exactly the same, since it was never about the type parameters.
- How many machine-code copies does the compiler emit for a helper instantiated with `*Foo`, `*Bar` and `int64`?Two. `*Foo` and `*Bar` share a GC shape — both are pointer-shaped — so they share one instantiation, distinguished at the call by different dictionaries. `int64` has its own shape and gets its own body. That is why generic code does not blow up code size the way full monomorphisation would.
- What do you check before concluding that the callback is the problem?Run a benchmark with `-benchmem` so there is a number to move, then `go build -gcflags=-m` to see the inlining and escape decisions. If the callback did inline, the profile is blaming the work inside it, not the call. Allocation of the result slice is often the larger term anyway.
- Is profile-guided optimisation worth trying here?Often yes, because it needs no source change. Building with a representative CPU profile lets the compiler inline more aggressively on hot paths and devirtualise some indirect calls. It has been generally available since Go 1.21. Treat it as a multiplier on a well-shaped program, not as a substitute for removing the per-element call.
The generic helper is a courier who fetches instructions from an envelope before each delivery; the inlined loop already knows the route. The envelope is cheap, but opening it a hundred million times is not.
saying these in an interview costs you the question
- Claims Go generics box values like an any-based helper
- Believes each type argument gets its own compiled copy
- Says generic code runs through reflection at run time
- Rewrites everything concrete without a benchmark
- Thinks the dictionary is consulted once per element
- Concludes generics are slow and reverts to interface{}