skip to content

A Go type parameter is claimed to be faster than an interface parameter — how would you check?

level: seniorimportance: nice to knowfreq 31%

answer

  1. assume nothing about which is faster
  2. two measurements, not one
  3. -benchmem prints allocs/op
  4. -gcflags=-m shows escape and inlining
  5. one body per GC shape, plus a dictionary

basics

~20 s

Measure both. Benchmark the two signatures with go test -bench and -benchmem, then compare the compiler's decisions from go build -gcflags=-m. Go compiles generic code per GC shape with a dictionary, so a type parameter does not automatically remove indirection.

solid answer

~40 s

I would write both versions and take two measurements. First `go test -run=NONE -bench=Stage -benchmem`, reading `allocs/op` and `B/op` as well as `ns/op` — allocations are where a real difference usually shows. Second `go build -gcflags=-m` on each version, to see which values still escape to the heap and which calls stopped being inlined; a type parameter can push a function past the inlining budget and cost more than it saves. The reason not to assume is how Go implements generics: instantiation is by *GC shape*, so every pointer-shaped type argument shares one compiled body and type-specific data travels in a dictionary, which keeps constraint method calls indirect. The genuine win is narrower — a non-pointer value that would have been boxed into an interface keeps its own type instead.

code

go · 17 lines
go
func writeAll(w io.Writer, batch [][]byte) error {
	for _, b := range batch {
		if _, err := w.Write(b); err != nil {
			return err
		}
	}
	return nil
}

func writeAllG[W io.Writer](w W, batch [][]byte) error {
	for _, b := range batch {
		if _, err := w.Write(b); err != nil {
			return err
		}
	}
	return nil
}

go deeper

for a junior

Know that the honest first answer is to measure, and that go test -bench with -benchmem is how you do it in Go. Being able to name allocs/op as the number to watch is enough here.

for a middle

Explain what a benchmark and go build -gcflags=-m each tell you, and where boxing a non-pointer value into an interface costs an allocation that a type parameter avoids.

for a senior

Show that you know why the intuition fails: GC-shape stenciling with dictionaries means pointer-shaped instantiations share a body and constraint method calls stay indirect. Then tie the decision back to a profile of the real workload.

for a principal

Weigh the measured win against what the signature costs everyone else — mockability, wrapping, heterogeneous storage — and set the bar for when a performance result may change a published API at all.

## Why the claim needs checking at all "Generics are faster than interfaces" comes from languages that monomorphise: one machine-code copy per concrete type, calls bound statically, everything inlinable. Go does not work that way, so the intuition imported from elsewhere is unreliable here. ## How Go actually compiles a generic function Go uses **GC-shape stenciling with dictionaries**. A GC shape groups types the garbage collector treats identically — crucially, *all pointer-shaped types share one shape*. The compiler emits one instantiation per shape, not per named type, and passes each instantiation a **dictionary**: a hidden argument carrying the type-specific information the shared body needs, including how to call methods required by the constraint. Two consequences follow: - Instantiating a generic function with `*Row`, `*Order` and `*Batch` produces **one** body, shared. Nothing was specialised for any of them. - A method call made through a constraint goes **through the dictionary**, which is an indirect call — structurally similar to an interface method call, and equally opaque to inlining. So replacing an `io.Writer` parameter with `[W io.Writer]` does not, on its own, turn a dynamic dispatch into a static one for pointer arguments. ## Where a type parameter genuinely wins The reliable win is **avoiding the conversion to an interface value**. Passing an `int`, a `time.Time` or a small struct to a parameter of interface type converts it, and that conversion may put the value on the heap. With a type parameter, the value is passed as its own type and can stay in the frame, so the allocation disappears. In a loop over a large batch, that difference is measurable and sometimes large. A secondary win: for non-pointer type arguments the instantiation is genuinely specialised to that shape, so arithmetic and field access compile as they would in a hand-written concrete version. ## The two measurements ### The benchmark ``` go test -run=NONE -bench=Stage -benchmem ./pipeline ``` `-run=NONE` keeps ordinary tests out of the run; `-benchmem` adds `B/op` and `allocs/op`. Read allocations first — a change from 1 alloc/op to 0 in a hot loop is a real result, while a 3% shift in `ns/op` on one run is noise. Run each version several times before believing a difference, and benchmark the *shape of data the service actually handles*, not a two-element slice. ### The compiler's own report ``` go build -gcflags=-m ./pipeline ``` This prints escape-analysis and inlining decisions to standard error: which values escape to the heap, which functions are inlinable and which are "too complex". Diff the two versions. Two things to look for: an escape that vanished (the win you were hoping for), and a call that stopped being inlined (a cost you were not looking for). `-gcflags=-m` is not a performance measurement — it explains a measurement, and the benchmark is still the arbiter. ## Writing the benchmark correctly Since Go 1.24 the loop is written `for b.Loop() { ... }`, which keeps the arguments and results of the benchmarked call alive so the compiler cannot delete the work you are trying to time. Older code uses `for i := 0; i < b.N; i++` and has to defeat dead-code elimination by assigning to a package-level sink. Benchmark the same work behind both signatures, with the same input, in the same package. ## Deciding with the numbers in hand Even a real win is not automatically worth an API change: - **Is the call hot?** If it does not appear in a CPU profile of the real workload, an allocation saved per call is an allocation nobody was waiting on. - **What does the signature cost readers and callers?** An interface parameter can be mocked, wrapped, stored in a slice of mixed implementations and satisfied by a type the caller has not written yet. A type parameter can do none of those things. - **Is the improvement reachable another way?** Reusing a buffer, sizing a slice with `make([]T, 0, n)`, or moving an allocation out of the loop often beats the signature change and touches nothing public. ## The answer to give "I do not know, and neither does anyone until we measure — and I know why the intuition is unreliable here: Go instantiates by GC shape with a dictionary, so pointer-shaped instantiations share a body and constraint method calls stay indirect. The place I would expect a win is boxing of non-pointer values, and `-benchmem` will show it as allocations per operation." That answer is worth more than either confident position.

  • Where does a type parameter genuinely avoid an allocation?
    When the value is a non-pointer type — an `int`, a `time.Time`, a small struct — that would otherwise be converted into an interface value. That conversion may put the value on the heap. With a type parameter the value keeps its own type and can stay in the frame if nothing else makes it escape.
  • Why does a type parameter not automatically devirtualise a method call?
    Because Go instantiates by GC shape: all pointer-shaped type arguments share one compiled body, and the type-specific information arrives in a dictionary passed to that instantiation. A method call required by the constraint goes through the dictionary, so it is indirect in much the same way an interface call is.
  • What would make you keep the interface version even when the generic one benchmarks faster?
    A win that does not appear in a CPU profile of the real workload. If the call is not hot, the interface version is the one callers can mock, wrap and store in a heterogeneous slice, and it costs every reader less. I want the improvement visible in production data before it changes a public signature.

saying these in an interview costs you the question

  • States that generics are always faster than interfaces in Go
  • Assumes Go fully monomorphises every instantiation
  • Reports only ns/op and never runs -benchmem
  • Treats -gcflags=-m output as a performance measurement
  • Changes a public signature on the strength of one microbenchmark