skip to content

In Go, what are open-coded defers and when does the compiler fall back to runtime defer records?

level: middleimportance: should knowfreq 40%

answer

  1. not every defer allocates anything
  2. the compiler can inline the whole thing
  3. one bit per defer statement
  4. a loop makes the count unknowable
  5. debug builds get the old path back

basics

~20 s

An open-coded defer is one the Go compiler inlines: it keeps the call in stack slots, sets a bit in a bitmask, and calls it at each return. Defers in a loop, or more than eight per function, fall back to runtime records.

solid answer

~50 s

Historically every `defer` statement asked the runtime to allocate a `_defer` record and push it onto a per-goroutine list, with `runtime.deferproc` doing the push and `runtime.deferreturn` popping and running them at the function's exit — tens of nanoseconds of work per defer. Since Go 1.14 the compiler can *open-code* defers instead: it keeps the deferred function value and its arguments in stack slots, sets one bit of a small bitmask when the `defer` statement executes, and emits inline code at every return point that tests the bits and calls the functions in reverse order. No record, no list, no runtime call, so the cost collapses to roughly a branch plus a direct call. Open coding requires that no `defer` in the function is inside a loop, that there are at most eight `defer` statements, and that optimizations are on — a debug build with `-gcflags=-N` gets the record path back. Semantics are identical either way.

code

go · 17 lines
go
func openCodable(mu *sync.Mutex, f *os.File) error {
	mu.Lock()
	defer mu.Unlock() // one of at most 8 statements, not in a loop
	defer f.Close()
	return nil
}

func notOpenCodable(paths []string) error {
	for _, p := range paths {
		f, err := os.Open(p)
		if err != nil {
			return err
		}
		defer f.Close() // in a loop: one runtime record per iteration
	}
	return nil
}

go deeper

for a junior

Know that modern Go usually compiles defer down to inline calls rather than runtime bookkeeping, so the old advice to avoid defer for speed is largely obsolete.

for a middle

Be ready to describe the mechanism: stack slots for the arguments, one bit per defer statement, and inline code at each return that tests the bits and calls in reverse order.

for a senior

Show that you know the fallback conditions and can check them: a defer in a loop, more than eight defer statements, or optimizations disabled, verified by looking for the runtime defer symbols in the assembly.

for a principal

Own the guidance your team writes down. Replace a blanket ban on defer in hot code with the narrow, still-true rule about loops, and require a measurement on an optimized build before anyone hand-unrolls cleanup.

## Three implementations, one meaning The observable behaviour of `defer` has never changed: arguments are evaluated when the `defer` statement runs, the calls happen last-in-first-out as the function finishes, and a deferred closure can still change a named result. What has changed, twice, is how the compiler and runtime deliver that. **Heap records (the original).** Each `defer` statement called into the runtime, which took a `_defer` record from a pool, filled in the function value and a copy of the arguments, and pushed it onto the current goroutine's linked list of pending deferred calls. At the function's exit the compiler had emitted a call to `runtime.deferreturn`, which walked that list, popped the records belonging to this frame, and ran them. Correct, uniform, and comfortably the most expensive thing in an otherwise trivial function. **Stack-allocated records (Go 1.13).** For defers that are not inside a loop, the compiler can reserve space for the `_defer` record in the caller's own stack frame instead of taking one from the heap. The record is still created and still linked onto the goroutine's list, so the bookkeeping remains, but the allocation disappears. This roughly halved the cost. **Open coding (Go 1.14).** The big one. If the compiler can see, statically, exactly which defers a function might have pending at each of its return points, it does not need a list at all. It allocates stack slots for each deferred call's function value and arguments, plus one small integer used as a **bitmask** — one bit per `defer` statement in the function. Executing a `defer` statement becomes: store the arguments into their slots and set that statement's bit. At each return, the compiler emits straight-line code that checks the bits from the highest-numbered statement down and calls the ones that are set. That preserves both LIFO order and the fact that a `defer` guarded by an `if` only runs when it was actually reached. The result is that in the overwhelmingly common case — a couple of defers near the top of a function — `defer` costs about as much as writing the calls by hand at each return. ## When open coding is not available The compiler gives up and uses runtime records when it cannot bound the set of pending defers statically or when the bookkeeping would not fit: - **A `defer` inside a loop.** The number of pending calls is a runtime quantity, so no fixed set of stack slots and no fixed bitmask can describe it. This is the case that costs real memory: one record per iteration, all alive until the function returns. - **Too many `defer` statements in one function.** The bitmask is small — eight statements is the limit — so a function with more than that falls back. - **Optimizations disabled.** Building with `-gcflags=-N`, as debugging builds do, turns open coding off. This matters when you benchmark: measuring a debug build tells you about the old code path, not the one that ships. ## Correctness on the non-return path A function can stop executing without reaching one of its return statements, and the deferred calls still have to run. With runtime records that is easy, because the goroutine's list is a data structure the runtime can walk. Open-coded defers live in stack slots and a bitmask that only the compiled code understands, so the compiler additionally emits function metadata describing where those slots and that mask are. The runtime reads that metadata when it has to unwind a frame itself, and runs the same calls in the same order. The optimisation is invisible from the outside — which is the whole point of allowing it. ## How you can tell which path a function got Compiled code that still uses records contains calls to the runtime's defer machinery: a call at the `defer` statement to push the record, and `runtime.deferreturn` at the exits. Open-coded functions contain neither; the deferred calls appear as ordinary calls guarded by bit tests. Dumping the assembly for a function with `go build -gcflags=-S` and looking for those runtime symbols is a direct answer to "did this get open-coded?" — far more reliable than reasoning about the source. ## Why this matters in practice Mostly because of the folklore it retires. "Never use `defer` in a hot path" was reasonable advice against the original heap-record implementation and is close to meaningless against open coding. What survives is the narrower and still-true version: a `defer` **inside a loop** is both a resource-lifetime problem and the one shape that reliably allocates per iteration. If a profile ever puts the defer machinery in front of you, that loop is where to look first.

  • Why does a defer inside a loop specifically defeat open coding?
    Open coding needs a fixed set of stack slots and one bit per `defer` *statement*, decided at compile time. A defer in a loop can be pending an unbounded number of times, so there is no fixed number of slots or bits that could represent it. The compiler falls back to runtime records, one per iteration, linked onto the goroutine's pending list until the function returns.
  • If open coding is invisible, how would you actually verify a function got it?
    Dump the function's assembly with `go build -gcflags=-S` and look for the runtime's defer symbols. A function that still uses records calls into the runtime to push each one and calls `runtime.deferreturn` at its exits; an open-coded function has neither, just the deferred calls guarded by bit tests. Do this on an optimized build — `-gcflags=-N` disables open coding.
  • Does open coding change when the deferred call's arguments are evaluated?
    No. Arguments are evaluated when the `defer` statement executes, exactly as before; open coding just stores those evaluated values into stack slots instead of copying them into a runtime record. Every user-visible rule — argument evaluation time, LIFO order, access to named results — is identical across all three implementations. Only the cost changed.

saying these in an interview costs you the question

  • Says every defer allocates on the heap
  • Claims open coding changed defer's semantics or ordering
  • Thinks the compiler open-codes defers inside loops too
  • Benchmarks a debug build and concludes defer is slow
  • Says defer is always tens of nanoseconds in current Go