skip to content

Your team bans `defer` in hot Go functions. How would you check whether that rule still pays?

level: seniorimportance: nice to knowfreq 30%

answer

  1. measure before you inherit an opinion
  2. benchmark the function, not the language
  3. optimizations off changes the answer
  4. look for the runtime defer symbols
  5. a defer in a loop is the real cost

basics

~20 s

Measure rather than argue. Benchmark the real function with and without the defer on an optimized build, then read its assembly for the runtime defer symbols. On current Go the rule usually only survives for defers inside loops.

solid answer

~50 s

Start by reproducing the claim: write a benchmark around the actual hot function, run it with `-benchmem`, and compare it against a version with the cleanup written out at each return. Then look at what the compiler did — dump the function with `go build -gcflags=-S` and check for the runtime's defer symbols, `runtime.deferproc` at the statement and `runtime.deferreturn` at the exits. Their absence means the defers were open-coded and cost roughly a bit test plus a direct call, so the rule is buying nothing. Their presence points at a real cause, almost always a `defer` inside a loop or more than eight `defer` statements in one function. Make sure you measure an optimized build: `-gcflags=-N`, which debug builds use, disables open coding and will manufacture evidence for a rule that does not hold in production. Then weigh what the ban costs: hand-written cleanup at every return is exactly how an unlock or a close gets missed when someone adds an early return later.

code

go · 19 lines
go
func BenchmarkWithDefer(b *testing.B) {
	var mu sync.Mutex
	for i := 0; i < b.N; i++ {
		func() {
			mu.Lock()
			defer mu.Unlock()
		}()
	}
}

func BenchmarkWithoutDefer(b *testing.B) {
	var mu sync.Mutex
	for i := 0; i < b.N; i++ {
		func() {
			mu.Lock()
			mu.Unlock()
		}()
	}
}

go deeper

for a junior

Know that a performance rule should come with a measurement, and that on current Go a plain defer outside a loop is usually too cheap to notice.

for a middle

Be able to describe the two checks: a benchmark of the real function with allocation reporting, and a look at the compiled output for the runtime defer symbols.

for a senior

Show that you weigh both sides. Name the build flag that would fake the result, identify the loop case as the real culprit, and account for the correctness cost of hand-written cleanup on every exit path.

for a principal

Own the outcome as policy. Turn folklore into a narrow written rule with evidence attached, decide what threshold justifies hand-rolling cleanup, and make sure new joiners inherit the reasoning rather than only the prohibition.

## The rule you are auditing "Do not use `defer` in hot paths" is a genuine piece of Go folklore with a real origin. In the original implementation, every `defer` statement asked the runtime to create a pending-call record and push it onto the goroutine's list, and every exit called into the runtime to pop and run them. In a function that otherwise does almost nothing, that overhead was easy to see in a profile. Teams wrote the rule down, and the rule outlived the implementation it was written against. Since Go 1.14, the compiler open-codes defers in the common case: it stores the deferred call's values in stack slots, sets one bit of a small per-function bitmask when the `defer` statement executes, and emits inline code at each return that tests the bits and calls in reverse order. No record, no list, no runtime call. The cost of the mechanism collapses to about a branch. So the honest answer to "does the rule still pay?" is a measurement, and the measurement has two halves. ## Half one: benchmark the real thing Write a benchmark around the function that is actually hot, not around a synthetic `func() { defer nothing() }`. Produce two variants — one with `defer`, one with the cleanup written by hand at every exit — and run them with `-benchmem` so you see allocations as well as time. Two failure modes to avoid: - **Measuring a debug build.** `-gcflags=-N` turns off optimizations and therefore turns off open coding. If your benchmark harness or your editor's test runner passes it, you are measuring the old implementation and will "prove" the rule. - **Measuring the mechanism instead of the workload.** A microbenchmark of nothing but a defer can show a difference that is real and simultaneously irrelevant next to the work the function actually does. Compare at the level of the function you care about. ## Half two: read what the compiler emitted Benchmarks tell you *whether*; the assembly tells you *why*. Dump the function with `go build -gcflags=-S` and look for the runtime's defer symbols. A function that still uses records calls `runtime.deferproc` where the statement is and `runtime.deferreturn` at its exits. An open-coded function contains neither, only the deferred calls themselves guarded by bit tests. If you find those symbols, you have a specific, fixable cause rather than a language-level indictment: - **A `defer` inside a loop.** The pending count is unknown at compile time, so open coding is impossible and each iteration creates a record. This is also the shape that holds resources far too long, so fixing it improves correctness and cost together, usually by extracting the loop body into a function that is called once per iteration. - **More than eight `defer` statements in the function.** The bitmask does not stretch. A function with that many deferred calls is usually asking to be split anyway. ## Half three: what the ban itself costs This is the part an engineer new to the codebase tends to skip. Removing `defer` means every exit path has to repeat the cleanup by hand, and functions grow exit paths over time. The failure mode is not theoretical: someone adds an early `return` for a validation case six months later, does not repeat the unlock or the close, and a mutex is held forever or a descriptor leaks. `defer` exists because that class of bug is common and expensive, and a blanket ban trades a real correctness guarantee for a cost that in most functions is now unmeasurable. ## Landing the recommendation A good outcome is not "the rule is wrong", it is a narrower rule with evidence behind it. Something like: `defer` is fine everywhere; never put one inside a loop; if a function is measured hot *and* its assembly shows the runtime defer symbols, restructure that specific function and record the number in the commit message. That version is enforceable, it survives contact with new joiners, and it does not ask anyone to give up automatic cleanup on faith. ## What an interviewer is listening for That you reach for a measurement before an opinion; that you know the mechanism well enough to name what you would look for in the compiled output; that you know the one build flag that would invalidate the measurement; and that you weigh the correctness cost of the rule and not only its performance benefit. Answering only "defer is fast now" is true and half the answer.

  • Your benchmark shows defer costing nothing. What do you propose the rule becomes?
    Replace the blanket ban with two enforceable clauses: never put a `defer` inside a loop, and restructure a specific function only when it is measured hot and its compiled output shows the runtime defer symbols. That keeps automatic cleanup everywhere it protects correctness, and reserves hand-written cleanup for functions where someone has a number. Record the number alongside the change so the next reader does not re-litigate it.
  • What build flag would silently invalidate this whole measurement?
    `-gcflags=-N`, which disables optimizations and therefore disables open-coded defers. A build made that way uses runtime defer records everywhere, so it reproduces the pre-1.14 cost and appears to confirm the ban. Debug and some editor-driven test runs pass it by default, so check how the binary under test was built before trusting any comparison.
  • If the assembly does show the runtime defer symbols, what do you look at first?
    Whether a `defer` sits inside a loop in that function. That is by far the most common reason open coding is unavailable, and it creates one pending record per iteration on top of holding whatever resource until the function returns. Extracting the loop body into a per-iteration function usually removes the records and the resource-lifetime problem in the same edit.

saying these in an interview costs you the question

  • Repeats the rule without measuring anything
  • Benchmarks a build made with optimizations disabled
  • Removes defer everywhere and leaves exit paths uncovered
  • Assumes every defer allocates a runtime record
  • Microbenchmarks defer alone and ignores the surrounding work