skip to content

A TestMain does defer lis.Close() then os.Exit(m.Run()), yet the listener leaks. Why, and what is the fix?

level: seniorimportance: should knowfreq 48%

answer

  1. the process does not unwind on the way out
  2. the cleanup is registered, never invoked
  3. one of defer and os.Exit has to move
  4. keep the code from m.Run, exit last
  5. put the defers in a helper returning int

basics

~20 s

os.Exit terminates the process immediately and runs no deferred function, so the deferred Close never fires. Capture the code from m.Run, run the teardown, and only then call os.Exit - or move the defers into a helper that returns the code.

solid answer

~50 s

`os.Exit` does not unwind: it ends the process on the spot without running any deferred function, so a `defer lis.Close()` written at the top of TestMain never executes. Everything the suite opened - the listener, a spawned helper process, a temp tree - survives the run, and on a developer machine or a long-lived CI agent those accumulate until the next run cannot bind the port. There are two fixes and I prefer the second. Either drop the `defer` and order it explicitly: `code := m.Run(); lis.Close(); os.Exit(code)`. Or keep `defer` working by putting the body in a helper - `func run(m *testing.M) int` that defers its cleanup and returns `m.Run()` - and write `os.Exit(run(m))` in TestMain, so `os.Exit` is called outside the frame holding the defers. The helper scales better once there are three or four resources to release in reverse order.

code

go · 13 lines
go
func TestMain(m *testing.M) {
	os.Exit(run(m))
}

func run(m *testing.M) int {
	lis, err := net.Listen("tcp", "127.0.0.1:0")
	if err != nil {
		fmt.Fprintln(os.Stderr, "listen:", err)
		return 1
	}
	defer lis.Close() // runs: os.Exit is called in the caller
	return m.Run()
}

go deeper

for a junior

Remember the rule: os.Exit ends the process straight away and skips every deferred call. Put cleanup between m.Run and os.Exit rather than in a defer.

for a middle

Explain why the defer is registered but never invoked, and show both fixes - explicit ordering, or a helper function that defers cleanup and returns the code that TestMain passes to os.Exit.

for a senior

Diagnose it from the symptom: orphaned processes, ports that will not rebind, temp directories accumulating on CI agents. Then design cleanup that is best-effort - port 0, per-run names, sweep on startup - because timeouts and kills skip teardown anyway.

for a principal

Set the standard shape once and make it reviewable, and decide how much a suite is allowed to depend on external state at all. A fixture that cannot leak beats a teardown path that usually runs.

## What actually happens `os.Exit` is documented to terminate the program immediately, and **deferred functions are not run**. It is not a return; nothing unwinds. So this shape: ```go func TestMain(m *testing.M) { lis, _ := net.Listen("tcp", "127.0.0.1:0") defer lis.Close() // never runs os.Exit(m.Run()) } ``` looks like every other Go function you have written and behaves unlike all of them. The `defer` compiles, the linters are quiet, the tests pass. The only evidence is what is left behind. This bites specifically in TestMain because TestMain is one of the very few functions in ordinary Go code that ends in `os.Exit`. The habit of pairing acquisition with `defer` is correct everywhere else, which is exactly why it survives code review here. ## How you notice The symptom is environmental, not a test failure: - a helper process the suite started is still running after `go test` returns; - the next run fails to bind because the previous one still holds the fixed port; - temporary directories pile up under the system temp directory on every developer's machine; - a long-lived CI runner degrades over days and gets "fixed" by a reboot. The engineer who inherits a suite like this usually finds it from the other end: `go test ./...` fails to start on a machine that has run the suite before, and works on a fresh one. ## Fix one: order it explicitly ```go func TestMain(m *testing.M) { lis, err := net.Listen("tcp", "127.0.0.1:0") if err != nil { fmt.Fprintln(os.Stderr, "listen:", err) os.Exit(1) } code := m.Run() lis.Close() os.Exit(code) } ``` Simple and obvious for one resource. It degrades badly with several: every early-exit path in setup now has to release whatever was already acquired, by hand, in the right order. ## Fix two: keep defer, move os.Exit out ```go func TestMain(m *testing.M) { os.Exit(run(m)) } func run(m *testing.M) int { lis, err := net.Listen("tcp", "127.0.0.1:0") if err != nil { fmt.Fprintln(os.Stderr, "listen:", err) return 1 } defer lis.Close() dir, err := os.MkdirTemp("", "transport") if err != nil { return 1 } defer os.RemoveAll(dir) return m.Run() } ``` `run` is an ordinary function, so its defers run when it returns, in reverse order, on every path including early returns. `os.Exit` is called afterwards, in TestMain, where there is nothing left to defer. This is the pattern to reach for the moment more than one resource is involved, and it is the same trick used in `main` functions that need both cleanup and a real exit status. (A third option exists: let TestMain simply return instead of calling `os.Exit`, since the generated test main exits with the status `m.Run` recorded - which makes deferred cleanup in TestMain work. It is correct but easy to break later, because the first person who adds an `os.Exit` back silently disables every defer above it.) ## What no fix covers Even a correct teardown path does not run when the process dies rather than returns: - the run exceeds `go test -timeout` and the binary is killed after dumping stacks; - someone interrupts the run from the terminal; - the process is killed by the CI system, or the machine dies. So cleanup is best-effort, and robust suites are designed not to need it: - **bind to port 0** and let the kernel assign a free port, instead of a fixed one that a corpse can hold; - **name external resources uniquely per run**, so a leftover never collides with a new one; - **sweep at startup**: if the fixture can detect and remove its own leftovers before creating new ones, a leaked resource costs one dirty run instead of every future run; - **prefer in-process fixtures** where possible - a listener in the same process dies with it, a spawned helper does not. ## The review rule In a function that ends in `os.Exit`, treat every `defer` above it as dead code. In TestMain specifically: if you see `defer` and `os.Exit` in the same function body, one of them is wrong.

  • Would returning from TestMain instead of calling os.Exit make the deferred Close run?
    Yes - a normal return unwinds the frame, so the defers execute, and the generated test main then exits with the status `m.Run` recorded. It works, but it is fragile: the next person who adds an `os.Exit` for an early setup failure silently disables every defer above it. The helper-returning-int shape is harder to break.
  • Which cleanup failures survive even a correct teardown path?
    Anything that kills the process instead of returning: exceeding `go test -timeout`, an interrupt from the terminal, or the CI system killing the job. Design so leftovers are harmless - bind to port 0, name resources per run, and sweep leftovers at startup - rather than assuming teardown always gets its turn.
  • How would you stop this pattern from coming back in review?
    Make the rule mechanical: in any function ending in `os.Exit`, every `defer` above it is dead code, so `defer` and `os.Exit` must not appear in the same body. Standardise on `os.Exit(run(m))` with all fixture code inside `run`, so there is one shape reviewers recognise across every package.
  • Why is a fixed shared port worse than binding to port 0 here?
    A fixed port turns one leaked listener into a permanently broken machine: every later run fails to bind until someone hunts down the process. Binding `127.0.0.1:0` lets the kernel pick a free port and the suite reads the real address back from the listener, so a leftover from an earlier run costs nothing.

os.Exit is pulling the power cord, not walking out of the room: nothing on the way to the door gets done.

saying these in an interview costs you the question

  • Believes os.Exit runs deferred functions first
  • Puts teardown in a defer above os.Exit in TestMain
  • Thinks the leak is a testing package bug
  • Assumes teardown always runs, even on timeout kills
  • Uses a fixed shared port that a leaked process can hold