skip to content

Would you run a goroutine-leak check in TestMain or in each Go test, and why?

level: seniorimportance: nice to knowfreq 33%

answer

  1. one net, or one alarm per room
  2. who pays when it goes red
  3. os.Exit runs no deferred calls
  4. process-wide counts and parallel tests
  5. an upper bound and a lower bound

basics

~20 s

TestMain gives one cheap package-wide check between m.Run and os.Exit, but names no culprit. Per-test checks attribute the leak exactly, cost a poll each and rule out t.Parallel. Most suites use TestMain plus per-test checks on the few goroutine-owning components.

solid answer

~50 s

In `TestMain` the check runs once: take a baseline before `m.Run()`, compare after it returns, and fail the package by changing the exit code you hand to `os.Exit`. It costs one poll for the entire package and nothing else — but it reports that *something* in the package leaked, not what, and it turns one bad test into a red package that has to be bisected. A per-test check names the culprit immediately, but the count is process-wide so those tests cannot call `t.Parallel()`, and each one pays its poll. The shape I would ship is both: a `TestMain` net that catches anything, plus explicit per-test checks on the handful of components that actually own long-lived goroutines — subscriptions, background pollers, connection readers. Both should print the goroutine profile's stacks on failure, because a package-level failure that says only "3 goroutines left" is a bisect, whereas a repeated parked stack is a fix.

code

go · 13 lines
go
func TestMain(m *testing.M) {
	warmUp() // create the lazily started background goroutines first
	before := pprof.Lookup("goroutine").Count()

	code := m.Run()

	if after := waitForGoroutines(before, 2*time.Second); after > before && code == 0 {
		_ = pprof.Lookup("goroutine").WriteTo(os.Stderr, 1)
		fmt.Fprintf(os.Stderr, "leaked goroutines: %d -> %d\n", before, after)
		code = 1
	}
	os.Exit(code)
}

go deeper

for a junior

Know that TestMain is where package-wide setup and teardown lives, that the tests run only when you call m.Run, and that os.Exit ends the process without running deferred functions.

for a middle

Explain the mechanics of each placement: where the check sits relative to m.Run and os.Exit, how the exit code must be preserved, and why a process-wide count cannot be compared inside a parallel test.

for a senior

Argue the allocation — broad net in TestMain, precise checks on the components that own goroutines — and insist the failure prints grouped stacks. Be able to contrast a noisy upper bound against the leak profile's high-confidence lower bound.

for a principal

Own the durability question: a leak check that flakes or reports numbers gets loosened and then deleted. Decide what the suite pays in runtime and lost parallelism, and how leak failures get triaged rather than silenced.

## The two placements ### In TestMain ```go func TestMain(m *testing.M) { warmUp() before := pprof.Lookup("goroutine").Count() code := m.Run() // check here os.Exit(code) } ``` `TestMain` runs instead of the tests; nothing happens until you call `m.Run()`, and nothing after the `os.Exit` call runs at all — `os.Exit` does not run deferred functions. So the check has exactly one place to live: between the return of `m.Run()` and the exit, and it must fold its verdict into the exit code rather than replacing it, or a real test failure gets masked by a passing leak check. **What it buys.** One check for the whole package, one poll budget, no interaction with any individual test. It sees leaks that no single test would attribute to itself, including a goroutine started by package initialisation. **What it costs.** No attribution. It knows the package ended with goroutines it did not start with; it cannot say which of two hundred tests is responsible. If the suite is shuffled or parallel, the leak may not even reproduce in the same place. And when it goes red, every developer in that package is blocked by one person's bug. ### In each test Baseline at the top, comparison deferred at the bottom, poll to a deadline. **What it buys** is precision: the failing test *is* the answer, and the stacks it prints are from a process with far less noise in it. **What it costs** is real: the count is process-wide, so such a test cannot run under `t.Parallel()` — a parallel sibling's goroutines land in the count and the failure becomes order-dependent. Each check spends a poll. And it must be maintained on every test, where it will be copy-pasted into tests that do not need it. ## How to choose The question is really about who pays for the failure. A `TestMain` check pays out in coverage and charges the whole package for triage; a per-test check pays out in a precise message and charges the suite in runtime and in lost parallelism. Since leaks are concentrated — almost always in the components that own long-lived goroutines — the efficient allocation is not uniform: - **Per-test checks on the goroutine-owning components.** For a fan-out service holding one subscription goroutine per connected client, the subscribe/unsubscribe/shutdown tests get an explicit check. Those are the tests where a leak is both likely and expensive. - **A `TestMain` net for everything else**, so a leak introduced somewhere unexpected is not invisible. - **Neither check anywhere without stacks.** Both should write `pprof.Lookup("goroutine").WriteTo(w, 1)` on failure, which groups goroutines by identical stack — the format that makes forty leaked subscription goroutines read as one bug. ## The lower bound and the upper bound A count-or-profile comparison is an **upper bound** on leaks: everything still alive is reported, including background goroutines that were always going to stay, so it produces false positives and needs a warm-up baseline and an allowlist of benign stacks. Go 1.27 shipped a goroutine-leak profile in `runtime/pprof`, which reports goroutines the runtime can *prove* will never resume. That is the opposite bias: essentially no false positives, but it stays silent about a goroutine that is merely idle and could in principle be woken — one blocked on a channel something still holds a reference to, or looping on a live ticker. It is a **lower bound**. The two are complementary: the profile-difference check tells you what is suspicious, the leak profile tells you what is certainly broken. ## What actually makes the check survive The failure mode of leak checking is social, not technical. A check that fires occasionally with an unhelpful message gets an allowlist entry, then a raised threshold, then a deletion, and the team is back to discovering leaks from a memory graph after an out-of-memory restart. So the criteria are: it must be deterministic (poll, do not sleep), it must be specific (stacks, not counts), it must be cheap where it is applied broadly, and its allowlist must be an explicit list of stacks someone can read in review rather than a tolerance number nobody can justify.

  • What must a TestMain leak check be careful about around os.Exit?
    `os.Exit` runs no deferred functions, so the check cannot be deferred — it has to sit between `m.Run()` returning and the exit call. It also has to combine with the run's own result rather than replace it: exit with the code `m.Run()` returned unless the leak check itself wants to fail, otherwise a genuine test failure is masked by a clean leak check, or a leak masks which tests actually failed.
  • How does the Go 1.27 goroutine-leak profile change this picture?
    It reports goroutines the runtime can prove will never resume, so it has essentially no false positives and needs no warm-up or allowlist. But it is a lower bound: a goroutine merely idling on a channel something still references, or looping on a live ticker, is not provably stuck and does not appear. Use it as the high-confidence signal that always deserves a fix, and keep the count-or-profile comparison as the broader, noisier net.
  • What would make you not add a leak check to a package at all?
    When the package owns no long-lived goroutines — pure functions, parsers, table-driven logic — the check can only ever report other people's background goroutines, so it is pure cost and pure noise. Leaks concentrate in the code that starts goroutines it must later stop. Spending the check there and nowhere else keeps the suite fast and, more importantly, keeps the signal high enough that nobody reaches for the delete key.
  • Why report stacks rather than a count when the check fails?
    Because a count is a bisect and a stack is a fix. "Goroutines went from 12 to 52" starts a search across every test in the package; forty goroutines sharing one stack parked on a channel send names the line to change and simultaneously proves it is one bug, not forty. Writing the goroutine profile with debug level 1 groups them exactly that way, which is why the failure message should always carry it.

A smoke alarm in the corridor tells you the building is on fire; one in each room tells you which room. You want the corridor alarm everywhere and room alarms in the kitchen.

saying these in an interview costs you the question

  • Defers the leak check and calls os.Exit afterwards
  • Overwrites the exit code from m.Run with the check's result
  • Adds a process-wide count check to parallel tests
  • Fails with a bare number and no goroutine stacks
  • Puts the check on every test in the repository uniformly
  • Responds to a red leak check by widening the tolerance