skip to content

Why does a test asserting runtime.NumGoroutine() is unchanged fail intermittently?

level: middleimportance: must knowfreq 55%

answer

  1. the number is instantaneous
  2. signalling stop is not stopping
  3. background goroutines that legitimately stay
  4. one process, many parallel tests
  5. poll to a deadline, compare stacks

basics

~20 s

The number is process-wide and momentary: a goroutine told to stop may not have run its last statement, other packages start background goroutines lazily, and parallel tests share the count. Poll to a deadline and compare stacks.

solid answer

~50 s

`runtime.NumGoroutine()` counts every goroutine in the process at that instant, so an exact before-and-after equality is asserting something the runtime never promised. Three things break it. Shutdown is asynchronous: after you close a subscription the goroutine still has to be scheduled to run its final statements, so reading the count immediately is a race against the scheduler. Background goroutines are created lazily: the first HTTP request, the first timer, the first database call each leave long-lived goroutines alive that will still be there next time, so the very first test to touch them sees a permanent step up. And parallel tests share one process, so a sibling's goroutines land in your count. The fix is to retry the comparison against a deadline instead of asserting once, take the baseline after a warm-up so lazily started goroutines already exist, and compare goroutine *stacks* against a known-ignorable set rather than a raw number.

code

go · 16 lines
go
func assertNoExtraGoroutines(t *testing.T, before int) {
	t.Helper()
	deadline := time.Now().Add(2 * time.Second)
	for {
		after := runtime.NumGoroutine()
		if after <= before {
			return
		}
		if time.Now().After(deadline) {
			buf := make([]byte, 1<<16)
			n := runtime.Stack(buf, true)
			t.Fatalf("goroutines %d -> %d after 2s\n%s", before, after, buf[:n])
		}
		time.Sleep(10 * time.Millisecond)
	}
}

go deeper

for a junior

Remember that the count is taken across the whole process at one instant, and that a goroutine you just told to stop may not have finished yet. That timing gap is the usual reason the assertion flakes.

for a middle

Explain all three causes — asynchronous shutdown, lazily created background goroutines that legitimately stay, and parallel tests sharing the count — and describe the poll-to-a-deadline loop plus a warm-up baseline as the fix.

for a senior

Show that you would keep the check trustworthy: compare stacks not numbers, keep an explicit allowlist rather than a fudge factor, and be clear that a flaky leak detector is deleted within a week and therefore worse than none.

for a principal

Own the tradeoff between suite runtime and confidence: where the poll budget is spent, which components are worth checking at all, and how you stop the team from responding to a red leak check by loosening its threshold.

## The assertion that looks right and is not ```go before := runtime.NumGoroutine() doWork() if runtime.NumGoroutine() != before { t.Fatal("goroutine leak") } ``` This passes locally, passes ninety times in a row, and fails in the pipeline. The reason is that it treats a process-wide, instantaneous number as if it were a per-test, settled one. It is neither. ## Cause 1: goroutine shutdown is asynchronous Closing a channel, cancelling a `context.Context`, or setting a stop flag does not stop a goroutine. It makes the goroutine *runnable*. The goroutine still has to be picked up by Go's goroutine scheduler, run its remaining statements, run its deferred calls, and return. Reading `runtime.NumGoroutine()` on the very next line is a race against that, and the race is decided by machine load, by `GOMAXPROCS`, and by whether the test binary happens to be running other goroutines at that moment — all of which differ between a laptop and a busy CI runner. This is the dominant cause, and it is why the check must be a **poll**, not a single read: retry the comparison every few milliseconds until it succeeds or a deadline (a second or two is generous) expires, and only then fail. ## Cause 2: background goroutines that are created lazily and stay A lot of the standard library starts goroutines on first use and keeps them: - `net/http`'s transport runs a read loop and a write loop per pooled connection, and those survive the response so the connection can be reused; they go away when the connection is closed or idles out, or when you call `CloseIdleConnections` on the transport. - A `database/sql` pool keeps a connection opener and per-connection goroutines alive. - Timers, tickers and anything watching a signal or a file descriptor add their own. The first test that touches one of these sees the count step up permanently. That is not a leak; it is a one-time initialisation, and it is indistinguishable from a leak if you are only looking at a number. Two defences: take the baseline **after** a deliberate warm-up that exercises those paths once, and compare stacks so you can ignore the specific frames you have decided are benign. ## Cause 3: the count is process-wide, so parallelism poisons it `t.Parallel()` tests run concurrently in the same process. Every one of them sees every other one's goroutines. A per-test count check is therefore incompatible with parallelism — not flaky-in-principle, actively wrong — and the failure is order-dependent, so it moves around when the shuffle order changes. Either the leak-checking tests do not call `t.Parallel()`, or the check moves to a suite-level position where nothing else is running. ## Cause 4: a matching number is not a matching set Even when the count settles, equality is weak evidence. One goroutine returning while another leaks nets to zero. Comparing the goroutine profile's stacks, and reporting the ones that are new, is strictly stronger and produces a better failure message. ## What a robust check looks like 1. Warm up the paths that start persistent background goroutines. 2. Take a baseline — ideally the set of stacks from `pprof.Lookup("goroutine")`, not just `Count()`. 3. Run the code under test and shut it down. 4. Poll: re-read, compare, sleep a few milliseconds, repeat until a deadline. 5. On failure, write the profile with `WriteTo(w, 1)` or `WriteTo(w, 2)` so the message carries the parked call sites. ## Why not just sleep A fixed `time.Sleep` before the assertion is the tempting fix and the wrong one twice over: too short and it still flakes on a loaded machine, too long and every test in the suite pays for it. A poll with a deadline costs nothing in the common case, because the first read almost always succeeds, and it only spends the full budget when something is genuinely wrong. ## The failure this protects For a fan-out service that holds one subscription goroutine per connected client, this check is the only thing standing between a clean pipeline and an out-of-memory restart in production: each abandoned subscription keeps its goroutine and its per-subscriber buffer alive forever, and the number only ever goes up. A check that flakes gets disabled within a week, so making it deterministic is not polish — it is the difference between having the check and not.

  • Why is a fixed time.Sleep before the assertion the wrong fix?
    It buys latency without buying determinism. Any constant is simultaneously too short for a loaded CI machine — so the flake survives — and too long for every run where nothing was wrong, and that cost is multiplied by every test carrying the check. Polling against a deadline returns immediately in the common case and only spends the whole budget when the count genuinely refuses to settle, which is the case you wanted to fail anyway.
  • How do you keep a leak check from tripping on the standard library's own long-lived goroutines?
    Warm up first: exercise the paths that create them — one HTTP request, one database call — before taking the baseline, so those goroutines already exist on both sides of the comparison. Then compare stacks rather than counts and ignore a small, explicit list of frames you have decided are benign. An allowlist of stacks is auditable; a magic tolerance number added to the count is not.
  • Can a per-test goroutine-count check coexist with t.Parallel()?
    No. Parallel tests share one process and `runtime.NumGoroutine()` is process-wide, so each test sees the others' goroutines and the result depends on interleaving and on test order. Either the leak-checking tests stay serial, or the check moves to a point where nothing else in the package is running. Leaving both in place produces failures that move around whenever the shuffle order changes.

Counting cars in a car park the second the concert ends tells you nothing: people are still walking to them. Count again in two minutes, and note the registration plates rather than the total.

saying these in an interview costs you the question

  • Adds a time.Sleep and calls the flake fixed
  • Assumes a cancelled context stops the goroutine immediately
  • Treats every background goroutine as a leak
  • Runs a process-wide count check inside parallel tests
  • Widens the tolerance until the check stops failing