skip to content

Removing Nondeterminism

The things that make a Go suite flaky - real time, leaked goroutines, randomised iteration - and what Go offers against them: synctest bubbles, goroutine counts, -shuffle and -race.

part ofGo (Golang)overview, primer and where to startread it →
on this pageshow

explore

questions

12

Why does a test asserting runtime.NumGoroutine() is unchanged fail intermittently?

level: middleimportance: must knowfreq 55%

answer

  1. the number is instantaneous
  2. signalling stop is not stopping
  3. background goroutines that legitimately stay
  4. one process, many parallel tests
  5. poll to a deadline, compare stacks

basics

~20 s

The number is process-wide and momentary: a goroutine told to stop may not have run its last statement, other packages start background goroutines lazily, and parallel tests share the count. Poll to a deadline and compare stacks.

solid answer

~50 s

`runtime.NumGoroutine()` counts every goroutine in the process at that instant, so an exact before-and-after equality is asserting something the runtime never promised. Three things break it. Shutdown is asynchronous: after you close a subscription the goroutine still has to be scheduled to run its final statements, so reading the count immediately is a race against the scheduler. Background goroutines are created lazily: the first HTTP request, the first timer, the first database call each leave long-lived goroutines alive that will still be there next time, so the very first test to touch them sees a permanent step up. And parallel tests share one process, so a sibling's goroutines land in your count. The fix is to retry the comparison against a deadline instead of asserting once, take the baseline after a warm-up so lazily started goroutines already exist, and compare goroutine *stacks* against a known-ignorable set rather than a raw number.

code

go · 16 lines
go
func assertNoExtraGoroutines(t *testing.T, before int) {
	t.Helper()
	deadline := time.Now().Add(2 * time.Second)
	for {
		after := runtime.NumGoroutine()
		if after <= before {
			return
		}
		if time.Now().After(deadline) {
			buf := make([]byte, 1<<16)
			n := runtime.Stack(buf, true)
			t.Fatalf("goroutines %d -> %d after 2s\n%s", before, after, buf[:n])
		}
		time.Sleep(10 * time.Millisecond)
	}
}

go deeper

for a junior

Remember that the count is taken across the whole process at one instant, and that a goroutine you just told to stop may not have finished yet. That timing gap is the usual reason the assertion flakes.

for a middle

Explain all three causes — asynchronous shutdown, lazily created background goroutines that legitimately stay, and parallel tests sharing the count — and describe the poll-to-a-deadline loop plus a warm-up baseline as the fix.

for a senior

Show that you would keep the check trustworthy: compare stacks not numbers, keep an explicit allowlist rather than a fudge factor, and be clear that a flaky leak detector is deleted within a week and therefore worse than none.

for a principal

Own the tradeoff between suite runtime and confidence: where the poll budget is spent, which components are worth checking at all, and how you stop the team from responding to a red leak check by loosening its threshold.

## The assertion that looks right and is not ```go before := runtime.NumGoroutine() doWork() if runtime.NumGoroutine() != before { t.Fatal("goroutine leak") } ``` This passes locally, passes ninety times in a row, and fails in the pipeline. The reason is that it treats a process-wide, instantaneous number as if it were a per-test, settled one. It is neither. ## Cause 1: goroutine shutdown is asynchronous Closing a channel, cancelling a `context.Context`, or setting a stop flag does not stop a goroutine. It makes the goroutine *runnable*. The goroutine still has to be picked up by Go's goroutine scheduler, run its remaining statements, run its deferred calls, and return. Reading `runtime.NumGoroutine()` on the very next line is a race against that, and the race is decided by machine load, by `GOMAXPROCS`, and by whether the test binary happens to be running other goroutines at that moment — all of which differ between a laptop and a busy CI runner. This is the dominant cause, and it is why the check must be a **poll**, not a single read: retry the comparison every few milliseconds until it succeeds or a deadline (a second or two is generous) expires, and only then fail. ## Cause 2: background goroutines that are created lazily and stay A lot of the standard library starts goroutines on first use and keeps them: - `net/http`'s transport runs a read loop and a write loop per pooled connection, and those survive the response so the connection can be reused; they go away when the connection is closed or idles out, or when you call `CloseIdleConnections` on the transport. - A `database/sql` pool keeps a connection opener and per-connection goroutines alive. - Timers, tickers and anything watching a signal or a file descriptor add their own. The first test that touches one of these sees the count step up permanently. That is not a leak; it is a one-time initialisation, and it is indistinguishable from a leak if you are only looking at a number. Two defences: take the baseline **after** a deliberate warm-up that exercises those paths once, and compare stacks so you can ignore the specific frames you have decided are benign. ## Cause 3: the count is process-wide, so parallelism poisons it `t.Parallel()` tests run concurrently in the same process. Every one of them sees every other one's goroutines. A per-test count check is therefore incompatible with parallelism — not flaky-in-principle, actively wrong — and the failure is order-dependent, so it moves around when the shuffle order changes. Either the leak-checking tests do not call `t.Parallel()`, or the check moves to a suite-level position where nothing else is running. ## Cause 4: a matching number is not a matching set Even when the count settles, equality is weak evidence. One goroutine returning while another leaks nets to zero. Comparing the goroutine profile's stacks, and reporting the ones that are new, is strictly stronger and produces a better failure message. ## What a robust check looks like 1. Warm up the paths that start persistent background goroutines. 2. Take a baseline — ideally the set of stacks from `pprof.Lookup("goroutine")`, not just `Count()`. 3. Run the code under test and shut it down. 4. Poll: re-read, compare, sleep a few milliseconds, repeat until a deadline. 5. On failure, write the profile with `WriteTo(w, 1)` or `WriteTo(w, 2)` so the message carries the parked call sites. ## Why not just sleep A fixed `time.Sleep` before the assertion is the tempting fix and the wrong one twice over: too short and it still flakes on a loaded machine, too long and every test in the suite pays for it. A poll with a deadline costs nothing in the common case, because the first read almost always succeeds, and it only spends the full budget when something is genuinely wrong. ## The failure this protects For a fan-out service that holds one subscription goroutine per connected client, this check is the only thing standing between a clean pipeline and an out-of-memory restart in production: each abandoned subscription keeps its goroutine and its per-subscriber buffer alive forever, and the number only ever goes up. A check that flakes gets disabled within a week, so making it deterministic is not polish — it is the difference between having the check and not.

  • Why is a fixed time.Sleep before the assertion the wrong fix?
    It buys latency without buying determinism. Any constant is simultaneously too short for a loaded CI machine — so the flake survives — and too long for every run where nothing was wrong, and that cost is multiplied by every test carrying the check. Polling against a deadline returns immediately in the common case and only spends the whole budget when the count genuinely refuses to settle, which is the case you wanted to fail anyway.
  • How do you keep a leak check from tripping on the standard library's own long-lived goroutines?
    Warm up first: exercise the paths that create them — one HTTP request, one database call — before taking the baseline, so those goroutines already exist on both sides of the comparison. Then compare stacks rather than counts and ignore a small, explicit list of frames you have decided are benign. An allowlist of stacks is auditable; a magic tolerance number added to the count is not.
  • Can a per-test goroutine-count check coexist with t.Parallel()?
    No. Parallel tests share one process and `runtime.NumGoroutine()` is process-wide, so each test sees the others' goroutines and the result depends on interleaving and on test order. Either the leak-checking tests stay serial, or the check moves to a point where nothing else in the package is running. Leaving both in place produces failures that move around whenever the shuffle order changes.

Counting cars in a car park the second the concert ends tells you nothing: people are still walking to them. Count again in two minutes, and note the registration plates rather than the total.

saying these in an interview costs you the question

  • Adds a time.Sleep and calls the flake fixed
  • Assumes a cancelled context stops the goroutine immediately
  • Treats every background goroutine as a leak
  • Runs a process-wide count check inside parallel tests
  • Widens the tolerance until the check stops failing
open as a page

A worker-pool test sleeps 50ms before asserting and goes red in CI weekly — how do you prove the sleep is the cause?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Make it fail on demand rather than reasoning about it: narrow with -run, repeat with -count, build with -race, and squeeze the parallelism the way CI does. Then vary the sleep constant — if the failure rate tracks it, the sleep is the synchronisation.

open as a page

Why does re-running `go test` after a flaky failure print `(cached)`, and how do you force a real re-run?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Go's test cache stores the result of a run that succeeded and replays it when nothing relevant changed, printing (cached) without executing anything. Add -count=1 to force a real run, or discard the stored results with go clean -testcache.

open as a page

How can a Go test detect that the code it exercised left a goroutine still running?

level: juniorimportance: should knowfreq 45%

basics

~20 s

Sample runtime.NumGoroutine() before the exercised code and again after it should have shut down; a higher number means something never exited. Dumping the runtime/pprof goroutine profile instead of counting also shows the leftover stacks, so you learn where.

open as a page

What does testing/synctest's Test function give a Go test that real time.Sleep calls cannot?

level: juniorimportance: should knowfreq 26%

basics

~20 s

synctest.Test runs a test inside a bubble with a fake clock, so time.Sleep, timers and tickers jump forward instantly instead of really waiting. A five-minute expiry can be tested in microseconds, with no slow-machine flakes.

open as a page

Why can GOMAXPROCS differ between your laptop and a CI container, and how do you test both?

level: middleimportance: should knowfreq 42%

basics

~20 s

GOMAXPROCS bounds how many goroutines run Go code at the same time and defaults to the machine's usable CPUs, so a ten-core laptop and a CPU-limited CI container run different amounts of real parallelism. Use go test -cpu=1,2,8 to exercise several values.

open as a page

What does a Go test binary print when the go test -timeout deadline expires?

level: middleimportance: should knowfreq 58%

basics

~20 s

It panics with a message like "panic: test timed out after 10m0s", names the tests still running and how long they have run, and then prints the stack of every goroutine in the process. The package is reported as failed.

open as a page

When does synctest.Wait return, and which blocked goroutines count as durably blocked?

level: middleimportance: should knowfreq 38%

basics

~20 s

synctest.Wait returns once every other goroutine in the caller's bubble is durably blocked: blocked so that only another goroutine in that same bubble could unblock it, such as a bubble-created channel operation, time.Sleep, WaitGroup.Wait or Cond.Wait.

open as a page

Why does synctest.Test panic when a goroutine started inside the bubble is still alive at the end?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Because the bubble owns every goroutine started inside it and must be empty when the test function returns. A sweeper or ticker loop nobody stopped cannot be left running on a fake clock, so synctest.Test panics instead of passing quietly.

open as a page

What does `go test -shuffle=on` randomise, and how do you replay the order that failed?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

It randomises the execution order of the top-level tests and benchmarks inside each test binary, which exposes tests that secretly depend on one another. The seed used is printed in the output; pass it back as -shuffle=SEED to replay that exact order.

open as a page

Would you run a goroutine-leak check in TestMain or in each Go test, and why?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

TestMain gives one cheap package-wide check between m.Run and os.Exit, but names no culprit. Per-test checks attribute the leak exactly, cost a poll each and rule out t.Parallel. Most suites use TestMain plus per-test checks on the few goroutine-owning components.

open as a page

Why would a test using testing/synctest hang until the test timeout instead of its clock advancing?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

Because a goroutine in the bubble is blocked on something outside it — real I/O or a channel created before the bubble — so it is never durably blocked, the bubble never settles, and the fake clock will not move.

open as a page