skip to content

Under `go test -shuffle=on` a package's tests fail, though each passes alone. How do you find the cause?

level: seniorimportance: should knowfreq 45%

answer

  1. one test binary per package, one process
  2. capture the seed, then reuse it
  3. the cache can replay a stale pass
  4. bisect with -run until two names remain
  5. somebody wrote a package variable and left it

basics

~10 s

Reproduce with the seed the shuffle prints, add -count=1 to defeat the test cache, then bisect with -run. The culprit is usually a package-level variable that one test writes and never restores.

solid answer

~50 s

Start from the mechanism: `go test` builds one test binary per package and runs that package's tests in a single process, so every test sees the same package-level variables. `-shuffle=on` randomizes the order of the top-level tests and prints the seed it used; pass that seed back with `-shuffle=<seed>` and add `-count=1` so a cached result cannot hide the failure. Then bisect with `-run 'TestA|TestB'` until you have the smallest failing pair. The offender is a test that mutates shared state and does not put it back: a package-level client or configuration variable, a swapped clock function, an entry added to a global registry, or a handler registered on a shared default mux. Fix it structurally by constructing that state per test, and where a global genuinely must be touched, restore it with `t.Cleanup` rather than `defer`, since `t.Cleanup` on a parent runs after its parallel subtests finish.

code

go · 11 lines
go
// Shared by every test in this package's test binary.
var defaultTimeout = 5 * time.Second

func TestSlowPath(t *testing.T) {
	old := defaultTimeout
	defaultTimeout = time.Millisecond
	// Runs after the test ends, including after any parallel subtests.
	t.Cleanup(func() { defaultTimeout = old })

	// ... exercise the slow path
}

go deeper

for a junior

Be ready to say that all tests in one Go package run inside a single process and share that package's variables, so state left over from one test is visible to the next.

for a middle

Explain the reproduction: capture the shuffle seed, replay it, add -count=1 to bypass the cached result, then bisect with -run. Know why t.Cleanup is a better restore point than defer.

for a senior

Demonstrate the triage under pressure, especially when the job is green locally and red in CI. Name the usual writers of shared state and argue for the structural fix over restore-and-hope, including turning shuffling on in CI.

for a principal

Frame it as a policy question: which suites run shuffled, whether parallel tests are the default, and what you require of a test that touches process-wide state before it is allowed to merge.

## Why the failure exists at all `go test` compiles one test binary per package and runs all of that package's tests inside a single process. Package-level variables are therefore shared by every test in the package, initialized once before the first test function runs, and never reset. Tests in *different* packages get different processes, which is why this class of failure never crosses a package boundary — and why "it passes when I run just that file" is not evidence of anything. The default execution order is source order, which is stable enough that a suite can carry a hidden dependency for months. `-shuffle=on` randomizes the order of top-level tests and benchmarks, exposing it. Because the seed is random per run, the failure looks intermittent: green on your machine, red in the job that happened to draw a bad order. ## Making it reproducible first Do not start reading code. Get a deterministic reproduction: 1. Run with `-shuffle=on` and capture the seed it prints. 2. Re-run with that seed. The same order gives the same failure. 3. Add `-count=1`. Without it, a cached successful result can be replayed and you will chase a ghost. `-count=1` is the standard way to force the test to actually run. 4. If the failure still moves, suspect concurrency rather than ordering, and run the same command under the race detector. ## Narrowing it to a pair Order dependence is almost always a relationship between two tests: one leaves state behind, another reads it. Use `-run` with a regular expression to run subsets — `-run 'TestA|TestB'` — and bisect the failing order until you have the smallest set that still fails. Two names is usually the answer, and it is often obvious which of the two is the writer once you see them together. A useful shortcut: run the suspected victim immediately after each other test in turn. Whichever predecessor makes it fail is the writer. ## What the writer usually is In practice the shared state is one of a small set of things: - A package-level configuration or client variable that a test swaps for a fake. - A clock: a package-level `var now = time.Now` that a test replaces with a fixed function. - A global registry map that a test adds an entry to. - A handler registered on a shared default multiplexer, which a second registration of the same path will panic on. - Process-wide state that is not yours: the environment, the working directory, a temporary file at a fixed path. All of them have the same shape: something mutable that outlives a single test. ## Restoring correctly, if you must touch a global When a global genuinely has to be swapped, restore it with `t.Cleanup` rather than `defer`. Two reasons. `t.Cleanup` still runs when the test ends early through `t.Fatal`, which unwinds via `runtime.Goexit` — a `defer` in the test function does run in that case, but a `defer` in a helper the test called does not help you once the helper has returned. More importantly, a test that calls `t.Parallel` is paused and resumed only after its parent's function body has returned, so a `defer` in the parent restores the value *before* the parallel children run, while a `t.Cleanup` registered on the parent runs after they have all finished. For environment variables specifically, `t.Setenv` sets and restores for you, and it deliberately refuses to be used from a parallel test — because there is one environment per process and no way to isolate it. Be aware that `t.Parallel` turns order dependence into a genuine data race: two tests mutating the same variable at the same time. Running the suite under the race detector will find that, and it is worth doing on this class of bug even though the ordering failure itself is not a race. ## The structural fix Restoring state is damage control. The fix that ends the category is to stop sharing: make the thing a field on a struct that each test constructs, pass the clock in as a parameter, build a registry value per test, use a fresh multiplexer instead of the shared default one, and let each test own a temporary directory. Once no test can observe another test's writes, order stops mattering and tests become safe to parallelize — which is the real payoff, since a suite that can run in parallel is usually several times faster. ## Keeping it from coming back Turn shuffling on in CI so a new ordering dependency fails on the pull request that introduces it, not months later. Keep the printed seed in the job output so a red build is reproducible. And treat "this test needs to run first" in a review the way you would treat a global variable, because that is what it is. ## What to tell the person triaging If the job is green locally and red in CI, the first question is what differs: a random shuffle seed, a cold test cache, a different value of `GOMAXPROCS`, or parallelism settings. Order dependence explains all four, and the reproduction recipe above turns an intermittent red into a deterministic one within a few minutes.

  • Why does the same failure never appear across two different packages?
    Because go test builds and runs a separate test binary per package. Package-level state is per process, so a test in package A cannot observe what a test in package B left behind. That is also why splitting a package can appear to fix the problem while leaving the real dependency intact inside one of the halves.
  • Why is -count=1 part of the reproduction command?
    go test caches results for identical runs and will replay a previous success without executing anything. -count=1 forces the tests to run. Without it you can spend a long time convinced that a change fixed something when the binary never ran at all.
  • The suite also calls t.Parallel in places. How does that change the diagnosis?
    It turns an ordering bug into a data race: two tests can be mutating the same package variable simultaneously. Run the same reproduction under the race detector. It also changes cleanup, since a parallel test resumes after its parent's body returns, so restoration belongs in t.Cleanup rather than defer.
  • Would TestMain solve it?
    No. TestMain runs once for the whole package, before and after the entire set of tests, so it can set up process-wide fixtures but cannot isolate tests from each other. Using it to reset globals between tests is not possible; the reset has to happen per test, or the sharing has to go.

It is a shared workbench: every test tidies up only the tools it remembers using, and the failure appears when someone arrives after the one person who left a clamp on.

saying these in an interview costs you the question

  • It is flaky, just retry the job
  • Add a sleep and it stabilizes
  • Rename the tests so they run in the working order
  • The race detector will find it, it is a race
  • Each test function starts with fresh package state
  • Set -p 1 so nothing runs concurrently and move on