What does `go test -shuffle=on` randomise, and how do you replay the order that failed?
answer
- three settings, not two
- one binary per package is the scope
- subtests keep the order their parent chose
- the run prints something you must copy back
- a fixed number makes it deterministic again
basics
~20 sIt randomises the execution order of the top-level tests and benchmarks inside each test binary, which exposes tests that secretly depend on one another. The seed used is printed in the output; pass it back as -shuffle=SEED to replay that exact order.
solid answer
~50 s`-shuffle` takes `off` (the default), `on`, or an integer seed. With `on`, the go command seeds a randomiser from the clock and permutes the order of the top-level `Test` and `Benchmark` functions within each test binary; it does not reorder subtests relative to their siblings, and it does not change which packages run. The seed is reported in the output as a `-test.shuffle` line so the run is reproducible: re-run with `go test -shuffle=1725209476433829000 ./pkg` and you get the same permutation. What it finds is order dependence — a test that only passes because an earlier one populated a package-level variable, a global registry, a temp directory or a fake clock. A failure that appears only under shuffling is a real bug in the tests, not flaky infrastructure, and the seed is what makes it debuggable instead of a rumour.
code
text · 9 lines$ go test -shuffle=on ./internal/scheduler
-test.shuffle 1725209476433829000
--- FAIL: TestResultsMatchExpectations (0.00s)
FAIL
$ go test -shuffle=1725209476433829000 ./internal/scheduler
-test.shuffle 1725209476433829000
--- FAIL: TestResultsMatchExpectations (0.00s)
FAILgo deeper
Know that -shuffle=on exists, that it randomises the order tests run in within a package, and that the printed seed lets you run that same order again. Off is the default.
Be ready to say precisely what is permuted — top-level tests and benchmarks in one test binary, not subtests, not packages — and what kind of bug that catches: tests sharing package-level state.
Use it as a discriminator during triage. A failure reproducible at a fixed seed is coupling between tests; a failure that ignores the seed but needs -count is timing inside one test, and the two get completely different fixes.
Decide whether the suite runs shuffled by default and what the team does with a shuffle failure. Enabling it is cheap; the real commitment is the rule that such a failure is never retried away, and that logs are retained whole so the seed survives.
## The flag `go test -shuffle` accepts three forms: - `-shuffle=off` — the default. Tests run in the order they appear in the source files, files in the order the go command lists them. - `-shuffle=on` — a seed is chosen from the system clock and the order is permuted. - `-shuffle=N` — the integer `N` is used as the seed directly. In both randomising forms the seed is reported so the run can be reproduced. It appears in the test output on its own line, as `-test.shuffle` followed by the seed value. Copy that number back onto the command line — `go test -shuffle=1725209476433829000 -run 'TestA|TestB' ./internal/scheduler` — and you get the identical permutation. Without that step a shuffle failure is unactionable: you know something is order-dependent but you cannot get back to the order that showed it. ## What actually gets permuted The unit of shuffling is the **top-level** `TestXxx` and `BenchmarkXxx` functions registered in one test binary — that is, one package. Two things it deliberately does not do: - **It does not reorder subtests.** Subtests run in the order the parent calls `t.Run`, because the parent's own code decides that; a shuffle cannot reach inside your loop over a table. - **It does not shuffle across packages.** Each package is its own binary, and the go command already runs those independently. Package-level ordering is not something your tests may depend on in the first place. So the property `-shuffle` is testing is narrow and precise: *does any top-level test in this package depend on another top-level test in the same package having run first?* ## Why order dependence happens in Go specifically A Go test binary is one process per package, and everything at package scope is shared by every test in it for the whole run. The usual culprits: - a package-level variable a test mutates and a later one reads; - a registry or a `sync.Map` that one test populates via an `init`-like helper; - a shared temp directory or fixture file written by one test and read by the next; - a fake clock, a seeded random source, or a counter that one test advances; - an environment variable set by one test and never restored; - a table of expectations built once and consumed destructively. None of these fail in source order, which is exactly why they survive review. They fail the first time somebody adds a test above them, deletes one, or turns shuffling on. ## How it fits into flake triage When a package goes red intermittently in CI, there are two very different shapes of cause, and shuffling separates them cheaply: 1. **Order dependence between tests.** Reproduces under `-shuffle=on` at some seed and is then *deterministic* at that seed. Fix the coupling: give each test its own state, use `t.TempDir()` and `t.Setenv` so cleanup is automatic, stop mutating package-level variables. 2. **Timing or scheduling nondeterminism inside one test.** Does not care about order; it needs repetition (`-count=N`) and often a different parallelism level to show up, and it is never deterministic at a fixed seed. That second, seed-independent behaviour is the useful signal from a negative result. If a hundred shuffled runs at a hundred seeds are green but `-count=100` at one seed is red, you have learned that the problem lives inside a single test, not between tests, and you can stop reading the rest of the file. ## Operating it Turning `-shuffle=on` on permanently in CI is a reasonable, common policy — it converts a class of latent coupling into a loud failure early. The cost is that a failure now arrives with a seed you must not lose: capture the whole test log, not just the last twenty lines, or the seed is gone and the failure becomes irreproducible in exactly the way the flag was meant to prevent. Teams that enable it usually pair it with the rule that a shuffle failure is triaged as a real defect in the test suite rather than retried, because retrying at a fresh seed will usually pass and teach nobody anything.
- A test fails only at one shuffle seed. Is that a flaky test or a broken one?Broken. At a fixed seed the order is fixed, so the failure is deterministic — it is telling you one test depends on another having run first. Retrying at a new seed hides it; the fix is to remove the shared package-level state the two tests are passing between them.
- Why does shuffling not reorder subtests inside a table-driven test?Subtests exist only because the parent calls t.Run in its own loop, so their order is ordinary program control flow, not a list the test binary is free to permute. Only the top-level Test and Benchmark functions registered in the binary are shuffled.
- What should CI do with the seed when a shuffled run fails?Preserve it in the retained log or the failure annotation, because it is the only way back to that order. Truncating the output to the last few lines loses the -test.shuffle line and turns a fully reproducible defect into an anecdote nobody can chase.
saying these in an interview costs you the question
- Thinks -shuffle also randomises subtests
- Believes it reorders packages across the module
- Retries at a new seed instead of recording the failing one
- Calls a seed-reproducible failure flaky infrastructure
- Assumes shuffling helps with timing-based nondeterminism