A code generator's golden test flakes, and each go test -update run rewrites the golden differently. What is the likely cause?
answer
- same input, different bytes
- something in the generator has no fixed order
- Go randomises one very common loop
- sort the keys, not the output
- generate twice and compare to lock it in
basics
~20 sThe generator is almost certainly emitting in the order of a range over a Go map, and Go randomises that order deliberately. Sort the keys and emit in that fixed order, then assert determinism by generating twice and comparing.
solid answer
~50 sWhen the same input produces different bytes on different runs, the generator is reading order from something Go does not fix. The overwhelmingly common source is `for k := range m` over a map — Go randomises map iteration order on purpose, so a generator that walks a `map[string]FieldSpec` emits its fields in a fresh order each run. Other candidates are a `time.Now()` stamp in a header comment, an absolute path or a hostname baked into the output, and unsynchronised goroutines appending to a shared buffer. The fix is to make the generator deterministic rather than to make the test tolerant: sort the keys once and range over the sorted slice. Then lock it in with a test that runs the generator twice on the same input and compares the two outputs — that check catches the next nondeterminism at the source, where a golden diff only shows the symptom.
code
go · 9 lines// Before: field order follows Go's randomised map iteration.
for name, f := range schema.Fields {
emitField(&buf, name, f)
}
// After: one fixed order, every run, on every machine.
for _, name := range slices.Sorted(maps.Keys(schema.Fields)) {
emitField(&buf, name, schema.Fields[name])
}go deeper
Know that Go gives no order guarantee when you range over a map, so anything built by walking a map comes out arranged differently each run.
Explain the fix and why it belongs in the generator: collect the keys, sort them, emit in that order, so the output is a function of the input alone.
Name the whole family of causes - map range order, a time.Now stamp, an absolute path, unsynchronised goroutines - and add a two-runs-must-agree assertion that fails at the source.
Argue that reproducible output is a contract the tool owes every consuming repository, not a test convenience, and resist weakening the comparison to make a flaky fixture quiet.
## The symptom and what it means A golden test compares generated bytes against a checked-in file. If it goes red for a change that should not have touched the output, and a rerun with `-update` produces yet another arrangement of the same content, the input to the comparison is not a pure function of the test's input. The bug is in the generator, not in the fixture and not in the test. ## The first suspect: map iteration order Go randomises the iteration order of `for k, v := range m` over a map, and does so **deliberately** — the runtime starts each range at a random bucket so that no program can accidentally depend on an order the implementation never promised. It is not "random on some builds"; it is randomised on every range statement, in every run. A schema-driven code generator is the classic victim, because a schema is naturally modelled as `map[string]FieldSpec`: ```go for name, f := range schema.Fields { emitField(&buf, name, f) } ``` That emits the same set of struct fields in a different order every run. The generated code compiles and behaves identically each time, which is why nobody noticed until a byte-for-byte fixture started disagreeing with itself. The fix is one line of intent: choose an order. ```go for _, name := range slices.Sorted(maps.Keys(schema.Fields)) { emitField(&buf, name, schema.Fields[name]) } ``` On older toolchains the same thing is spelled by collecting the keys into a slice and calling `sort.Strings`. Either way, the order is now a property of the generator that a reviewer can see, rather than an accident of the runtime. One useful counter-example to keep straight: printing a map with `fmt.Printf("%v", m)` is **not** a source of nondeterminism. The `fmt` package sorts map keys before printing precisely so that output is reproducible. It is `range` that randomises, not formatting. ## The other suspects - **Timestamps.** A `// Code generated at <time.Now()>` banner makes every run differ. Generated-code headers should carry the `// Code generated ... DO NOT EDIT.` line and the source of truth, never a clock reading. - **Absolute paths.** Embedding the input file's absolute path, the module cache path, or a hostname makes the golden machine-specific: green locally, red in CI, different on every developer's laptop. - **Concurrency.** Fanning work out to goroutines and appending results as they finish reorders the output by scheduling. If the generator parallelises, it must reassemble results into a fixed index order. - **Randomness.** Any use of the global functions in `math/rand` gives a different sequence per process in modern Go; anything derived from it belongs behind an injected seed or not in generated output at all. - **Set-shaped intermediates.** A `map[string]struct{}` used as a set has the same problem as any other map the moment its contents are emitted. ## Why you fix the generator and not the test The tempting shortcuts are to sort the two byte slices before comparing, to compare only line sets, or to strip the offending lines with a regexp. All of them make the test weaker in a way that is invisible later: after the change, a genuine reordering bug in the generator passes. Worse, the *product* is still nondeterministic — two runs of the tool over the same schema produce diffs in the consuming repository, which means every downstream pull request carries phantom changes and no reviewer can tell signal from noise. Reproducible output is a property the generator owes its users; the golden test merely noticed it was missing. ## Lock it in Add an assertion the golden file cannot make: generate twice in one test and compare the two results to each other. ```go first := Generate(schema) second := Generate(schema) if !bytes.Equal(first, second) { t.Fatal("generator output is not reproducible for identical input") } ``` This is cheap, it fails at the source rather than at the fixture, and — unlike the golden comparison — an `-update` rerun cannot make it pass. Two runs in one process will not catch everything (a per-process random seed stays constant within the process), so pair it with the CI run, which is a fresh process on a different machine. ## Finally, normalise the formatting While you are in there: run generated Go source through `go/format`'s `Source` before writing or comparing. It removes whitespace-only churn from the diffs, so the golden's diff shows semantic changes only. It does not help with ordering — that is the generator's job — but it removes the second most common source of noisy golden diffs.
- Why not just sort both byte slices before comparing, so the order stops mattering?Because it hides a real defect in the product. Users of the generator get a differently ordered file on every run, so every downstream pull request contains phantom diffs. A sorted comparison also stops the test from catching a genuine ordering regression later. Fix the order in the generator and keep the comparison exact.
- Is printing a map with fmt also a source of nondeterministic output?No. The `fmt` package sorts map keys before printing a map, specifically so that formatted output is reproducible. The randomisation lives in `range` over a map, not in formatting. That distinction matters when hunting the cause: a generator that formats a whole map with `%v` is fine, one that ranges over it to emit lines is not.
- What else, besides map ordering, most often makes generated output differ between machines?A timestamp or hostname in the header banner, and any absolute path — the input file's full path or a module cache path — leaking into the output. Both are green locally and red in CI, or different for every developer. Generated headers should carry the DO NOT EDIT line and the logical source of truth, nothing environment-specific.
Photographing a bookshelf every morning is a fine way to spot a missing book — but only if nobody reshuffles the shelf overnight.
saying these in an interview costs you the question
- Blames CI or a flaky filesystem
- Sorts or normalises the bytes before comparing instead
- Thinks map order is stable within one program run
- Believes fmt printing a map is nondeterministic
- Regenerates the golden until CI happens to agree