skip to content

A Go code generator walking a map of type descriptors emits a different file each run. How do you make its output byte-identical?

level: seniorimportance: should knowfreq 40%

answer

  1. the diff churns but nothing changed
  2. count the maps in the emitter
  3. one sort at the top is a partial fix
  4. order belongs in the parsed data
  5. generate twice, compare the bytes

basics

~20 s

Every place the emitter ranges a map is a source of randomised order. Sort the keys at each emission site, or keep the schema's order in slices and use maps only for lookup, then assert byte-identical output by generating twice in one test.

solid answer

~40 s

The generator is walking maps — the symbol table of type names, each type's field descriptors, the set of imports — and Go randomises every map range, so each run emits the same declarations in a different arrangement. The symptoms are a generated file whose diff churns when the schema did not change, and a golden test that is red only sometimes. The fix is not one sort at the end: audit every point where the emitter iterates a map and give each an explicit order, sorting the keys or, better, carrying the schema's own declaration order in slices and demoting maps to lookup tables. Then make determinism testable — generate twice in one test and require the bytes to be equal, which fails every run instead of one in ten.

code

go · 9 lines
go
// fields is map[string]FieldSpec; emit it in a fixed order, not range order
names := make([]string, 0, len(fields))
for name := range fields {
	names = append(names, name)
}
slices.Sort(names)
for _, name := range names {
	fmt.Fprintf(&buf, "\t%s %s `json:\"%s\"`\n", name, fields[name].GoType, fields[name].JSONName)
}

go deeper

for a junior

Know the first move: any generator that walks a map produces output in a randomised order, so the keys must be collected and sorted before anything is written to a file.

for a middle

Be able to enumerate the several maps hiding in an emitter — the symbol table, each type's fields, the import set — and explain why fixing only the outer loop leaves a rarer version of the same failure.

for a senior

Demonstrate the operating judgment: replace the repeated-runs guess with a test that generates twice and compares bytes, restructure the parsed schema so order lives in slices, and rule out the neighbouring nondeterminism sources such as timestamps and paths.

for a principal

Own it as a build policy. Byte-identical generated output is what makes checked-in generated code reviewable and lets CI fail when it is stale; decide that determinism is a property of the generator, not something each test works around.

## The failure as it is actually experienced Someone regenerates code, commits, and the diff shows fifty moved lines and no semantic change. Later, CI goes red on a golden-file test; the engineer reruns it and it passes. That combination — a churning diff plus an intermittently red pipeline — is the signature of map iteration leaking into output. ## Find every iteration site, not just the obvious one A generator usually holds several maps, and each one is a separate source of disorder: - the **symbol table**, `map[string]TypeSpec`, driving which types get emitted and in what order; - each type's **fields**, if they were parsed into a `map[string]FieldSpec` rather than a slice; - the **import set**, typically a `map[string]bool` or `map[string]struct{}` used to deduplicate; - any **enum or constant set** collected the same way; - a **cache of rendered fragments** later concatenated. Fixing only the outermost loop is the classic partial fix: the file's top-level order stabilises, the diff gets smaller, and the intermittent failure survives at lower frequency, which is worse than before because it now looks like flaky infrastructure. ## Two ways to impose the order, and they are not equal **Sort at the emission site.** Extract the keys, sort them, then emit in that order: ```go names := make([]string, 0, len(fields)) for name := range fields { names = append(names, name) } slices.Sort(names) for _, name := range names { emit(name, fields[name]) } ``` This is correct and local, but it must be repeated everywhere, and a future contributor who adds a sixth map will not know the rule. **Carry order in the data structure.** Parse the schema into `[]TypeSpec` and `[]FieldSpec` in source order and keep a `map[string]int` index only for lookups. Now the emitter cannot produce disorder because it never ranges an order-bearing map, and the generated struct's field order matches the schema, which reviewers find far more readable than alphabetical. This is the design fix; the sorts are the patch. Imports are the one place alphabetical order is genuinely right, since that is the conventional layout anyway. ## Make determinism a test, not a hope Rerunning a flaky test with a higher count raises your chance of catching the bug but proves nothing: with three entries in a map, one run in six may still match by luck, and a passing suite tells you only that today's dice were kind. A determinism test should fail on the first execution: ```go func TestGenerateIsByteIdentical(t *testing.T) { first := Generate(schema) for i := 1; i < 20; i++ { if got := Generate(schema); !bytes.Equal(got, first) { t.Fatalf("run %d produced different bytes than run 0", i) } } } ``` Because each `Generate` call executes fresh range statements, twenty in-process runs exercise twenty independent randomisations and the assertion catches disorder essentially every time. Keep a golden file alongside it, regenerated by a flag, so that a *deliberate* change to the output shows up as a reviewable diff. Running `go test -count=10` is still useful when you suspect nondeterminism you have not localised, but as a permanent guard it is the weaker instrument. ## Other nondeterminism to rule out while you are here Map order is the most common cause but not the only one. Timestamps or a build host name embedded in a header, absolute file paths that differ between a laptop and a CI runner, concurrent workers appending results to a shared slice in completion order, and iteration over a set built from a directory walk whose order the filesystem chose — all produce the same churning diff. Once the output is byte-identical, all of these become visible immediately rather than through a statistical haze. ## Why byte-identical output is worth the effort It makes the generated file reviewable: a diff means something changed. It makes the build cacheable and comparable across machines. It lets you check the generated code into version control and verify in CI that regenerating produces no diff, which is the cheapest possible guard against a stale generated file. And it removes an entire category of intermittent CI failure that costs far more in trust and re-runs than the sorts cost in code. ## What an interviewer is listening for The candidate should reach map iteration quickly, then — and this is the discriminator — should say *every* map, not just the top-level one. They should prefer restructuring the data over sprinkling sorts, should propose a test that fails deterministically rather than relying on repeated runs, and should mention the neighbouring sources of nondeterminism that the same discipline exposes.

  • Why is running the test with -count=10 a weak guard against this?
    It only raises the odds of catching disorder. A map with few entries can repeat the same arrangement by chance, so the suite may pass while the bug remains, and it tells you nothing about which map is at fault. Generating twice in one test and comparing bytes fails every time and points at the artefact directly.
  • Should the generated struct's fields be alphabetical or in schema order?
    Schema order, carried in a slice, is usually better: it matches what the author wrote, reads naturally in review, and keeps the wire layout predictable. Alphabetical is a determinism patch that also reorders a human-authored structure. Imports are the exception, where alphabetical is the conventional layout anyway.
  • Beyond map ordering, what else makes generated output differ between runs?
    Embedded timestamps, a hostname or user name in a header comment, absolute paths that differ between a laptop and CI, concurrent workers appending results in completion order, and iteration over entries collected from a filesystem walk. All produce the same churning diff, and all become obvious once map order is fixed.
  • How do you stop a regenerated file from silently drifting out of date?
    Check the generated file in and add a CI step that regenerates and fails if the working tree differs. That check is only meaningful once output is byte-identical — with map-order churn it would fail on every run and be disabled within a week.

saying these in an interview costs you the question

  • Sorts only the outermost map and calls the problem solved
  • Reruns the failing test until it passes and moves on
  • Treats intermittent CI red as an infrastructure problem
  • Suggests seeding or disabling map randomisation
  • Relies on -count as the permanent determinism guard
  • Keeps declaration order in a map and sorts at render time by habit