skip to content

How would you make CI prove that committed generated Go files match `go generate ./...` output?

level: seniorimportance: should knowfreq 42%

answer

  1. regenerate, then ask version control
  2. a diff is the failure message
  3. new files are not tracked yet
  4. both machines need the same versions
  5. stage everything before comparing

basics

~20 s

On a clean checkout, run go generate over the module, stage everything, then run git diff with its exit-code flag. A non-empty diff means the committed generated files are stale, and the printed diff is the fix.

solid answer

~50 s

The check is three lines: clean checkout, `go generate ./...`, then `git add -A && git diff --cached --exit-code`. Staging first matters, because a plain `git diff --exit-code` cannot see files a generator newly created — untracked files are invisible to it, so a generator that adds a file would pass. Printing the diff into the job log makes the failure self-explanatory. For the check to be trustworthy rather than flaky, three things must hold: the generators are pinned so both environments run the same build, their output is deterministic (no timestamps, no absolute paths, no ordering that varies run to run), and the Go release is pinned too, since anything that emits formatted Go depends on the formatter's version. When it goes red on a pull request that touched no generated file, the cause is almost always one of those versions moving, not the author's change.

code

text · 11 lines
text
$ go generate ./...
$ git add -A
$ git diff --cached --exit-code
diff --git a/internal/api/codes.go b/internal/api/codes.go
--- a/internal/api/codes.go
+++ b/internal/api/codes.go
@@
-case StatusRetrying:
+case StatusRetrying, StatusPaused:
$ echo $?
1

go deeper

for a junior

Know why generated code is committed at all, and that a build can be green while a committed generated file is stale because the stale file still compiles.

for a middle

Explain the command sequence and why staging precedes the diff: untracked files are invisible to a plain diff, so a generator that adds a file would slip through.

for a senior

Show you have run this gate in anger: determinism requirements, pinned generators and pinned Go release, triage order when it reddens with no relevant change, and why a bot that pushes the fix is the wrong answer.

for a principal

Decide whether the repository commits generated code at all, who absorbs a repo-wide regeneration when a version moves, and whether this gate blocks merges or reports until its flakiness is provably zero.

## What the check is for Generated Go code is usually committed. That is a deliberate choice: it keeps `go build` free of a code-generation step, it lets readers and reviewers see what actually compiles, and it means a consumer of the repository needs no generator at all. The cost of committing it is that the committed copy can fall out of step with its inputs. Someone edits the source of truth, forgets to regenerate, and the repository now contains code that no longer corresponds to anything. The build is perfectly green, because the stale file compiles. The fix is a check that regenerates and compares. ## The shape of it ``` go generate ./... git add -A git diff --cached --exit-code ``` `git diff --exit-code` exits 1 when there is a difference and prints it, so the diff becomes the failure message. Two details in that snippet carry most of the value: - **Run it on a clean checkout.** The comparison is only meaningful if the tree started identical to the commit under test. In CI that is free; locally it is worth saying so in the docs. - **Stage before diffing.** A plain `git diff --exit-code` compares tracked files against the index and is blind to **untracked** files. A generator that emits a brand-new file therefore produces no diff and the check passes while the repository has an uncommitted artifact. `git add -A` first (or a `git status --porcelain` emptiness test) closes that hole. A third hole is worth checking once, by hand: if the generated files are listed in `.gitignore`, the check is vacuous — nothing is ever tracked, so nothing ever differs. A gate that cannot fail is worse than no gate, because people believe it. ## Making it deterministic The check compares bytes, so anything non-deterministic in the generated output turns it into a flaky gate that people learn to re-run. The recurring offenders: - **Timestamps and hostnames** written into a header banner. Remove them; the commit already records when and by whom. - **Absolute paths** from the machine that generated the file. - **Version strings** for the generator itself in the banner — harmless until the generator is upgraded, at which point every generated file in the repository churns. - **Ordering** that is not fixed. A generator that walks a map and emits in iteration order produces a different file each run, because Go randomises map iteration order. Sort before emitting. ## Making it version-stable The failure this gate produces most often is not stale code at all: it is **version skew**. Two things must match between the laptop and the pipeline. First, the **generators**. If each environment runs a differently-built copy of the same tool, the outputs differ and the diff blames whoever pushed last. Declaring the tools in `go.mod` so they are built from a pinned source removes this, and it turns an upgrade into a reviewable commit that carries the regenerated files with it. Second, the **Go release**. Anything that emits Go source usually formats it with the toolchain's formatter, and formatter output is not frozen forever — Go 1.19's reformatting of doc comments is the well-known example of a release that legitimately changed the bytes of correctly formatted files. A pipeline one release ahead of the laptops will therefore regenerate different bytes from identical inputs. Pin the Go version the same way you pin the tools, and treat a Go upgrade as a change that may carry a repo-wide regeneration commit. ## Triaging a red run When the gate fails on a pull request that touched nothing generated, work in this order: read the diff in the log — is it whole files, or one reformatted comment block? Compare the Go version and the tool versions between the two environments. Reproduce with the pipeline's toolchain. If the change is real and repo-wide, land it as its own commit, separate from feature work, so that reviewers can see a mechanical change for what it is. ## Why not have CI commit the result instead A tempting alternative is to let the job regenerate and push the result back to the branch. It looks friendlier and it is usually a mistake: it hides the drift instead of reporting it, it races with the author's own pushes, it can retrigger itself, and a bot commit lands code that no human reviewed on a branch that may have required checks. Failing with the diff in the log keeps the regeneration where it belongs — in the author's commit, under review, next to the change that caused it. ## The one-command rule Whatever the pipeline runs, an engineer must be able to run the identical command locally and see the identical result. A generation gate that can only be reproduced by pushing is the most expensive kind of check there is: every iteration costs a full pipeline round trip, and people respond by regenerating blindly until it goes green.

  • Why is `git diff --exit-code` on its own not enough after regenerating?
    It compares tracked files against the index and ignores untracked ones. A generator that emits a brand-new file therefore leaves an empty diff and the check passes while an unreviewed artifact sits in the tree. Staging with `git add -A` first, or testing that `git status --porcelain` is empty, makes new files count. It is also worth confirming the generated paths are not gitignored, which would make the gate incapable of ever failing.
  • The generate check goes red on a pull request that touched no generated file. How do you triage it?
    Assume version skew before assuming a bad change. Read the diff: whole files rewritten or one comment block reflowed points at a Go release difference, since formatter output changes across releases; scattered content changes point at a generator version difference. Compare the Go version and the pinned tool versions in both environments and reproduce with the pipeline's toolchain. If the change is genuine and repo-wide, land the regeneration as its own commit.
  • Why not have the job regenerate and push the result back to the branch?
    It hides the drift instead of surfacing it, races with the author's pushes, and can retrigger the pipeline against itself. It also lands code nobody reviewed onto a branch that may have required checks. A failing check with the diff printed puts the regeneration in the author's commit, beside the change that made it necessary, where a reviewer sees both.
  • What makes this check flaky rather than merely red?
    Non-deterministic output. Timestamps or hostnames in a banner, absolute paths from the generating machine, the generator's own version string, and emission driven by map iteration order all differ run to run — Go randomises map iteration deliberately, so a generator that walks a map without sorting produces a different file each time. Remove those and sort before emitting, or the team learns to re-run the job until it passes.

saying these in an interview costs you the question

  • Uses git diff --exit-code and never stages new files
  • Runs the check on a tree that was already dirty
  • Ignores that generated paths may be gitignored entirely
  • Blames the author when the Go version differs between environments
  • Lets the job push regenerated code back to the branch
  • Leaves a timestamp banner in generated output