Two CI runners build the same commit with different Go toolchains — how do you diagnose that drift?
answer
- the installed compiler and the running compiler differ
- ask each machine two questions, not one
- the setting resolves from three places
- the binary itself records who built it
- the fix is a setting in the image, not a rebuild
basics
~20 sCollect go version and go env GOTOOLCHAIN from every builder, then go version -m on each binary. The setting resolves from the environment, the go env -w file, then the built-in default, so identical installs still switch differently.
solid answer
~50 sTreat it as a configuration inventory problem, not a build problem. On each runner run `go version` (which toolchain is installed), `go env GOTOOLCHAIN` (the effective setting after the environment, the `go env -w` file and the built-in default are resolved) and `go env GOMODCACHE`. Then run `go version -m` on each produced binary to see which release actually stamped it — the artefact is the evidence, the runner is the hypothesis. The usual causes are: one image ships a repackaged Go whose default is not `auto`; a job sets `GOTOOLCHAIN` in its environment; one runner's Go is old enough to switch while the other's is already new enough, so only one of them upgrades; or a `go.work` in the checkout overrides the module's lines. Then remove the ambiguity: pin `GOTOOLCHAIN` explicitly in the image, so the answer stops depending on what was preinstalled.
code
text · 2 lines$ go version && go env GOTOOLCHAIN GOMODCACHE
$ go version -m ./svc | head -3go deeper
Know that two machines can build the same code with different Go versions, and that go version and go env GOTOOLCHAIN are the first two commands to run when versions are in doubt.
Explain why an older machine may switch upward while a newer one does not, and where the effective setting is resolved from before it takes effect.
Drive the diagnosis: gather an inventory across every builder, corroborate it with build information read out of the artefact, and close the incident with a pinned setting in the image rather than a manual fix.
Decide what the pipeline must assert about the toolchain before a release ships, and who is accountable when a compiler nobody chose produces a shipped binary.
## Why the drift is possible at all Because the go command may hand off to a different toolchain, "which Go built this" has two answers: the one installed on the machine, and the one that ran. Under `GOTOOLCHAIN=auto` a runner with an older installed Go will switch up to satisfy the repository, while a runner already at or above that version will not switch and will use its own — which may be a *different, newer* release than the one the first runner downloaded. Same commit, two compilers, no error anywhere. ## The inventory to collect Run the same three commands on every machine that builds, developer laptops included, and store the output next to the machine name: - `go version` — the toolchain that responded to the command. - `go env GOTOOLCHAIN` — the effective setting. This is the important one, and it must be `go env`, not the shell variable, because the value can come from three places. - `go env GOMODCACHE` — where downloaded toolchains land, which tells you whether a runner has been switching and caching releases at all. The expected shape of the result is boring uniformity. Any row that differs is your answer. ## Where the setting comes from Precedence: the process environment, then the go environment file written by `go env -w`, then the default compiled into the distribution. Three practical consequences: 1. A job or a shell profile that exports `GOTOOLCHAIN` beats everything, and is invisible unless you look at the job's environment. 2. An image built by someone who once ran `go env -w` carries that choice forever, in a file no Dockerfile mentions. 3. Official Go releases compile in `auto`, but a Go obtained through some other packaging may not. Two images that both report a plausible `go version` can still disagree on policy. ## Evidence from the artefact The machine can be reimaged; the binary cannot lie. `go version -m ./yourbinary` prints the build information embedded in a Go binary, including the Go version that produced it along with module and build settings. Comparing that across the two runners' artefacts confirms the drift independently of whatever the runners report today, and it is the check to bake into a release pipeline: assert that the shipped artefact was built by the release you intended. ## The repository side If the machines agree, look at the checkout. The `go` line sets the floor that triggers a switch; the `toolchain` line names the destination; and in a workspace, `go.work`'s own lines take precedence over the individual modules'. A `go.work` file present on one runner and not another — checked in on one branch, generated by a setup script on the other — will produce exactly this symptom. ## The failing-runner variant The same root cause has a louder form: one machine builds and another refuses. That is `GOTOOLCHAIN=local` (or a machine that cannot reach the module mirror to fetch the toolchain) meeting a repository whose `go` line has moved ahead. The error names the required version, the running version and the setting, so it diagnoses itself — the interesting part is that it is the *same* configuration drift, surfacing loudly on one machine and silently on the other. Preferring the loud form is a legitimate design choice. ## Fixing it for good The fix is to stop letting the answer depend on what happened to be preinstalled: - Set `GOTOOLCHAIN` explicitly in the build image — an exact version name for reproducibility, or `local` if the image is supposed to be the sole source of the compiler. - Make the image's installed Go and that setting one change: bumping the compiler should be a single reviewed diff. - Record the toolchain in the repository with a `toolchain` line if you want repositories, rather than images, to lead — but know that a pinned image overrides it. - Verify, don't assume: have the pipeline print `go version` and `go env GOTOOLCHAIN` into the build log, and check `go version -m` on the artefact before release. A one-line log statement turns this class of incident into a five-second diff. ## What good looks like in the interview Say that you would gather evidence from machines *and* artefacts, name the three-level precedence of the setting, and end with a policy change rather than a one-off fix. Blaming "the runner image" without saying how you would prove it is the weak version of this answer.
- Both runners report the same go version. Where do you look next?At the effective setting and at the checkout. `go env GOTOOLCHAIN` may differ because one image had a value written with `go env -w` or the job exports one. If those agree too, look for a `go.work` present on one runner only, since a workspace's own go and toolchain lines override the module's and can trigger a switch on just one machine.
- How do you prevent this recurring rather than just explaining today's build?Pin `GOTOOLCHAIN` in the build image to an exact version name so the compiler no longer depends on what was preinstalled, make the installed Go and that pin move in one reviewed change, and log `go version` plus `go env GOTOOLCHAIN` in every build. Then assert the expected version from `go version -m` on the artefact before a release ships.
- One runner fails outright instead of drifting. What is the likely cause?It is not switching. Either `GOTOOLCHAIN=local` is in force, so a repository whose go line moved ahead of the installed Go is an error, or the machine cannot fetch the toolchain it wants. The error text names the required version, the running version and the setting, which is enough to tell those apart quickly.
- Why is go version -m on the binary worth running when you already have the runners' output?Because it is evidence from the artefact rather than a claim about the machine. Runners get reimaged, environments change between the failing build and your investigation, and a cached build can outlive its builder. The build information embedded in the binary records the Go version that produced it, so it settles the question after the fact.
saying these in an interview costs you the question
- Assumes go version alone identifies the compiler that ran
- Reads the shell variable instead of go env GOTOOLCHAIN
- Never inspects the produced binary for build information
- Overlooks a go.work file overriding the module's lines
- Fixes one runner by hand and calls the incident closed