skip to content

A CI runner fails with a module checksum mismatch on a commit that built yesterday — how do you triage it?

level: seniorimportance: should knowfreq 38%

answer

  1. the reflex fix and the attack agree
  2. capture the two hashes before touching the cache
  3. diff this runner's settings against a healthy one
  4. one runner versus all runners narrows it
  5. a moved tag is benign but not automatic

basics

~20 s

Treat it as a security event until proven otherwise. Capture the module version and both hashes, diff the runner's GOPROXY, GOSUMDB, GOPRIVATE and GOFLAGS against a healthy runner, then reproduce with an empty module cache. Never retry it away.

solid answer

~60 s

The first decision is not technical: a checksum mismatch is the one error you must not retry away, because retrying and clearing caches are exactly what destroys the evidence. I capture the failing `module@version` and both hashes from the error, then dump that runner's `GOPROXY`, `GOSUMDB`, `GOPRIVATE`, `GONOPROXY` and `GOFLAGS` and diff them against a runner that succeeded — a surprising share of mismatches are one host fetching through a different path, such as a mirror that repacked the archive or a direct VCS fetch bypassing the proxy. Then I reproduce on a clean machine with an empty `GOMODCACHE` where the checksum database is reachable. The blast pattern narrows it fast: one runner failing points at that host's cache or config, every runner failing points upstream or at the mirror. Benign explanations exist — a maintainer moving a released tag, a hand-edited `go.sum` from a bad merge — but each of those still needs a decision recorded, not a workaround. Only after the cause is named do I change anything.

code

text · 6 lines
text
verifying example.internal/[email protected]: checksum mismatch
	downloaded: h1:9lD6Y...=
	go.sum:     h1:JQ0Zm...=

SECURITY ERROR
This download does NOT match an earlier download recorded in go.sum.

go deeper

for a junior

Know that a checksum mismatch means the fetched dependency does not match the recorded hash, and that the right reaction is to escalate rather than retry the job or edit go.sum.

for a middle

Be able to list the plausible causes — a moved upstream tag, a mirror serving repacked archives, a corrupted cache entry, an edited go.sum — and say which evidence separates them.

for a senior

Show the triage discipline: capture the hashes and the resolved settings before touching the cache, compare failing and healthy runners, reproduce clean, and read the blast radius as a signal.

for a principal

Frame it as a control that only works if the organisation never routes around it, and be ready to say what your platform does to make the safe response easier than clearing a cache.

## Why this error is different Most CI failures are worth a retry. This one is not, and saying so is half the answer. A checksum mismatch means the bytes the go command just fetched are not the bytes that were recorded as correct. Two of the three natural reflexes — retry the job, clear the module cache, add the path to an exemption variable — remove the evidence, and the third, editing `go.sum` to match, converts a detected substitution into an accepted one. The instinctive fix and the attacker's goal are the same action, which is what makes this a discipline problem rather than a debugging problem. ## Step 1: capture before you touch anything The error names the module path, the version, the hash that was downloaded and the hash that was expected. Save that verbatim, along with the job's commit and timestamp. Then run `go mod verify` on the runner while the cache is still in the state that failed, and capture its output next to a dump of the environment the fetch actually ran under: `GOPROXY`, `GOSUMDB`, `GOPRIVATE`, `GONOPROXY`, `GOFLAGS` and `GOMODCACHE`. That pair — the verify output beside the resolved settings — is the artefact everything else is argued from, and it is unrecoverable once someone wipes the cache. ## Step 2: is this runner different from the others? Compare that environment dump against a runner that built the same commit successfully. In a self-hosted fleet this finds the cause more often than any upstream theory does, because the mismatch usually means *this host fetched from somewhere else*: - The mirror served a repacked archive rather than the bytes it received, so the hash differs from the published one. - `GOPROXY` fell through to `direct` on this host and a VCS export produced a different archive than the proxy's canonical zip. - `GOFLAGS` carries something injected per-runner that changed resolution. - The host's cache holds a partially written or corrupted entry from an earlier interrupted download. The blast radius reads the same way. One runner failing and eleven succeeding is a host problem. Every runner failing on a commit that built yesterday points at the mirror or at upstream, because the only thing that changed is outside your fleet. ## Step 3: reproduce cleanly On a machine with an empty `GOMODCACHE` and the checksum database reachable, fetch the module version again. Three outcomes, three conclusions: - The clean fetch matches `go.sum` — the failing runner's copy or path is at fault, not the dependency. - The clean fetch matches the *downloaded* hash from the error and the database agrees with it — the `go.sum` line in your repository is wrong, typically hand-edited or mangled in a merge; find the commit that changed it. - The clean fetch matches `go.sum` but your mirror keeps serving something else — the mirror is the problem, and that is an incident. A fourth case deserves naming: if the module path falls inside a `GOPRIVATE` pattern there is no outside record to appeal to at all, so the only recourse is the origin repository itself. That is a good moment to notice how wide the exemption patterns have grown. ## Benign causes that still require a decision The most common genuinely benign cause is an upstream maintainer moving an already-published tag onto new commits. The correct response is not to accept the new hash quietly: read what changed, then either pin to a version that was never moved or accept the new bytes deliberately, with the reasoning recorded in the change that updates `go.sum`. A hand-edited or badly merged `go.sum` is likewise benign in origin and still needs the offending commit identified rather than the file regenerated. "It went away when we cleared the cache" is not a conclusion, and a fleet that has learned to respond to this error by clearing caches has effectively turned the check off. ## What good looks like at the end A named cause, an artefact trail (the error, the verify output, both environments), a change that fixes the cause rather than the symptom, and — if the cause was a mirror or a fetch path — a check that the same condition is not silently in effect on the other runners. If the answer turns out to be an exemption someone widened months ago to make a similar error go away, that finding is more valuable than the incident itself.

  • An engineer resolves the mismatch by adding the module's prefix to GOPRIVATE and the build goes green. What has actually happened?
    The check was removed for every module under that prefix, not just this one, and the mismatch that prompted it was never explained. Any future substitution under that prefix now passes silently, and the change usually outlives the incident because nobody revisits an exemption that made a build green. Treat the widened pattern as the finding and revert it once the real cause is named.
  • Eleven runners build the commit cleanly and one reports the mismatch. Where do you look first?
    At that host: its module cache and its resolved fetch settings. A per-host divergence — a stale or corrupted cache entry, a GOPROXY that fell through to direct, an injected GOFLAGS value, a different mirror endpoint — explains a single-host failure far better than any upstream theory, since the other eleven prove the dependency and go.sum are consistent.
  • How would you tell an upstream re-tag apart from a compromised mirror?
    Fetch the version on a clean machine where the checksum database is reachable and see who agrees with whom. If the database's record matches the newly downloaded bytes, the published version itself changed — a re-tag. If the database agrees with your go.sum while only the mirror serves different bytes, the mirror is substituting content, which is an incident rather than a dependency-hygiene problem.

saying these in an interview costs you the question

  • Retries the job or clears the cache as the first move
  • Edits go.sum to match the downloaded hash
  • Adds the path to GOPRIVATE to unblock the build
  • Sets GOSUMDB=off across the fleet during triage
  • Calls it flaky infrastructure without naming a cause