A Go service untouched for years is rebuilt on a modern toolchain and one test now fails: how do you triage it?
answer
- suspect your own code first
- release notes have a compatibility section
- was it ever promised at all?
- bisect across toolchain versions
- pin to buy time, not to fix
basics
~20 sAssume the code relied on something Go never promised. Read the compatibility notes for each skipped release and classify the failure: unspecified behaviour, a fixed bug, a newly run vet check, or changed tool output. Then fix the dependency, not the toolchain.
solid answer
~50 sMy prior is that the toolchain is right and the code was depending on something never promised, because the Go 1 promise is narrow but well kept. I would read the compatibility section of the release notes for every version we skipped and match them against the single failing test. The usual classes: an assertion on something unspecified, such as map iteration order or a standard-library error string; a latent bug that a compiler or runtime fix exposed — Go 1.25 fixed a Go 1.21 compiler bug that delayed nil-pointer checks, so code that silently "worked" began panicking; a vet check that `go test` now runs by default, which is a tool change and outside the promise; or a `gofmt` output change tripping a CI diff. Then I install a couple of intermediate toolchain versions to bisect which release flipped it, and fix the dependency rather than pinning forever.
go deeper
Know the first move: read the release notes for the versions you jumped over before you debug anything, and do not assume the upgrade is at fault.
Explain the classification — unspecified behaviour, a bug fix exposing a latent bug, a tool change, a tightened behaviour — and that each one lands the fix in a different place.
Show the method: bisect across intermediate toolchain versions, name the change from that release's notes, and repair the dependency instead of freezing the version.
Own the upgrade posture — a cadence that never skips years, CI that tolerates tool-output changes, and a standard that forbids depending on behaviour the language never specified.
## Start from the right prior A service that has not been rebuilt for years and now fails one test after a toolchain jump is, overwhelmingly, a service that depended on something the Go 1 compatibility promise never covered. The promise binds the language specification and the documented standard-library API; everything else — unspecified behaviour, tool output, performance, and the correction of genuine bugs — was always free to move. Starting with "Go broke us" wastes the afternoon and usually ends in an embarrassing retraction, so start with "which unpromised thing did we lean on?" ## Read before you debug Each Go release ships notes with a compatibility section listing exactly this class of change. Read the notes for **every** version you skipped, not just the one you landed on, and read them against the specific symptom: one failing test is a narrow target and the notes are a short list. This is much faster than stepping through code, because the change you are hunting was documented deliberately. ## The four classes, and the different fix each needs **1. You asserted on unspecified behaviour.** A test comparing an error's message text, depending on map iteration order, on which ready `select` case fires, on goroutine interleaving, or on the order of equal elements from a non-stable sort. The fix is in the test: match errors with `errors.Is` or `errors.As`, sort before comparing, remove the ordering assumption. Nothing about the service is actually broken. **2. A bug fix exposed a latent bug in your code.** The promise explicitly allows fixing bugs in the compiler, runtime and libraries. Go 1.25 fixed a Go 1.21 compiler bug that had been delaying nil-pointer checks, which means some programs that appeared to work began panicking at a dereference they always contained. The fix is in your code, and the failing test just did you a favour. **3. A tool changed.** Tools were never in scope. `go test` runs vet analysers by default and gains more of them over time — as of Go 1.27 it runs the `stdversion` check — so a build can fail with no source change. `gofmt`'s output has changed, notably the doc-comment reformatting in Go 1.19, and that shows up as a CI diff rather than a test failure. The fix is to accept the tool's new output, or address the real issue it now reports. **4. A behaviour was tightened, often for security or correctness.** Standard-library packages do get stricter about malformed input. Here the fix is to stop producing or accepting the malformed thing. ## Bisect to name the change If the notes do not identify it, bisect. Install a handful of toolchain versions spanning the gap, run only the failing test with each, and narrow to the release where it flips. Then re-read that one release's notes; the change is almost always named there. Bisecting across releases is far cheaper than bisecting your own history, because your history did not change. ## Fix the dependency, do not freeze the toolchain Staying on the old toolchain is a legitimate way to buy a week, and nothing more. It costs you security fixes, it compounds — the next upgrade will have to cross an even larger gap with even more notes to read — and it leaves the actual defect in place. ## What to tell the architect The question behind the question is usually "is this language safe to build on for the next five years?" The honest answer is that the promise is narrow, written down and well honoured: upgrades are recompiles, and the failures that do occur cluster in a small, enumerable set of unpromised behaviours. The cost is a code standard rather than a migration budget — keyed struct literals, no assertions on error text or map order, CI that tolerates tool-output changes, and an upgrade cadence that never lets the gap grow to years. A team that pays that small tax rebuilds on a new toolchain in an afternoon; a team that skips it discovers all of it at once, exactly as this service just did.
- The failing assertion compares the text of an error returned by the standard library. Who is wrong?The test is. Error message wording is unspecified and free to change between releases, so asserting on it was never safe. Rewrite the assertion against the error value — `errors.Is` for sentinels, `errors.As` for a concrete type, or `errors.AsType` since Go 1.26 — and the test survives future upgrades.
- The release notes do not obviously explain the failure. How do you find which release changed it?Bisect on the toolchain rather than on your code. Install a few versions spanning the gap, run only the failing test under each, and narrow to the release where behaviour flips. Then re-read that single release's notes, where the change is almost always described — far faster than stepping through unchanged source.
- Is pinning the old toolchain an acceptable resolution?Only as a scheduling device. Pinning forfeits security fixes and compounds the problem, because the next attempt must cross an even wider gap with more notes to read and more latent assumptions to unwind. Pin to buy a sprint, file the ticket, and fix the code that depended on something unpromised.
- The architect asks whether Go is safe to depend on for the next five years. What do you say?That the promise is narrow, written down and well honoured, so upgrades are recompiles rather than projects. The residual risk is entirely in code that leans on unpromised behaviour, which a small code standard removes: keyed literals, no assertions on error text or map order, CI tolerant of tool-output changes, and an upgrade cadence measured in months.
saying these in an interview costs you the question
- Declares Go broke compatibility before checking what was promised
- Pins the old toolchain permanently and calls it resolved
- Skips the release notes and starts debugging blind
- Treats a test asserting error text as a valid regression
- Jumps years of releases at once with no way to bisect