For 40 Go services on an automated bump pipeline, how do you set upgrade cadence and who may force a transitive bump?
answer
- size of the diff, not the frequency
- two schedules, not one
- small steps often, big steps deliberately
- the indirect block is where the surprise hides
- permissive on who, strict on recording it
basics
~20 sRun patch-only upgrades on a frequent schedule so diffs stay reviewable, take minor upgrades deliberately per service before a release, and let any engineer pin a transitive module upward under an advisory, provided the pin is recorded and revisited later.
solid answer
~50 sSplit the cadence by blast radius. A frequent `-u=patch` sweep produces small, semantically bounded diffs a service team can merge on green tests, and stops drift compounding into an unreviewable jump later. Full `-u` moves whole minor versions across the build list, so it belongs to a service's own release window with a human reading the `go.mod` diff — including the `// indirect` lines, where transitive movement is recorded. Forcing one transitive module upward is a different act: it is cheap, often the only lever when upstream has not released, and should not need a committee. Any engineer may do it under an advisory, provided the pin records the advisory and is revisited once the graph requires that version anyway. Decide in advance who breaks the tie when security's clock and a release window disagree.
go deeper
Know that not all upgrades are equal: a patch bump is a small, routine change, while a minor bump deserves someone to read the diff and the release notes.
Be ready to explain why patch-only and full upgrades belong on different schedules, and what a reviewer should actually look at in the go.mod diff.
Argue the operational side: how a bump pipeline keeps reviews meaningful, how a forced pin is verified and tested, and how you avoid closing an advisory by merging an unreviewed fleet-wide upgrade.
Own the policy and the escalation: who breaks the tie between an advisory deadline and a release window, what evidence justifies an exception, and what metric would make you change the cadence.
## What the decision actually is The question is not "should we upgrade" — everyone agrees on that — but **how much change per pull request, how often, and who is allowed to bypass the schedule**. Those three settings interact, and getting them wrong is expensive in opposite directions: too slow and every upgrade is a migration; too fast and 40 repos produce a stream of PRs nobody reads, which is worse than not upgrading because it manufactures a rubber stamp. ## Cadence by blast radius **Patch sweeps, frequently.** A scheduled `-u=patch` run per module produces the smallest honest upgrade: bug fixes within the minor version already selected. The diff is short, the semantic promise is narrow, and green tests are usually sufficient evidence. Run it often — the value is that drift never accumulates, so no single PR is ever large enough to be scary. **Minor upgrades, deliberately.** `go get -u ./...` walks the whole set of modules behind your imports and can move dozens across minor versions at once. That is a real change, and its natural home is a service's own pre-release window where a human reads the diff, the changelogs for the moves that matter, and the tests run with time to react. Doing this on the same schedule as the patch sweep is the classic mistake: it makes every week's PR indistinguishable from every other week's, and the review degrades to a merge button. **Major versions, never automatically.** A new major version is a different module path and an import edit; it is project work, tracked as such. ## The review problem the pipeline creates With graph pruning, `go.mod` carries an explicit requirement for every module providing a transitively imported package, so one direct upgrade rewrites several `// indirect` lines. A reviewer who reads only the first require block sees a tidy one-line change and misses the transitive movement underneath. If you run a bump pipeline, make the PR body say what actually moved — the before/after of the whole build list, not just the direct requires — because that is the artefact people will decide on. `go list -m -u all` before the run gives you the survey to include. ## Who may force a transitive bump Be permissive here, and pay for it with bookkeeping rather than with gatekeeping. Raising one transitive module ahead of what its consumer requires is a one-line change, it is the only lever available when upstream has not shipped, and requiring approval for it just means the advisory sits open. So: any engineer may do it. In exchange, three obligations. Record the advisory that motivated the pin in the PR, so the line is not a mystery in six months. Take the **smallest** version that carries the fix, not the newest, so a security change does not smuggle in a behavioural one. And put the pin on a list to be revisited, because once the intervening library requires that version on its own the explicit line is redundant and should go with the next tidy. ## The tie you should break in advance The genuine organisational conflict is between an advisory's deadline and a service's release discipline. Security owns the clock; the service team owns the blast radius and the on-call consequences. Decide *before* an incident who breaks that tie and on what evidence. A defensible answer: the service team chooses **how** (a targeted pin, an upgrade of the direct dependency, or in the rare case a temporary mitigation in code), security owns **by when**, and an extension needs an explicit exception with an owner and a date rather than silence. What you want to avoid is the failure mode where the deadline is met by merging a fleet-wide `-u` nobody reviewed — the advisory is closed and the risk has simply changed shape. ## What would change my mind This policy is calibrated for services you deploy yourself. For a **library** other teams import, the calculus inverts: raising a floor in your `go.mod` propagates to every consumer, so a transitive pin there is a decision with an external audience and belongs in the release notes. And for a fleet where tests are weak, faster cadence buys less than it costs — the honest first move is the test suite, not the schedule, because the whole argument for frequent small upgrades rests on green tests meaning something.
- Why not simply run the full upgrade weekly and trust the test suite?Because a weekly stream of large diffs trains reviewers to approve without reading, and tests only cover behaviour you thought to assert. Minor releases change defaults, logging, error text and performance in ways a service suite rarely catches. Keeping the scheduled sweep small preserves the meaning of a green review; the large upgrade earns a human because it deserves one.
- How would you measure whether the cadence is working?Track the age of the oldest selected version per service and the time from an advisory's publication to the fix being deployed — those two say whether drift and response are under control. Also watch the merge-without-comment rate on bump PRs: if it approaches 100%, the review is ceremonial and the diffs are either too large or too frequent.
- What changes if the module being upgraded is a library your other teams import?A requirement in a library's go.mod becomes a floor for every consumer, so raising one is an externally visible decision, not a private one. Take the smallest version that solves the problem, say so in the release notes, and avoid pinning transitive modules on a consumer's behalf unless a security fix genuinely requires it.
saying these in an interview costs you the question
- One cadence for every kind of upgrade
- Requires central approval for a one-line security pin
- Treats a green pipeline as a substitute for reading the diff
- Ignores indirect requirement movement in review
- Leaves forced pins in place with no record of why