skip to content

A monorepo enforces a single version of a widely-used internal logging library across 300 consuming projects. The library's maintainers want to ship a breaking change to its initialization API. What does a safe 'lockstep upgrade' process for this look like in practice, and what makes it different from just bumping the version and letting the build fail until people fix it?

level: seniorimportance: must knowfreq 60%

answer

  1. additive-then-migrate-then-remove
  2. codemod rewrites call sites
  3. never a global red build
  4. LSC tooling / large-scale-change process

basics

~20 s

A safe lockstep upgrade updates the library AND every one of the 300 callers together in a controlled, often automated way (codemods, staged batches), instead of just breaking everyone's build at once and hoping they fix it themselves.

solid answer

~50 s

A lockstep upgrade treats "update the library" and "update every consumer" as one coordinated project, not two separate events. In practice that means: first, ship the new initialization API alongside the old one (so nothing breaks yet) and mark the old path deprecated; second, use tooling — often an automated codemod, sometimes assisted large-scale-change infrastructure — to mechanically rewrite call sites across all 300 consumers to the new API, landing the changes in batches or one big changelist gated by full-graph CI; third, once every consumer is confirmed migrated, remove the old API in a final, separate commit. This differs sharply from "bump the version and let the build fail" because that approach turns 300 teams' CI red simultaneously with no guidance on the fix, creates a scramble under pressure, and — in a single-version-policy repo — can block unrelated work for every team behind that build until it's resolved.

go deeper

for a junior

Should recognize that just flipping the library to a breaking version and letting everyone's build fail is risky and disruptive, even without knowing the proper staged process.

for a middle

Should be able to describe the coexistence idea — old and new API available together for a while — as safer than an instant breaking cutover.

for a senior

Should lay out the full additive-migrate-remove sequence, explain why it avoids a globally red build in a single-version-policy repo, and reason about when the overhead is and isn't worth it relative to blast radius.

for a principal

Should discuss the organizational tooling this requires at scale (codemods, batched landing gated by affected-graph CI, owner review of automated changes) and be able to cite or describe a real large-scale-change process, plus articulate the calendar-time-vs-blast-radius trade-off as a deliberate engineering decision, not just 'best practice.'

## What lockstep means 'Lockstep' describes an upgrade discipline where the library and all of its consumers are treated as advancing together, in step, rather than the library racing ahead and consumers catching up whenever they get around to it. In a monorepo enforcing a single-version policy, this isn't merely a nicety — it's close to a structural requirement, because the moment the library's version changes for anyone, it has effectively changed for everyone building against the same resolved graph. If the change is breaking, an unmanaged version bump means the maintainer flips one flag and 300 previously-green builds turn red at once. That's not really an upgrade process; it's an **outage generator with extra steps**. ## The three phases A well-run lockstep upgrade for a breaking API change typically goes through three phases, and the ordering is the whole point. 1. **Phase one is additive:** the new initialization API is introduced alongside the old one, so the library now supports both simultaneously and nothing currently building against it breaks. This might mean adding a new constructor or builder method while leaving the old one in place but marked deprecated (with a compiler warning, a lint rule, or both). 2. **Phase two is the migration itself:** every one of the 300 consumers needs its call sites rewritten from the old API to the new one. Doing this by hand, 300 times, by 300 different owning teams on their own schedule, is exactly the polyrepo-style drift the monorepo is trying to avoid — so mature setups instead use an automated **codemod**, a script that parses each call site (via an AST-aware tool, not naive text substitution) and mechanically rewrites it to the new form. This can be landed as one very large changelist, or more commonly as a batch of changelists grouped by project/ownership boundary, each gated by the full CI of the affected subgraph, sometimes with individual project owners required to approve the change to their own directory even though a bot or a central team authored it. 3. **Phase three**, only once telemetry or a repo-wide search confirms zero remaining call sites use the old API, is the actual breaking removal: delete the deprecated path in its own commit. At that point the 'breaking change' has already happened invisibly, absorbed entirely by phase two, and the final commit is nearly risk-free because nothing depends on what it's deleting. ## Why not just bump the version Why go to this trouble instead of just bumping the version? Because 'bump and let it fail' inverts **who bears the cost and when**. It turns a plannable, staged migration into an emergency for every downstream team simultaneously, right at the moment the maintainer chooses, with no guarantee those teams have capacity that week. In a single-version-policy monorepo it's worse than that: if the build system enforces one version repo-wide and the new version doesn't build for anyone yet, unrelated changes by unrelated teams can get blocked behind the broken build until someone fixes it, because CI for the whole graph (or a large affected slice of it) is now red. The lockstep, phased approach avoids ever having a repo-wide red state by construction — at every point in the migration, `HEAD` builds and passes, because the old and new APIs coexist until the very last, low-risk deletion step. ## The trade-off The trade-off is calendar time and process overhead versus speed. - A 'just bump it' approach is a single afternoon of the maintainer's time. - A proper lockstep migration for 300 consumers might take days or weeks of codemod development, staged landing, and confirmation, plus ongoing maintenance of two parallel APIs during the transition window. For a small, low-risk internal utility with three consumers, that overhead may not be worth it and a direct coordinated bump (update the library and the three call sites in one PR) is entirely reasonable — this is squarely the atomic cross-project change pattern rather than a phased migration, and it's the right tool when blast radius is small. Lockstep, phased migration earns its cost specifically when blast radius is large enough that a single atomic commit touching all consumers is impractical to review, land, and verify safely in one shot. ## The canonical example The canonical real-world example is Google's **large-scale-change (LSC)** process: breaking API changes to foundational internal libraries are almost never landed as a single flag flip. They go through exactly this additive-then-migrate-then-remove lifecycle, with tooling mechanically generating and landing the call-site rewrites across the codebase in batches, each batch small enough to be reviewed and tested independently, so that at no point does the build go globally red and no single team is surprised by someone else's API change breaking their build without warning.

  • What's the risk of skipping phase one (the additive/deprecation period) and going straight from 'old API only' to 'migrate everyone then delete'?
    Without the coexistence window, the library maintainer has to migrate all 300 consumers before they can ship anything, meaning the old API stays load-bearing (and can't be touched) for the entire migration duration, and any consumer added mid-migration might still get written against the soon-to-be-deleted API. The additive phase decouples 'the new capability exists' from 'everyone has moved to it,' which is what makes a staged, low-pressure migration possible at all.
  • Why would a monorepo require individual project owners to approve an automated codemod's changes to their own directory, rather than letting the codemod land everywhere unilaterally?
    Because a mechanical rewrite, however well-tested, can still produce a semantically odd result in an edge case the codemod author didn't anticipate — an owner review catches those before they land, and it also keeps ownership and accountability intact so teams aren't surprised by changes to code they're responsible for. It's a safety and trust mechanism, not just a formality.
  • How does a team know it's actually safe to move to phase three (deleting the old API) rather than just assuming the migration is done?
    They typically do a repo-wide search/build-graph query for any remaining reference to the deprecated symbol, sometimes backed by runtime telemetry (e.g., a deprecation warning logged whenever the old path is actually invoked) to catch dynamically-constructed call sites that static search might miss. Only when that comes back empty across a full build of the affected graph is the deletion considered safe.

Like replacing a bridge's roadway while keeping traffic flowing — you build the new lane alongside the old one, shift traffic over lane by lane, and only tear out the old lane once nothing is using it, instead of closing the whole bridge at once and telling every driver to find their own detour.

saying these in an interview costs you the question

  • thinks bumping a version and fixing broken callers reactively is fine at any scale
  • no mention of a deprecation/coexistence window before removal
  • assumes one engineer manually editing 300 call sites is a reasonable plan
  • doesn't recognize that a single-version-policy repo can globally block on one broken upgrade

context