skip to content

You're responsible for a widely-consumed internal library published to hundreds of services inside a large engineering org. Design a policy that lets SemVer's MAJOR/MINOR/PATCH numbers actually stay trustworthy over years of contributions from dozens of different teams, given that no single person reviews every change. What mechanisms would you put in place, and what's the failure mode of skipping each one?

level: principalimportance: should knowfreq 35%

answer

  1. automated API-diffing converts self-reported compliance into a build-time fact
  2. explicit public-API boundary makes 'breaking' checkable, not a debate
  3. separate approval gate for flagged-breaking changes, not ordinary code review
  4. deprecation window gives consumers lead time before a MAJOR removal
  5. changelog generated from the same classification data, not hand-written after the fact

basics

~20 s

You need more than good intentions: a clear written definition of what's 'public API,' automated checks that catch accidental breaking changes before merge, a required human sign-off for anything flagged as breaking, and a changelog that's actually generated from real diffs. Skip any of these and the version number slowly stops meaning what it claims.

solid answer

~50 s

Trustworthy SemVer at org scale needs to move the compatibility check out of individual judgment and into process/tooling, because no single reviewer sees every change over years of multi-team contribution. Key mechanisms: (1) an explicit, documented boundary of what counts as the public API so 'breaking' has a checkable definition; (2) automated API-diffing in CI that flags a PR as breaking/additive/patch-level automatically rather than trusting the author's self-classification; (3) a required approval gate from an API-owner role specifically for anything flagged breaking, separate from normal code review; (4) a deprecation-window policy so removals are never a surprise MAJOR bump with zero lead time; (5) changelogs generated from the same classification, not hand-written after the fact. Skipping automated diffing reverts you to trusting every contributor's manual judgment; skipping the deprecation window turns every removal into an unannounced MAJOR bump; skipping changelog automation lets the human-written summary drift from what the version number actually signals.

go deeper

for a junior

Not expected to design this; should be able to recognize, when told the policy, why each piece (diff tool, deprecation window) matters.

for a middle

Should understand why manual review alone doesn't scale and be able to name at least one category of automated tooling that helps.

for a senior

Should be able to propose 2-3 of the concrete mechanisms unprompted and reason about their individual failure modes.

for a principal

Should design the full policy end to end, including ownership/governance, tooling investment trade-offs, and how to roll it out across an org with existing multi-team contribution habits without stalling velocity.

## Why one person's judgment stops scaling At small scale, SemVer compliance is a single maintainer's judgment call, exercised consistently because one person (or a tiny core team) reviews every change against a mental model of the whole public API. That model breaks down completely at organizational scale: once dozens of teams contribute to a shared internal library over years, no individual reviewer has full visibility into every consumer's dependency on every corner of the API, and 'did I just break something' becomes a question that can't be reliably answered by inspection alone. Making SemVer trustworthy at this scale means replacing individual judgment with process and tooling wherever possible, and reserving human judgment for the cases tooling genuinely can't resolve. ## The five mechanisms, and what skipping each costs 1. **The first mechanism is a written, checkable definition of the public API surface itself** — without this, 'breaking change' is undefined, because reviewers disagree about whether an internal helper class that some other team imported anyway counts. Concretely this means a convention (an internal/impl package namespace, explicit annotations, or a maintained allowlist of exported symbols) that both humans and tooling can check against, so the compatibility promise has a bounded, unambiguous scope rather than covering 'anything anyone might be depending on,' which is unbounded per Hyrum's Law and impossible to fully protect. 2. **The second, and highest-leverage, mechanism is automated API-diffing in CI**: a tool that compares the previous published API surface against the PR's resulting surface and classifies the change as breaking, additive, or internal-only, independent of what the author believes they did. Ecosystems have mature tooling for this — binary/source compatibility checkers for compiled languages, and schema-diff tools for network APIs like OpenAPI or protobuf — and the mechanism matters because it converts SemVer compliance from a self-reported claim into a build-time fact. Skipping this step means the org is back to trusting every individual contributor's classification, which historically fails the way the automated-upgrade-trust problem does more broadly: well-intentioned engineers mislabel changes, especially ones involving default values or validation tightening that don't look like 'API changes' from a casual diff view. 3. **The third mechanism is a distinct approval gate**, separate from ordinary code review, for anything the diff tool flags as breaking — typically owned by a small API-owner or architecture-review role rather than whichever team happens to be touching the code that week. This matters because a routine PR reviewer optimizing for 'does this code work' is not the same lens as 'should this ship as a MAJOR bump and who does it affect downstream' — the second question requires visibility across consumers that an individual code reviewer usually doesn't have, so it needs to be routed to whoever does, or to tooling that can enumerate known internal consumers via a reverse-dependency graph. 4. **The fourth mechanism is a deprecation-window policy**: mark a symbol deprecated with a machine-readable annotation and a target removal version/date, keep it functioning for a minimum supported period, and only then remove it in a MAJOR bump that consumers had advance warning about. Skipping this turns every removal into a same-release surprise — technically still a correctly-labeled MAJOR bump, but one that gives consumers zero lead time to migrate, which defeats much of the practical value of the compatibility promise even when the version number is technically accurate. 5. **The fifth mechanism is generating the changelog from the same classification data the diff tool produced**, rather than a hand-written summary assembled after the fact from memory of what changed. Hand-written changelogs drift — engineers forget to mention a default-value change because it didn't feel like 'a change' worth writing up, and the gap between the changelog's story and the version number's actual meaning grows over years until consumers stop trusting either. ## The trade-off The overarching trade-off across all five mechanisms is **upfront tooling/process investment versus long-run trust**: each one requires real setup cost that's easy to defer when the library is small and the org is small. ## What deferring all of it looks like The failure mode of deferring all of it isn't a single dramatic incident — it's **gradual erosion**: version numbers technically increment correctly by author intent, but enough misclassified changes accumulate that consumer teams stop trusting MINOR/PATCH bumps to be safe, start manually reviewing every dependency bump regardless of version number, and the entire point of having a SemVer contract — safe, low-friction automated upgrades — quietly stops paying off, even though nothing about the process ever visibly 'broke.'

  • What tooling category actually detects an accidental breaking change automatically, and what can't it catch?
    API/schema diffing tools reliably catch structural breaks - removed/renamed symbols, changed signatures, changed field types or required-ness. They generally cannot catch pure behavioral breaks with an unchanged signature, like a changed default value's runtime effect, which is why behavioral changes still need human judgment or targeted contract tests as a complement.
  • Why route breaking-change approval to a dedicated role instead of the normal code owners of that file?
    Because assessing 'is this breaking, and who does it affect' requires visibility across the consumer graph that a file's usual code owner typically doesn't have; a dedicated API-owner role or a reverse-dependency-aware tool can see the blast radius across the org, whereas a local reviewer can only see whether the code itself looks correct.
  • How long should a deprecation window realistically be for a library depended on by hundreds of internal services?
    There's no universal number, but it should be driven by realistic consumer migration capacity, not maintainer convenience - commonly at least one full release cycle plus enough calendar time for teams to schedule the change into their own sprints, with the deprecation warning surfaced somewhere consumers actually see it, such as a build-time warning, not just changelog prose.

It's like building code inspection for a large apartment complex: at a single-family-home scale, the owner's own judgment about what's safe is enough; at hundreds-of-units scale, you need a written building code, an inspector role separate from the contractor doing the work, and a permit/notice process before anyone tears out a load-bearing wall - otherwise 'looks fine to the person who touched it' quietly stops meaning the building is actually safe.

saying these in an interview costs you the question

  • Proposes relying purely on code review/trust with no automated API-diff tooling
  • Doesn't distinguish structural breaking changes (catchable by tooling) from behavioral ones (need human/test judgment)
  • Suggests removing deprecated symbols immediately without a lead-time window
  • Has no answer for 'who has authority to say this is breaking' beyond 'whoever wrote the PR'
  • Treats the changelog as a documentation afterthought rather than a source-of-truth artifact

context