skip to content

At an organization with dozens of teams sharing internal library modules, how do you keep the api/implementation split from eroding over time, and what tooling or process actually enforces that a published api module's contract doesn't silently break its consumers?

level: principalimportance: should knowfreq 30%

answer

  1. japicmp/revapi bytecode diffing
  2. semver gates on api modules
  3. ArchUnit module-boundary CI checks
  4. structural convention on api artifact naming
  5. manual discipline doesn't scale, automate it

basics

~20 s

You need automated checks, not just good intentions - tools comparing a new API module version's bytecode against the previous one, flagging anything that would break callers, plus a versioning rule enforcing the intended split.

solid answer

~50 s

At scale, informal discipline like 'just be careful what you expose' reliably fails, so organizations pair the api/impl module split with automated binary and source compatibility checking - tools such as japicmp or revapi for JVM languages diff two versions of a jar and fail the build if a public method signature changed or a public class was removed incompatibly - alongside a strict semantic-versioning policy tied to those checks, where only a major version bump may introduce a breaking API change. Many organizations also add a build-system-level convention or custom Gradle or Maven plugin that only lets modules following an api naming pattern be depended on by other teams, forbidding implementation modules from being depended on directly, plus a dependency-graph or ArchUnit-style linter running in CI, since code review alone does not scale reliably across dozens of independent teams.

go deeper

for a junior

Not expected to know this in depth; a fair answer just recognizes that more teams sharing a module means you need more automated checks, not just trust.

for a middle

Should mention semantic versioning as part of the answer, even without naming specific compatibility-diffing tools.

for a senior

Should describe at least one concrete enforcement mechanism, such as a compatibility-diff tool or an architecture-boundary CI test, beyond simply 'being careful.'

for a principal

Should synthesize multiple layers - structural convention, automated compatibility diffing, semver policy, and CI boundary tests - discuss their trade-offs and failure modes, and reference a plausible real-world-scale example.

## The enforcement layers Multiple enforcement layers work together at this scale. 1. **The first is a structural convention**: dedicated api artifacts, sometimes reinforced by a custom build-tool plugin that fails the build if an implementation artifact appears as a dependency anywhere outside the module that owns it - for example, a Gradle plugin that inspects the resolved dependency graph and errors out if `foo-impl` shows up as a dependency of any module other than foo's own runtime wiring code. 2. **The second layer is automated compatibility diffing**: tools like `japicmp` or `revapi` for Java compare a previously released jar's public signatures against a candidate build and categorize each change as source-compatible, binary-compatible, or breaking, failing CI on a breaking change unless the change is accompanied by an explicit major-version acknowledgment. 3. **The third layer is semantic versioning itself**, used as the communication contract between teams: a `MAJOR.MINOR.PATCH` scheme where only a major version increment may contain a breaking API change, letting consumers safely auto-upgrade minor and patch versions through their dependency manager while a major bump becomes a deliberate, visible signal requiring active review of a migration guide. 4. **A fourth layer, architecture-boundary tests** such as ArchUnit rules or a module-boundary verification tool, runs in CI to catch code reaching into another module's internal package even inside the same monorepo, independent of whether a separate published artifact exists at all. ## Why the layering is necessary This layered approach exists because manual vigilance does not survive organizational scale. - With two teams, a Slack message and a code review are enough; with dozens of teams releasing independently, no single reviewer sees every consumer of a shared module, and per **Hyrum's Law**, informal discipline fails exactly when it matters most - once a module has enough users, undocumented behavior gets depended on regardless of anyone's intentions. - The organization needs the api/impl separation's guarantee to hold without relying on any individual engineer's memory or goodwill, which in practice means encoding the rule as an automated, CI-enforced gate rather than a policy document nobody reliably reads. ## The trade-offs of the tooling The trade-offs of this tooling are real. - **Compatibility checkers** need baseline configuration and occasionally produce false positives - a change that is technically ABI-breaking but practically harmless, such as widening a parameter type, still needs a manual override or suppression annotation reviewed by a human. - **The semver discipline** itself imposes genuine process overhead, since every breaking change now needs a deprecation cycle, migration documentation, and cross-team coordination, which slows down API evolution compared to a single-team codebase that can just change everything atomically in one commit. - **Overly strict enforcement** can also push teams toward additive-only API design that never removes anything, accumulating years of cruft, because a proper deprecate-then-remove cycle is expensive to execute across an entire organization. ## Failure modes at this scale Failure modes at this scale are distinct from the single-team case. 1. A team may ship what they believe is a minor internal cleanup, but the compatibility checker was never configured to catch a specific class of breaking change - for instance, altering a default method's runtime behavior without touching its signature, which a purely signature-based checker misses entirely - and dozens of downstream services break simultaneously in production. 2. Conversely, teams learn to **game the checker** by only ever adding and never removing or changing anything, producing API modules cluttered with numerous overlapping deprecated methods that nobody dares delete. 3. And as a monorepo grows large enough, ArchUnit-style boundary tests can become slow or flaky, tempting teams to disable them under deadline pressure, which silently reopens exactly the leakage the whole system was built to prevent. ## How this looks at platform scale This is essentially how large platform organizations manage API evolution in practice: a public/internal split enforced structurally, compatibility diffing gating every release, and a documented semantic-versioning policy that lets many independent consuming teams upgrade a shared dependency with reasonable confidence that nothing will silently break - precisely because the enforcement is automated in CI rather than trusted to any individual engineer remembering the convention. The pattern generalizes beyond Java: any ecosystem publishing a shared library to many independent internal consumers eventually needs the same three ingredients - - a structural boundary, - an automated compatibility gate, - and a versioning policy that turns 'breaking change' into a visible, deliberate signal rather than a silent surprise. The underlying principle is the same one that motivates the api/impl split at the level of a single module, just applied recursively at organizational scale: the fewer things a team promises to keep stable, and the more automatically that promise is checked, the more freedom that team retains to keep improving what's underneath it.

  • What's the difference between a binary-incompatible change and a source-incompatible-only change, and why does that distinction matter for tooling?
    A source-incompatible change breaks compilation of consumer code that must be recompiled against the new version, such as adding a new required method to an interface. A binary-incompatible change breaks already-compiled consumer bytecode even without recompilation, such as removing a class entirely or changing a method in a way that alters its erasure signature. The distinction matters because some consumers run against a dependency without recompiling everything, relying on the JVM's own runtime linking, so binary compatibility is the stricter and sometimes hidden category that a check limited to 'does it still compile' can miss entirely.
  • Why would a compatibility checker sometimes let a technically-breaking change through with an override or annotation?
    Because some flagged changes are false positives relative to real-world risk - for example, altering the internals of a method that has been marked deprecated for two versions already, or a class explicitly marked unstable or experimental. Tooling therefore typically supports an explicit suppress-and-acknowledge mechanism reviewed by a human, rather than unconditionally blocking every technically-detectable change regardless of its actual practical impact.

Building codes and inspectors for a city with hundreds of independent contractors: one inspector can personally vouch for a single building, but a city of thousands of buildings needs codified rules plus automated permit checks, because no individual can personally verify that every contractor followed every rule.

saying these in an interview costs you the question

  • Believes code review alone scales to dozens of independent teams sharing a module
  • Unaware of any compatibility-diffing tool category, even without needing to name a specific tool
  • Conflates semantic versioning with simply bumping the version number whenever convenient
  • Never mentions automating the enforcement, relying purely on documentation or policy

context