skip to content

At an organization with dozens of independently-deployed services, what practices let a platform team enforce contract compatibility and coordinate breaking changes across teams, without a central team reviewing every API change by hand?

level: principalimportance: should knowfreq 40%

answer

  1. CI-enforced compatibility checks, not manual review
  2. policy set centrally, checking runs per-team
  3. usage telemetry to know who depends on what
  4. deprecation window as explicit policy
  5. review gate reserved for breaking/new-public contracts only

basics

~20 s

Automate the checks instead of relying on manual review: run compatibility checks in CI, block deploys that break a registered consumer contract, track who's still using an old version with telemetry, and set an org-wide deprecation policy everyone follows.

solid answer

~50 s

At scale, manual review doesn't work - no central team can review every API diff across dozens of services fast enough, and it creates a bottleneck that teams route around. The scalable answer is to encode compatibility rules as automated gates in each team's own CI/CD pipeline: a schema-registry compatibility check or a consumer-driven-contract 'can-i-deploy' gate that runs on every PR/deploy and blocks it locally, without a human in the loop. This decentralizes enforcement while keeping the rules consistent, because the rules themselves are set centrally as policy, even though the checking runs per-team. Coordination for genuine breaking changes still needs process: an explicit deprecation policy, usage telemetry per consumer so a team knows who to notify before flipping a switch, and often a lightweight architecture/API review only for new public contracts or genuinely breaking changes, not routine additive ones. The org-wide investment is in tooling and policy, not headcount for manual gatekeeping.

go deeper

for a junior

Understands that big orgs need some automated check rather than a person reviewing every change, even without design details.

for a middle

Can name specific automated mechanisms (schema registry, CDC gate) as the enforcement layer and knows manual review doesn't scale.

for a senior

Separates policy definition from checking execution, designs the CI gates, and knows usage telemetry is required to safely retire versions.

for a principal

Designs the org-wide governance model: what's centralized (policy, tooling ownership) versus decentralized (checking execution), scopes human review narrowly to genuine breaking/new-public-contract cases, and builds in periodic auditing against governance decay.

## Why manual review stops scaling Once an organization has more than a handful of services and teams, manual, human review of every API contract change becomes a bottleneck that doesn't scale: a small central 'API review board' cannot keep pace with dozens of teams shipping changes daily, review turnaround becomes the limiting factor on deploy velocity, and teams predictably start routing around it - either by avoiding review for changes they wrongly believe are safe, or by batching changes into big, hard-to-review 'version 2' rewrites instead of small incremental ones. ## Automated, decentralized enforcement The mechanism that scales instead is shifting compatibility enforcement from a central human gate to automated, decentralized CI/CD checks, while keeping the rules that define compatibility centralized as policy. Concretely: - A **schema registry** with an org-mandated compatibility mode (e.g., FULL) rejects incompatible schema registrations automatically, wherever they're attempted, without anyone from a central team looking at the diff. - A **consumer-driven-contract system** with a 'can-i-deploy' gate blocks any service's deploy pipeline automatically if it would violate a registered consumer's expectations, again without a human reviewer. - **Linters or service-template defaults** (e.g., banning strict-deserialization-by-default JSON configs, mandating default branches on enum switches) encode the tolerant-reader pattern into what a new service gets out of the box, rather than relying on every engineer independently knowing the pattern. ## The two kinds of work it separates The reason this decentralization-with-centralized-policy model exists is that it separates two different kinds of work that don't scale the same way, policy versus execution: - **Defining what's safe** - a policy question, naturally centralized. - **Checking whether a specific change is safe** - an execution question that must scale with the number of changes, and therefore must be automated and run locally in each team's own pipeline, not queued through a shared bottleneck. A platform or API-governance team's real leverage point becomes the tooling and defaults rather than personally reviewing diffs. ## The trade-off The trade-off is that automation can only catch what it's told to check. Compatibility tooling reliably catches syntactic/structural breakage (removed fields, changed types, violated schema rules) but is **blind to semantic breakage** - a field that's still present, still the same type, but whose meaning quietly changed (e.g., a 'discountPercent' field that used to mean 0-100 and now means 0-1). Catching semantic drift still requires some human judgment, typically reserved for a lightweight review specifically scoped to new public contracts or explicitly-flagged breaking changes, rather than routine additive changes, which is what keeps the human-review bottleneck small and proportionate instead of universal. There's also a genuine cost to maintaining the automation itself: someone has to run the schema registry or contract broker as reliable infrastructure, keep compatibility-mode policy up to date as the org's risk tolerance evolves, and periodically audit that teams haven't quietly set checks to a permissive mode to unblock themselves under deadline pressure - automation reduces but doesn't eliminate governance work, it relocates it from per-change review to periodic policy auditing. ## Failure modes Failure modes at this scale are mostly about governance decay rather than any single technical bug. 1. **Weakened checks.** A common one: individual teams discover they can bypass or weaken an automated check under deadline pressure, and because enforcement is decentralized, there's no single point where someone would necessarily notice - this is why periodic audits or dashboards showing which services/subjects have weakened settings matter as much as the initial tooling rollout. 2. **Coordination breakdown.** A second failure mode is coordination breakdown on genuinely breaking changes: even with perfect automated checking of individual changes, retiring an old major version still requires knowing who's still using it, and without consistent usage telemetry tagged by consumer across all services, a platform team ends up guessing, either leaving dead versions running indefinitely out of caution or breaking a forgotten consumer when they finally remove one. ## The paved road in practice A recognizable real-world pattern combining these ideas is how large platform organizations (Netflix and Uber's internal engineering blogs have both described versions of this) run an internal developer platform that provides contract testing, schema registries, and service scaffolding as default-on paved-road tooling: teams get compatibility enforcement 'for free' by using the standard service template and CI pipeline, a small central platform team owns and evolves the policy and tooling itself, and a much smaller architecture-review function is reserved only for genuinely new external-facing contracts or explicitly-declared breaking migrations - not for the routine, additive majority of API changes that automated checks handle unattended.

  • What kind of contract breakage can automated compatibility checks (schema registries, CDC tests) NOT catch, and how do organizations typically handle that gap?
    They can't catch semantic drift - a field that keeps its name and type but silently changes meaning, like a percentage field switching from a 0-100 scale to a 0-1 scale. Organizations typically handle this with a narrowly-scoped human review reserved for new public contracts or explicitly-flagged breaking changes, rather than reviewing every routine additive change, keeping the human bottleneck small and proportionate.
  • Why is decentralizing compatibility checking while centralizing compatibility policy more scalable than either fully centralizing or fully decentralizing both?
    Fully centralizing both makes a small team the review bottleneck for every change across the org, which doesn't scale with deploy frequency. Fully decentralizing both means every team invents its own definition of 'safe,' producing inconsistent standards and integration surprises; splitting the two lets one team or committee reasonably own the policy definition while the actual checking runs automatically, in parallel, inside every team's own pipeline.
  • What governance risk does automated, decentralized enforcement introduce that centralized manual review doesn't have, and how is it typically mitigated?
    Because enforcement runs inside each team's own pipeline, an individual team under deadline pressure can quietly weaken a check with no central point automatically noticing. This is typically mitigated with periodic audits or dashboards that surface which services or subjects have non-standard, permissive settings, rather than relying on the automation alone.

Like a city that enforces building codes through automated permit-checking software builders must pass before construction, rather than sending a single inspector to personally approve every nail in every building in the city.

saying these in an interview costs you the question

  • proposes a central team manually review every API change as the scaling solution
  • doesn't distinguish semantic breakage from structural/syntactic breakage that tooling can catch
  • assumes automated compatibility checks are sufficient with no ongoing governance/auditing
  • has no mechanism for knowing which consumers still use a version before retiring it
  • treats platform tooling as a one-time rollout rather than something requiring ongoing policy maintenance

context