skip to content

Inter-Service Contracts & Versioning

Independent deployment only works if changing a service does not break its callers, so contracts need compatibility rules and evolution strategy. You will cover semantic versioning, tolerant readers, schema registries and consumer-driven contract testing with tools like Pact.

part ofMicroservices architectureoverview, primer and where to startread it →
on this pageshow

questions

6

In a microservices system, what does it mean for an API change to be 'backward compatible', and what is a concrete example of a change that breaks it?

level: juniorimportance: must knowfreq 80%

answer

  1. additive-only = safe
  2. remove/rename = breaking
  3. semver MAJOR.MINOR.PATCH
  4. optional->required is breaking
  5. old consumer must still parse new response

basics

~20 s

A backward-compatible change lets old clients keep working without updates, like adding a new optional field. A breaking change forces every client to update at once, like renaming or removing a field they depend on.

solid answer

~40 s

Backward compatibility means a new version of a provider's API can be deployed without forcing existing consumers to change or redeploy immediately. Practically: only add new optional fields/endpoints, never remove or rename existing ones, don't change field types or make optional fields required, don't tighten validation, and keep response structures additive. A breaking change is anything that makes an old consumer's parsing or validation logic fail or misbehave: removing a field a client reads, renaming a field, changing a field's type (string to int), making an optional field mandatory in a request, changing enum values, or altering error codes/status semantics a client branches on. Semantic versioning maps this to MAJOR (breaking), MINOR (backward-compatible additions), PATCH (bug fixes, no contract change) - a MAJOR bump signals consumers must review before upgrading.

go deeper

for a junior

Can define backward vs breaking change with a correct example (e.g., adding vs removing a field) and knows the basic MAJOR.MINOR.PATCH meaning.

for a middle

Applies the additive-only rule to both request and response shapes, recognizes optional-to-required as breaking, and knows why rolling deploys need compatibility even within a single version.

for a senior

Distinguishes backward vs forward compatibility, designs API changes to satisfy both, and proactively plans deprecation windows instead of discovering breakage from consumer complaints.

for a principal

Sets org-wide policy on what counts as breaking, mandates compatibility checks in CI, and balances the cost of eternal backward compatibility against the cost of coordinated breaking migrations across dozens of teams.

## What backward compatibility is **"Backward compatibility"** is the property that lets a service provider ship a new version of its API while every consumer that was written against the old version keeps working, unmodified and unaware anything changed. It is the mechanism that makes **independent deployability** possible in a microservices architecture: dozens of teams can release on their own schedules only if none of them can silently break someone else's running code. ## The rule of thumb: additive only The practical rule of thumb is "additive only, never subtractive". You may: - add a new optional request field; - add a new field to a response; - add a new endpoint; - add a new enum value if consumers are required to ignore unknown values. You may not: - remove a field, rename a field, change a field's type; - tighten a previously loose validation rule (e.g., make an optional field required, lower a max-length); - change the meaning of an existing status code; - remove an endpoint that's still in use. ## Forward compatibility, the mirror image **Forward compatibility** is the mirror image: it protects a consumer (or an older provider instance) from breaking when it receives a payload shaped by a newer version of the contract than it was built against - critical during rolling deploys, where old and new instances of the same service run side by side for minutes to hours. A consumer built to be forward-compatible ignores fields it doesn't recognize rather than rejecting the whole payload. ## Semantic versioning as the shared vocabulary Semantic versioning (`MAJOR.MINOR.PATCH`) is the common vocabulary teams use to communicate the compatibility class of a change before anyone has to read the diff. | Bump | What it signals | |---|---| | **PATCH** | no visible contract change - internal refactor or bug fix | | **MINOR** | an additive, backward-compatible change - safe to pull automatically | | **MAJOR** | a breaking change - consumers must consciously opt in and test before adopting it | This convention only has value if it's enforced honestly; a team that ships a field rename as a MINOR bump destroys the trust that lets other teams auto-upgrade, which is the whole point of the scheme. ## The trade-off The trade-off is that strict backward compatibility is **not free**. Keeping deprecated fields alive, running dual code paths, and never tightening validation accumulates technical debt: the schema grows crustier every year, and the provider carries the maintenance cost of every historical shape it ever exposed. The alternative - periodic breaking MAJOR versions - shifts that cost onto every consumer team, who must schedule migration work, and it requires the provider to support at least two live major versions during the transition window, which is itself extra operational cost. There is no version of this that's free; the only choice is who pays and when. ## Failure modes in production Failure modes show up in specific, recognizable ways in production. 1. **Silent breakage** - the most common. A provider ships what its own team believed was a "safe" change - dropping a field they thought was unused - and a consumer three teams away, using it in a way the provider never audited, starts throwing null-pointer exceptions or silently miscalculating totals. This is why field removal should generally require either a deprecation period with usage monitoring or a genuine MAJOR bump with an explicit migration announcement. 2. **Asymmetric compatibility during rollout** - a second common failure. A provider deploys a MINOR change that's backward compatible for consumers but not forward compatible for its own older instances mid-rollout, causing a fleet of half-upgraded pods to reject each other's traffic for the duration of the deploy. 3. **Validation tightening** - a third. Adding a new required field to a request payload looks harmless to the provider team but immediately rejects every existing client that doesn't send it. ## Backward compatibility taken to the extreme A concrete, widely cited example is Stripe's API versioning model: - Stripe **pins every account to a specific dated API version** at the time the integration was built, and it will forever serve that account requests shaped according to that version's contract, even as the "current" version moves forward with breaking changes over the years. - New accounts get the latest version by default; existing integrations are **never force-migrated**. This is backward compatibility taken to its logical extreme - Stripe absorbs the cost of maintaining many historical response shapes internally so that no external consumer is ever surprised by a breaking change they didn't opt into. It's expensive operationally, but for a payments API a broken integration means a merchant can't take money, so the cost of surprise breakage vastly outweighs the cost of maintaining old shims.

  • Why is making a previously optional request field required considered a breaking change even though the field already existed?
    Because existing consumers that never sent that field will now get their requests rejected, even though nothing about the field's name or type changed. Breaking compatibility is about the consumer's ability to keep operating unmodified, not just about the shape of the schema - a stricter validation rule is just as breaking as a removed field.
  • How does forward compatibility differ from backward compatibility, and why do you need both in a microservices mesh?
    Backward compatibility protects old consumers when the provider changes; forward compatibility protects a provider (or a consumer) when it receives a payload from a newer version it hasn't been updated to understand yet, such as during a rolling deploy where old and new instances run side by side. You need both because deployments are never atomic across dozens of services - for a window, old and new code coexist, so messages must survive being read by either.
  • If a MAJOR version bump is needed, how do teams typically avoid a hard cutover for all consumers at once?
    They run the old and new major versions side by side for a deprecation window - via separate URI paths, header-based routing, or content negotiation - and migrate consumers incrementally, monitoring usage of the old version until traffic drops to zero before retiring it.

Like adding a new optional topping to a pizza menu (fine, old orders still work) versus renaming 'pepperoni' to 'spicy meat' on the menu (every regular customer's order breaks).

saying these in an interview costs you the question

  • says removing an unused field is always safe without checking consumers
  • treats semver as just a number bump ritual with no compatibility discipline behind it
  • doesn't distinguish request-schema compatibility from response-schema compatibility
  • assumes all consumers upgrade at the same time as the provider deploys
  • conflates 'backward compatible' with 'never changes'

context

open as a page

How does consumer-driven contract testing (e.g., with a tool like Pact) work, and what problem does it solve that end-to-end integration tests don't?

level: middleimportance: must knowfreq 70%

basics

~10 s

Each consumer writes down exactly what it expects from a provider's API. The provider runs those expectations as tests before deploying, catching breakage early without needing a slow, flaky full end-to-end environment.

open as a page

When a public-facing service needs to expose multiple versions of its API at once (e.g., /v1/orders vs /v2/orders, or via an Accept header), what are the main versioning strategies, and what's the operational cost of each?

level: seniorimportance: must knowfreq 75%

basics

~20 s

You can put the version in the URL path, in a request header, or in a query parameter. URL versioning is the simplest and most visible; header-based versioning keeps URLs stable but is less discoverable. All of them mean you're running and maintaining more than one version of your API at once.

open as a page

What is the tolerant reader pattern in service-to-service integration, and what specifically should a consumer's parsing code do (and not do) to implement it?

level: middleimportance: should knowfreq 55%

basics

~20 s

A tolerant reader only looks at the specific fields it actually needs from a response and ignores everything else, so it doesn't break when the provider adds new fields or extends the data it sends.

open as a page

When a schema registry enforces 'BACKWARD' compatibility mode on a service's request/response schema, what exactly does it check before allowing a new schema version to be registered, and how does that differ from 'FORWARD' and 'FULL' compatibility modes?

level: seniorimportance: should knowfreq 45%

basics

~20 s

A schema registry checks new versions of a data format before allowing them to be published, using rules like 'can old readers still read new data' (forward) or 'can new readers read old data' (backward), and 'full' means both must hold at once.

open as a page

At an organization with dozens of independently-deployed services, what practices let a platform team enforce contract compatibility and coordinate breaking changes across teams, without a central team reviewing every API change by hand?

level: principalimportance: should knowfreq 40%

basics

~20 s

Automate the checks instead of relying on manual review: run compatibility checks in CI, block deploys that break a registered consumer contract, track who's still using an old version with telemetry, and set an org-wide deprecation policy everyone follows.

open as a page