When two teams integrate over an API, what makes a good API contract, and how do you evolve it without breaking existing consumers?
answer
- additive vs breaking changes
- consumer-driven contract testing
- version the API, don't mutate it
- Pact
- schema as source of truth
basics
~20 sA contract is the agreed shape of requests and responses between two systems - like a form both sides fill out the same way. To change it safely, add new optional fields instead of removing or renaming old ones, so old callers keep working.
solid answer
~40 sAn API contract defines the interface shape (endpoints/fields/types), the semantics (what a field means, what's required, what error codes mean), and guarantees (SLAs, idempotency, ordering) producer and consumer agree to. A good contract is explicit rather than implied by current behavior, and distinguishes backward-compatible changes (adding optional fields, new endpoints, new enum values consumers must tolerate) from breaking changes (removing/renaming fields, changing types, tightening validation, changing semantics). Safe evolution means: default to additive changes; when a breaking change is unavoidable, version the API (URI, header, or schema version) and run old and new in parallel until consumers migrate; and use contract testing (e.g., consumer-driven contracts) so a producer's CI catches a break before it ships, rather than discovering it in production.
go deeper
Understands that changing a response field can break callers and knows to add fields rather than rename them.
Can classify a proposed change as backward-compatible or breaking, and knows to version an API when a break is unavoidable.
Designs the versioning/deprecation strategy for a service, sets up contract or schema tests in CI, and negotiates migration windows with consumer teams.
Sets org-wide API governance: contract review gates, consumer-driven contract tooling adoption, deprecation policy and timelines, and decides when the cost of running N parallel versions outweighs forcing a coordinated migration.
## What an API contract is An API contract is the explicit, agreed-upon interface between a producer (the system exposing the API) and its consumers: - which operations exist - what data shape each request and response takes - which fields are required versus optional - what each field means semantically - what error codes signify - and often non-functional guarantees like latency SLAs and idempotency behavior The contract is what lets two teams build and deploy independently — the consumer team codes against the contract, and as long as the producer honors it, the consumer's code keeps working regardless of internal changes on the producer's side. This is the entire point: a contract is a boundary that decouples two teams' release schedules. ## Where the contract actually lives The artifact defining a contract varies by protocol: - OpenAPI/Swagger for REST - Protobuf `.proto` files for gRPC - a GraphQL schema - or an AsyncAPI spec for event-driven contracts In all cases it's a machine-readable schema, checked into version control, that both sides can validate against — rather than the contract living only as tribal knowledge or 'whatever the current response happens to look like.' ## Classifying every change before shipping Contract evolution is where most integration pain actually happens, because a producer team changing a field can silently break every consumer that depends on it, and the producer often has no visibility into who's calling them or how. The core discipline is classifying every change as backward-compatible or breaking before shipping. **Backward-compatible changes include:** - adding a new optional response field (old consumers ignore fields they don't know about) - adding a new endpoint - adding a new enum value consumers are contractually required to handle gracefully - relaxing validation **Breaking changes include:** - removing or renaming a field - changing a field's type or meaning - tightening validation on an input the producer used to accept - changing default behavior consumers implicitly relied on ## Versioning when a break is genuinely necessary When a breaking change is genuinely necessary, the standard technique is **versioning**: expose the change as a new version (a `/v2/` URI path, an `Accept` header version, or a new schema version), keep the old version running and fully supported, and give consumers a real migration window before deprecating it. This costs real operational overhead — running two code paths, two sets of tests, monitoring which consumers remain on the old version — but the alternative (breaking consumers with no warning) is far more expensive in trust and incident response. ## Proving it with consumer-driven contract testing The verification side is contract testing, specifically consumer-driven contract testing (**Pact** is the best-known tool): each consumer publishes the exact subset of the contract it actually relies on, and the producer's CI runs those consumer-published expectations against every build. If a producer change would break a real consumer's actual usage, CI fails immediately, before the change ships — a much tighter feedback loop than discovering the break via a support ticket after production starts erroring. This also solves a subtler problem: without consumer-driven contracts, a producer team is often overly conservative (afraid to change anything) or overly reckless (assuming nobody relies on a field that turns out to be load-bearing). ## Failure modes when the discipline is missing The failure modes when contract discipline is missing are familiar: | The producer's change | What the consumers hit | |---|---| | A producer renames a field | three downstream consumers throw errors in production simultaneously | | A producer adds a required field with no default | every existing caller's requests start failing validation | | A producer changes an enum's value set | a consumer's switch statement silently falls into a default branch that mishandles the new case | These often surface hours or days after deploy, since the producer's own tests pass (tested against their new contract) while consumers, tested against the old contract, are the ones who break — and because producer and consumer are separately deployed, the person who made the change is rarely the person who gets paged. ## A concrete example A concrete example: Stripe's API is a well-known case of contract discipline done well — they version by date (a version header, e.g., `2023-10-16`), keep every old version running indefinitely, and document exactly which changes are considered backward-compatible (adding a field, a new webhook event type) versus which require a new dated version. This lets thousands of independent integrations keep working for years without a coordinated migration, at the cost of Stripe supporting many parallel API versions in production simultaneously.
- Why is adding a new required field to an API request always a breaking change, even if you also update the documentation?Existing consumers built their request payloads against the old contract and have no code path that populates the new field, so their requests fail validation the moment the field becomes mandatory. Documentation doesn't get read or re-implemented automatically — only a live migration by every consumer fixes it, so a required-field addition needs a version bump or a default value, not just a doc update.
- What does consumer-driven contract testing catch that a producer's own unit and integration tests won't?The producer's own tests validate against the producer's current understanding of its contract, which is exactly what's changing — they won't flag a break in behavior the producer didn't realize a consumer depended on. Consumer-driven contracts capture each consumer's actual real-world usage pattern and run it against every producer build, catching breaks the producer team has no way to anticipate on their own.
- When is it acceptable to skip versioning and just change an existing endpoint's behavior directly?Only when the change is provably additive/backward-compatible for every real consumer — e.g., adding an optional field, loosening validation, or fixing a bug where the old behavior itself violated the documented contract. Even then, teams often stage the rollout and monitor consumer error rates, because 'provably' backward-compatible in theory can still surprise a consumer that depended on undocumented behavior.
A contract is like a printed form: as long as you only add new optional boxes, everyone who already knows how to fill out the old form can keep using it. Renaming or removing a box means every old form is suddenly wrong.
saying these in an interview costs you the question
- Treats 'the tests pass' as proof a change is backward-compatible
- Renames or removes a field and calls it a minor change
- Has no mechanism to know which consumers depend on which fields
- Assumes documentation updates alone will keep consumers in sync
- Adds a required field without a default and doesn't consider existing callers