skip to content

What does an OpenAPI spec differ such as oasdiff actually compare, and what verdicts can it report?

level: middleimportance: must knowfreq 62%

answer

  1. Two files, never a running service
  2. Baseline revision versus candidate revision
  3. Walks paths, schemas, required-ness, enums
  4. Every difference gets a severity
  5. Breaking, non-breaking, informational

basics

~20 s

A spec differ compares two revisions of one machine-readable API description - a baseline OpenAPI document against the candidate - and classifies each difference as breaking, non-breaking or informational. It reads documents only; it never calls a running service.

solid answer

~40 s

`oasdiff` takes two OpenAPI documents, a **baseline** revision and the **candidate** revision from the branch under review, and walks them structurally: paths, operations, parameters, request and response schemas, required-ness, enum members, security schemes. Each difference is emitted as a typed change with a severity, and the breaking-changes mode keeps only the ones the rules classify as breaking, so the step can fail the job. The comparison is **document versus document**: no service runs, no traffic is replayed, no consumer is consulted. That is the sharp contrast with a Pact file, which pairs interactions a named consumer recorded with the responses a running provider actually returns. A differ answers *did the described interface shrink?*; Pact provider verification answers *does this provider still satisfy this consumer?*

code

bash · 5 lines
bash
# compare the published baseline document with the candidate on this branch
oasdiff breaking published/openapi.yaml candidate/openapi.yaml

# reports only the differences its rules classify as breaking;
# CI is wired so that a non-empty result fails the job

go deeper

for a junior

Recall what the check reads: two revisions of an API description file, not a live service. Knowing that CI can compare yesterday's OpenAPI document with today's and flag removals is enough at this level.

for a middle

Explain the inputs and the verdict classes - baseline versus candidate, breaking versus non-breaking versus informational - and why the order of the two arguments changes the question being asked.

for a senior

Show what the verdict is worth in production: a clean diff constrains the description, not the deployment, and you should be able to name the evidence you put beside it before calling a release safe.

for a principal

Own what the gate's baseline represents across many services, and be ready to argue what it costs an organisation when a differ's clean verdict gets read as a deployment decision.

## What a spec differ takes as input A **spec differ** compares two revisions of one machine-readable interface description and reports how the second departs from the first. For a REST API that description is an OpenAPI document and the usual tools are `oasdiff` and `openapi-diff`; the Protobuf equivalent is `buf breaking`, and a graph API has its platform's schema check. All of them take exactly two artefacts: - the **baseline** - the revision that is already published or already deployed, normally read from a released git ref or from a store that holds the last approved revision; - the **candidate** - the revision produced by the branch under review. Nothing else takes part. No provider process starts, no HTTP request is sent, no consumer's test suite runs, no production traffic is inspected. That is worth stating plainly, because the phrase "contract check" gets used for both this and provider verification against a recorded pact, and the two read completely different inputs. ## What it walks, and how it classifies what it finds The differ parses both documents into a structure and walks it: paths and operations, parameters and their required-ness, request bodies, response codes, response schemas and the types inside them, enum members, and security schemes. Every difference becomes a **typed change** - an entry naming what changed, where in the document it changed, and which of the tool's rules it matched. Each typed change then gets a **verdict**. The vocabulary varies between tools, but the classes are consistently three: | Verdict class | What it covers | Usual CI treatment | | --- | --- | --- | | Breaking (error) | a change the tool's rules say can break an existing caller of the described interface: a removed operation, a response field that disappears, a request parameter that becomes required | fail the job | | Non-breaking (warning) | a change the rules treat as safe for a tolerant caller but worth a human glance: a new optional parameter, an added response field | annotate, do not fail | | Informational | text carrying no interface obligation: descriptions, summaries, examples | report only | The taxonomy those rules encode is API-versioning practice. What matters about the tool is that it applies that taxonomy **mechanically, to a document**, with no judgement about whether the flagged change matters to anyone alive. Two mechanical consequences follow, and both surprise people: 1. **The comparison is not symmetric.** Which revision you pass as the baseline defines the question being asked. Baseline-then-candidate asks whether callers written against the old description still fit the new one. Swap the two and you are asking the opposite question, and you get a different set of changes back. 2. **Every verdict is about the description, not the deployment.** If the running service already disagreed with its own OpenAPI document, both revisions can be equally wrong and the diff will still come back clean. ## How that differs from what a consumer-recorded pact compares A pact file is a JSON artefact that a consumer's own test run produced: a list of interactions, each carrying a request the consumer said it would send, the response it said it expected, and a named provider state. Provider verification replays those interactions against a **running provider** and checks the responses that actually come back. So the two checks answer different questions from different inputs: | | Spec differ | Pact provider verification | | --- | --- | --- | | Inputs | two revisions of one description | one consumer's recorded interactions plus a running provider | | Question answered | did the described interface shrink? | does this provider still satisfy this consumer? | | Needs another team? | no | yes, a consumer must have written and published a pact | | Sees real responses? | no | yes | | Names who is affected? | no | yes, by consumer name, for consumers that have pacts | | Blind to | meaning, usage, the running service | anything no consumer wrote an expectation for | That difference decides where each is useful. The differ's strength is that it needs nobody's cooperation and covers the whole described surface, including operations nobody has ever called; its weakness is that it has no idea who calls what. Verification's strength is that a failure names a consumer; its weakness is that coverage stops at the expectations consumers actually recorded. ## Reading the result in CI What the build gets back is a report of typed changes plus a process exit status, and the step is normally configured so a breaking classification fails the job while everything else is annotation. Three things are worth doing with that output rather than just reading red or green: - keep the report as a build artefact, so the difference between two releases can be reconstructed months later; - record an accepted breaking change as an explicit, reviewed entry in the repository, rather than rerunning with the check switched off; - remember that a clean result is a claim about two files, and say so out loud when someone reads it as clearance to deploy.

  • If you swap which OpenAPI document is passed as the baseline, does the verdict change?
    Yes. The comparison is not symmetric. Baseline-then-candidate asks whether callers written against the old description still fit the new one; reversing the order asks whether callers written against the new description would have fitted the old one. Different question, different set of reported changes. Passing them the wrong way round is a common reason a gate looks quiet when it should be loud.
  • Can a spec differ tell you whether anyone actually calls the endpoint it flagged?
    No. It reads two descriptions and knows nothing about traffic. It flags the removal of an operation no client has called in a year exactly as loudly as one every client uses. Deciding that a flagged change is acceptable needs evidence the differ cannot supply, such as usage telemetry or a consumer-recorded contract that names who depends on what.
  • Does a clean spec diff prove the provider still returns what its document says?
    No. The differ never contacts the service; it compares two documents. If the implementation had already drifted from its own description, both revisions can be equally wrong and the diff is still clean. A clean diff constrains the description, not the deployed build.

It is a proofreader comparing two printings of a catalogue: it can tell you a page was removed, but not whether anyone was reading that page.

saying these in an interview costs you the question

  • Thinks the differ calls the running API and compares responses
  • Says a clean spec diff proves no consumer is affected
  • Treats the comparison as symmetric, so argument order looks irrelevant
  • Confuses it with replaying a consumer's recorded interactions
  • Assumes every reported difference should fail the build