How would you stop published OpenAPI documents from drifting away from the running services?
answer
- Drift is a missing feedback loop
- Two sources of truth will disagree
- Fail the build on the regenerated diff
- Publish from the deploy, not by hand
- Compare against the last published version
basics
~20 sMake the document impossible to bypass: one generation path per service, a CI check that fails when the regenerated spec differs from the committed one, publication from the deploy pipeline rather than by hand, and runtime checks that responses still match the schemas.
solid answer
~50 sDrift is not a documentation problem, it is a missing feedback loop. Four mechanisms, in order of leverage. First, **remove the second source of truth** — either the server implements interfaces generated from the spec so a mismatch is a compile error, or the spec is generated from the code and committed with a CI job that regenerates and fails on any diff. Second, **publish from the pipeline**, tagged with the build that deployed, so no human step can be skipped and the published document always corresponds to something actually running. Third, **verify at runtime**: contract tests that exercise the API through a generated client, or schema validation of real responses in a test environment, catch what static checks cannot. Fourth, **gate on comparison** — diff each new spec against the last published one and require an explicit decision on breaking changes. Then make it organisational: a catalogue with named owners and a visible staleness signal.
go deeper
Understand that a published document can silently stop matching the service, and that automated checks — not memory — are what keep the two aligned.
Explain the concrete controls: regenerating the spec in CI and failing on a diff, and generating server interfaces so a contract change becomes a compile error.
Design the pipeline end to end — generation, lint, breaking-change diff, contract tests through a generated client, and publication tied to the deploy that produced it.
Own it at fleet scale: a catalogue with named owners, centrally configured required checks, a visible staleness signal, and a deliberate decision about which APIs justify the heavier controls.
## Why drift happens A published document is accurate on the day it is written. It decays because the *cost of updating it* and the *cost of not updating it* are paid by different people at different times. A developer adds a field; nothing breaks; the spec is now slightly wrong. Multiply by a hundred changes and the document becomes something consumers have learned not to trust — at which point they read the code or ask in chat, and the spec stops being maintained at all. So the design goal is not diligence. It is arranging things so that a divergence is *detected automatically* and, ideally, cannot be introduced in the first place. ## Level 1 — eliminate the second source of truth Drift requires two artefacts that can disagree. Collapse them. If the contract is authoritative, generate server interfaces from it and have the implementation implement them. A change on either side that breaks agreement is then a compile error — the strongest, cheapest, earliest signal available. Its limit is that it constrains shapes and routes, not semantics: nothing stops a handler returning an empty list where the contract implies a populated one. If the code is authoritative, generate the document from it, **commit the generated file**, and have CI regenerate and fail on any difference. The document then cannot be stale by construction, and every contract change surfaces as a diff in code review — which also fixes the quieter problem that code-first APIs change without anyone reviewing the change. Anything else — a hand-written spec beside a hand-written implementation, with goodwill in between — will drift. ## Level 2 — publish from the pipeline A manual publication step is a drift source of its own: the repository is right and the developer portal is six months old. Publish as part of deployment, from the same commit, and record which build the published document came from. Two useful properties fall out: the portal can never be ahead of or behind production by more than a deploy, and "which version of the contract is live in this environment" becomes an answerable question. ## Level 3 — verify against the running service Static agreement between a spec and the code that generated it is not agreement with observed behaviour. Two complementary checks: - **Contract tests through a generated client.** Drive the service in a test environment using the SDK generated from the published document. Deserialisation failures, missing fields and unexpected nullability show up as test failures, and you exercise exactly the code path consumers use. - **Response validation.** In test environments, validate real responses against the declared schemas. This catches the cases a generator cannot see — a field that is documented as required but sometimes absent, an enum that has grown a value the spec does not list, an error shape that does not match the declared envelope. These are the only mechanisms that catch semantic drift, and they are worth the cost on APIs with external consumers. ## Level 4 — gate on comparison, not just on validity A spec can be perfectly fresh and still ruin someone's week. Diff each candidate document against the last published one and classify the change; tools such as oasdiff exist for exactly this. Breaking changes should require an explicit, recorded decision rather than being discovered by a consumer. Additive changes can flow through automatically. This turns "did anyone think about compatibility" from a review habit into a pipeline state. ## Level 5 — make ownership and staleness visible At fleet scale the technical controls need an organisational frame: - A **catalogue** listing every service's published document, its owning team and the deploy it came from. Without it, nobody can even enumerate what exists. - **Required checks** — lint, diff, contract tests — configured centrally so a new repository inherits them rather than opting in. - **A staleness signal** that is visible to the owning team: a document whose last publication predates the service's last deploy is suspect by definition, and that comparison is trivially automatable. - **Clear ownership of the shared machinery** — the ruleset, the generation tooling, the publication path — because a pipeline nobody owns degrades quietly. ## What to accept Perfect fidelity is not the target; cost-effective detection is. Descriptions and examples will always lag reality somewhat, and chasing that with process is a poor trade. Concentrate enforcement on the parts consumers' code depends on — routes, status codes, field names, types, nullability, enum values, error shapes — and let prose be improved opportunistically. The measure of success is not that no document is ever wrong, but that a wrong one is found by a pipeline rather than by a consumer.
- Which single control gives the most drift protection for the least effort?Committing the generated document and failing CI when a regeneration differs from it. It is a few lines of pipeline, it makes staleness structurally impossible, and it turns every contract change into a reviewable diff — which also catches the quieter problem of public fields being renamed without anyone deciding to.
- Why aren't static checks enough on an API with external consumers?Because they prove the document matches the code that produced it, not that observed behaviour matches the document. A field documented as required can still be absent at runtime, an enum can grow a value, an error path can return a different shape. Contract tests through a generated client and response validation in a test environment are what catch that.
- How do you decide what level of enforcement a given API deserves?By blast radius. An internal service with one consumer in the same repository needs little more than the diff gate. A public or partner API — many consumers, generated SDKs, no ability to coordinate a fix — justifies contract tests, breaking-change gating and pipeline publication. Uniform maximum enforcement across a fleet mostly buys resentment.
- What signal tells you a service's published document is probably stale?Its last publication predating the service's last deployment. That comparison needs only a catalogue entry recording the build a document came from, and it is fully automatable — it will not catch semantic drift, but it reliably surfaces documents that were never republished after a change shipped.
saying these in an interview costs you the question
- Treats drift as a discipline problem to be solved by reminders
- Publishes documentation manually after deployment
- Assumes a code-generated spec needs no verification
- Runs contract checks nightly instead of per change
- Enforces maximum ceremony on every service regardless of consumers