skip to content

How do you decide whether to extract a schema into OpenAPI components or leave it inline?

level: principalimportance: should knowfreq 28%

answer

  1. A shared ref means shared fate
  2. Same meaning, not same fields today
  3. Named components become named SDK types
  4. Creation payload is not the resource
  5. Cross-service sharing is a coordination cost

basics

~20 s

Extract when several operations genuinely describe the same thing and should change together; leave it inline when the shape is incidental to one operation. A shared component is a coupling decision — every consumer of every referencing operation inherits its changes.

solid answer

~50 s

Reuse in a spec is not free deduplication; it is a statement that these positions must evolve together. Extracting a schema into `components.schemas` gives one name, one place to change, and — importantly — one generated class in every SDK, which is what you want for a genuine domain concept appearing in many operations. But it also means a field added for one endpoint appears in the model every other endpoint hands its clients. The signal to extract is *sameness of meaning*, not sameness of shape today: two objects that happen to have the same three fields but answer different questions will diverge, and merging them buys a future migration. The commonest concrete trap is sharing one model between a request and a response when the server owns half the fields; `readOnly` and `writeOnly` handle mild cases, and separate `OrderCreate`/`Order` components handle the rest. Inline shapes stay legitimate for wrappers and one-off envelopes.

go deeper

for a junior

Know that putting a schema under components and referencing it means every user of that reference gets the same definition, including future changes.

for a middle

Explain the concrete consequences: one named type in generated SDKs, one place to change, and a field added for one endpoint showing up in all the others.

for a senior

Show the read/write asymmetry judgment — why a create payload and a resource representation usually deserve separate components, and where readOnly/writeOnly stops being enough.

for a principal

Own the registry as a coordination budget: which components are org-wide and vendored as pinned artefacts, which stay service-local, and how naming and unused-component cleanup are enforced.

## Reuse is a coupling decision Every `$ref` is an assertion: *whatever happens to this definition happens here too*. That is exactly what you want when many operations genuinely describe the same concept, and exactly what you do not want when two shapes merely rhyme. The naive rule — "extract anything that appears twice" — treats a spec like source code being DRY-ed. But a spec is a contract read by many consumers, and consolidating two definitions merges their change histories permanently. A schema used by fifteen operations cannot be adjusted for one of them. ## What extraction actually buys **One place to change.** The error envelope, the pagination wrapper, the money type: these must be identical everywhere or the API is inconsistent, and a component makes that structural rather than aspirational. **One name.** Named components become named types in generated SDKs and named anchors in rendered documentation. An inline schema gets a synthesised name from most generators — derived from the operation and property path — which is ugly, unstable across regenerations, and impossible for a consumer to talk about. If you want consumers to have a `Money` class rather than an anonymous nested type, `Money` must be a component. **A vocabulary.** A well-curated `components.schemas` list *is* the domain model as the API presents it. Reviewers can read it in isolation and ask whether the concepts are right. ## What extraction costs **Ripple.** Adding a field to a shared schema adds it to every operation that references it. That is usually additive and harmless — but not always: a field meaningful in one context is noise or, worse, an information leak in another. **False sameness.** Two schemas with identical fields today may represent different concepts. Merge them and the first divergence forces either an awkward optional field, a composition workaround, or a split that breaks consumers who have already generated code against the merged type. **Read/write asymmetry.** The most frequent concrete failure. A creation request and the resource representation are not the same object: the server assigns the id, timestamps, computed totals and status; the client supplies a subset. Sharing one schema forces either the id to be optional in the response — weakening the contract for every reader — or a required id that clients cannot supply. `readOnly` and `writeOnly` express the mild version of this, and generators honour them unevenly, so for anything beyond a couple of fields distinct `OrderCreate` and `Order` components are clearer and safer. **Over-composition.** Chains of composed fragments assembled to avoid repeating four fields produce a document nobody can read and generated types nobody recognises. Reuse should reduce the reader's work; when a reader must open five definitions to learn what one response looks like, it has stopped doing that. ## A usable decision rule Ask three questions, in order. 1. **Does it have a name a domain expert would recognise?** `Money`, `Address`, `Problem`, `PageInfo` — yes, extract. A nameless "the object inside the `filters` key of this one search endpoint" — leave inline. 2. **Would a change to it need to reach all uses?** If the answer is "yes, and if it didn't we'd have a bug", extract. If it is "it depends on the endpoint", do not. 3. **Do consumers need to hold it as a value?** If clients will pass it around, store it, or write a function over it, it needs to be a named type in their SDK, so it must be a component. Envelopes, one-off wrapper objects and single-use request bodies fail all three and are fine inline. Repeated *primitive* shapes — a `string` with a `format` — are usually not worth a component either; the indirection costs more than it saves. ## Ownership at scale Across many services the question becomes organisational. A shared component library — an error envelope, pagination, common identifier types — is worth publishing as a pinned, versioned artefact that services vendor in, not referenced live from a URL, so each team upgrades deliberately and the change appears as a reviewable diff. Everything else should stay service-local: a component shared across service boundaries is a coordination point, and the number of those is a design budget. Finally, curate. Unreferenced components rot, and many generators emit a class for every entry in the registry, so dead definitions become dead code in shipped SDKs. A lint rule for unused components and a review expectation that new components are named deliberately keep the registry meaningful.

  • Should a create request and the resource it returns share one schema?
    Usually not. The server owns the id, timestamps and computed fields; the client owns a different subset. Sharing forces those fields to be optional, which weakens the contract for every reader. `readOnly` and `writeOnly` cover a field or two, but generator support is uneven — beyond that, distinct `OrderCreate` and `Order` components are clearer and safer.
  • Two operations currently have identically shaped objects. Is that enough reason to extract one component?
    No. Identical shape is not identical meaning. Ask whether a change to one should necessarily reach the other; if the honest answer is "it depends", they are two concepts that happen to coincide. Merging them makes the first divergence a breaking split for consumers who already generated code against the merged type.
  • How would you share components across services without coupling their release cycles?
    Publish the shared set — error envelope, pagination, common identifier types — as a pinned, versioned artefact that each service vendors into its repository, rather than referencing a live URL. Each team upgrades deliberately, the change lands as a reviewable diff, and no service can be broken by an edit made elsewhere.

saying these in an interview costs you the question

  • Extracts every shape that appears twice on principle
  • Treats a spec like code to be DRY-ed
  • Shares one model between create requests and responses
  • Builds deep composition chains to avoid repeating fields
  • References another team's live spec URL to share types

context