In OpenAPI, is a self-referential schema legal, and what breaks with circular $ref chains?
answer
- Legal — trees are real data
- Ask whether an instance can terminate
- Required all the way round is unsatisfiable
- Pointers survive cycles; copies do not
- Depth limits belong in prose and code
basics
~20 sYes — recursive references are legal and normal for tree-shaped data, provided the cycle passes through an optional or array property so an instance can terminate. Cycles through required properties describe unsatisfiable data, and any tool that inlines references instead of keeping them recurses forever.
solid answer
~50 sA schema referencing itself, directly or through a chain, is valid: a comment with `replies` that are comments is the canonical case. What matters is whether the cycle can *stop*. If the self-reference sits behind an optional property or an array, a finite instance exists — an empty array ends the recursion. If every hop in the cycle is `required`, no finite document can satisfy the schema, and you have written a contract nothing can fulfil. Separately, the cycle stresses tooling: resolvers that keep the `$ref` in place handle recursion fine, while anything that dereferences — full inlining, naive example generation, some mock servers — expands forever until it hits a depth limit or dies. If a recursive model is breaking your pipeline, the fixes are to bundle rather than inline, to cap example depth, or to break the cycle in the model by referencing a child by id instead of embedding it.
code
yaml · 32 linescomponents:
schemas:
# SAFE: the cycle passes through an array, so [] terminates it
Comment:
type: object
required: [id]
properties:
id:
type: string
replies:
type: array
items:
$ref: '#/components/schemas/Comment'
# BROKEN: every hop is required, so no finite instance validates
Node:
type: object
required: [id, child]
properties:
id:
type: string
child:
$ref: '#/components/schemas/Node'
# ALTERNATIVE: break the cycle with an identifier
Category:
type: object
properties:
id:
type: string
parentId:
type: stringgo deeper
Know that a schema may reference itself — a comment containing replies of the same type is normal — and that the recursion needs an optional or array edge to stop.
Explain why an all-required cycle admits no finite instance, and why keeping references as pointers works while copying targets into place cannot terminate.
Diagnose it in a real pipeline: identify inlining versus bundling, sample-generation depth limits, and the unbounded-nesting request that overloads a parser.
Decide when recursion belongs in the contract at all — embedded trees versus identifier links — weighing payload bounds, tool support across every consumer and the denial-of-service surface.
## Recursion is a feature, not an accident Plenty of real payloads are recursive: comment threads, org charts, category trees, filter expressions, file-system nodes. OpenAPI's Schema Object inherits JSON Schema's reference model, and JSON Schema explicitly permits a schema to reference itself. So `Comment` with a `replies` array whose `items` is `#/components/schemas/Comment` is legal and idiomatic. Cycles can also be indirect: `A` references `B`, `B` references `A`. Nothing forbids that either. ## The question that decides validity: can an instance terminate? A recursive schema describes a set of documents. That set is non-empty only if the recursion can bottom out. Two shapes matter. **Terminating.** The self-reference is reached through an optional property, or through an array's `items`. An empty array, or an absent property, ends the descent. `{ "id": "c1", "replies": [] }` satisfies the recursive `Comment` schema, so the contract is satisfiable. **Non-terminating.** Every step of the cycle is mandatory: `Node` has `required: [child]` and `child` is a `Node`. Now every valid instance must contain another valid instance, forever. No finite JSON document satisfies it. This is not an error a validator is obliged to catch — most will happily accept the document as a *spec* and then reject every payload you send. It is a modelling bug, and the symptom in the wild is "the schema validates but nothing I send ever passes". The same trap appears through composition: a schema that composes itself unconditionally, rather than behind an optional property, cannot terminate either. ## What breaks in tooling The spec permits recursion; individual tools vary in how well they cope. **Inlining resolvers loop.** Anything that replaces a `$ref` with a copy of its target cannot finish on a cycle. A full dereference of a recursive spec either detects the cycle and bails, or expands until it exhausts memory. This is the single most common way a recursive model breaks a build, and it is why you bundle (rewrite references to local pointers) rather than inline. **Example and mock generation.** Producing a sample payload from a recursive schema means choosing a depth. Well-behaved renderers cut off after a couple of levels and show a placeholder; less careful sample generators recurse until they hang the page or emit a megabyte of nested objects. **Code generation.** Most target languages handle a self-referential type naturally — a class with a list of its own type is ordinary. Trouble concentrates where a generator flattens or synthesises names for inline schemas, or where a cycle spans multiple files and a bundler renames one participant, silently producing two types where you meant one. **Validators and diff tools.** Anything doing a recursive walk needs a visited set. Tools that lack one stack-overflow on cyclic input; the giveaway is a crash with a deep, repeating stack rather than a validation message. ## How to keep recursion safe *Name every participant in a cycle.* A cycle made of named components under `components.schemas` is easy for tools to represent as a pointer; a cycle involving inline anonymous schemas is much harder for them to name and reuse. *Make the recursive edge optional or an array.* This is what makes the contract satisfiable and simultaneously gives every generator a natural termination point. *Bound the depth in the contract's prose.* Schemas cannot express "at most five levels", so state the server's limit in the `description` and enforce it in the implementation. Unbounded nesting is also a denial-of-service surface: a client can post a deeply nested body specifically to blow up your parser. *Prefer identifiers over embedding when the depth is unbounded in practice.* Returning `parentId: string` instead of an embedded `parent` object removes the cycle entirely and usually reflects how clients consume the data anyway — they fetch what they need. This is the standard fix when recursion causes more pain than it saves. *Test the pipeline, not just the document.* If your spec is recursive, run the actual generator, the actual docs build and the actual mock server in CI. A recursive schema that validates cleanly can still be the thing that hangs your documentation job.
- How do you tell a satisfiable recursive schema from an unsatisfiable one?Follow the cycle and ask whether any hop can be omitted. If the self-reference is reached through an optional property or an array's `items`, an empty array or absent field terminates it and finite instances exist. If every property in the cycle is listed in `required`, every valid instance must contain another one, so no finite payload can ever validate.
- Your docs build hangs on a spec with a recursive Comment schema. What do you check first?Whether the pipeline is inlining rather than bundling. Full dereference expands a cycle forever; bundling rewrites external references to local pointers and leaves the cycle intact and representable. After that, check the renderer's sample-generation depth limit, since naive example builders recurse independently of how references were resolved.
- When would you remove recursion from the model rather than fix the tooling?When nesting is unbounded in practice and clients do not consume it depth-first. Replacing an embedded `parent` object with a `parentId` string removes the cycle, keeps responses small and bounded, and matches how clients actually navigate — fetching the parent when they need it. It also removes an unbounded-nesting denial-of-service surface from request bodies.
saying these in an interview costs you the question
- Claims OpenAPI forbids self-referential schemas
- Thinks any cycle is a bug to be removed
- Misses that an all-required cycle is unsatisfiable
- Blames the spec when the tool was inlining refs
- Assumes generated clients cap nesting depth for you