When would you not put a strict JSON schema on an OpenAI call, and what do you use instead?
answer
- strong defaults still have edges
- prose does not want a wrapper
- cached only if the schema repeats
- the dialect has hard caps
- several providers, several dialects
basics
~20 sSkip strict schemas when the output is prose, when the shape is decided per request so every call pays fresh grammar compilation, when the schema cannot fit the restricted dialect or its size caps, or when portability across providers outweighs the guarantee.
solid answer
~50 sStrict Structured Outputs is the right default for known, stable shapes, but it is not free. Four situations argue against it. **Prose deliverables** — wrapping an essay in a one-field object buys nothing and can cost quality. **Dynamic schemas**, generated per user or per tenant: a previously unseen schema pays a compile step before the first token, so a long tail of one-off schemas turns a cached cost into a per-request one. **Schemas the dialect cannot hold** — very deep or very wide structures hit nesting, property-count and size caps, and cross-field business rules were never expressible anyway; there the answer is to split the work into staged calls and validate semantics in code. **Portability**, when the same logical call must run against several vendors whose strict dialects differ; a thinner contract plus your own validator may be cheaper than maintaining per-vendor schemas. The fallback ladder is: strict schema, then a schema-shaped prompt with json_object and validation, then plain text.
go deeper
Know that a strict schema is the normal choice when you know the output shape, and that plain text is fine when the answer is prose a person will read.
Be able to name the concrete costs — a restricted dialect, size and nesting caps, extra tokens because every field must be emitted — and say that business rules still need code-level validation.
Show operational judgment: decompose oversized schemas into staged calls, watch first-token latency when schemas are dynamic, and design nullable or unknown options so the model can be honest about missing data.
Own the ladder and the policy behind it — where response shapes are defined and versioned, how failures are classified and handled, and what multi-provider routing costs in schema maintenance versus a thinner shared contract.
## Framing the decision Strict Structured Outputs converts a probabilistic formatting problem into a deterministic one, and for most production calls that is an obvious trade. The principal-level question is where the trade stops paying, because the costs are real: a restricted schema dialect, a compile step for unseen schemas, extra output tokens for fields that must always be emitted, and coupling to one vendor's dialect. ## Case 1 — the output is prose If the deliverable is a paragraph, a draft email, or an explanation, forcing it through `{"text": "..."}` adds structure tokens, invites escaping problems, and constrains the decoder for no benefit. Use plain text. The exception is when prose travels alongside metadata the caller genuinely branches on — a confidence label, a category — and then the schema earns its place because of the metadata, not the prose. ## Case 2 — the schema is dynamic A new schema must be compiled into a grammar before decoding can start, and that compiled form is cached. Steady-state traffic against a fixed set of schemas amortises this to nothing. A product where each tenant defines their own extraction shape, or where a schema is assembled per request from user-selected fields, defeats the cache: many schemas are seen once. Mitigations before abandoning strict mode: canonicalise schema generation so semantically identical shapes serialise identically and share a cache entry; collapse a family of near-identical schemas into one superset schema with nullable fields; or warm the common shapes. ## Case 3 — the schema does not fit Two distinct limits. **Structural**: nesting depth, total property count and overall schema size are capped, and a schema that models an entire domain object graph will hit them. The fix is almost always decomposition — several narrow calls, each returning a small object, orchestrated by your code — which usually improves quality too, because a model asked for eight fields does better than one asked for eighty. **Semantic**: the dialect covers structure and a subset of type-level validation, not invariants. "The total equals the sum of the line items," "the end date follows the start date," "this identifier exists in our catalogue" — none of these are schema-expressible at any level of strictness. That is not an argument against strict mode; it is an argument that strict mode is one layer of a validation stack, never the whole of it. ## Case 4 — portability Several vendors offer schema-constrained output, and the accepted dialects and enforcement guarantees are not identical. If a call must run against more than one provider — for cost routing, redundancy, or regional requirements — you are choosing between maintaining a per-vendor schema mapping and adopting a thinner shared contract that every provider can honour, with your own validator downstream. Neither is wrong; the decision depends on how many providers, how stable the shape, and how expensive a malformed response is. What is wrong is discovering the difference during a failover. ## Case 5 — the schema is fighting the task A subtler one. Over-constraining can degrade the answer. If the schema has no way to say "absent from the source," a required field forces a fabrication; give it a nullable type or an explicit unknown member. If the conclusion field is declared before the working, the model commits to an answer and then narrates support for it, because decoding runs left to right and property order is generation order. Both are schema-design bugs that present as model-quality complaints, and both are worth naming in an interview because they show you have debugged this rather than only configured it. ## The fallback ladder 1. **Strict `json_schema`** — known, stable shape; the default. 2. **`json_object` plus a schema-shaped prompt and your own validation** — dynamic or unsupported shapes; you accept repair-and-retry as the cost of flexibility. 3. **Plain text** — prose deliverables and anything a human reads directly. Across all three, the invariant is that your own validation layer never goes away. Strict mode changes how often it fires, not whether you need it. ## Organisational angle Decide once, centrally, where response shapes are defined and how they version, so a shape change is one edit with type errors pointing at every consumer. Decide the failure policy — retry, dead-letter, human review — per outcome class rather than per call site, because refusals, truncations and validation failures want different treatment. And instrument the compile-cost tail if you are anywhere near dynamic schemas: it shows up as first-token latency, not as an error, so nothing alerts on it unless you look. ## Interview framing Answer with the ladder and one concrete cost per rung. The failure mode to avoid is dogma in either direction: "always use strict schemas" ignores dynamic and portability cases; "schemas are overkill, just parse it" ignores that repair heuristics are the thing this feature deleted.
- Your product lets each tenant define their own extraction fields. How do you keep strict mode viable?Canonicalise generation so identical shapes serialise byte-identically and share a compiled-grammar cache entry, and collapse near-identical tenant schemas into one superset with nullable fields. If the tail is still mostly one-off shapes, accept json_object plus your own validator for that tail and keep strict mode for the common shapes.
- A schema that models your whole domain object exceeds the size and nesting caps. What is the fix?Decompose the call. Several narrow calls, each returning a small object your code assembles, stay inside the caps and usually improve quality, because a model asked for eight fields outperforms one asked for eighty. Fan out where the sub-extractions are independent, and sequence only where a later call genuinely needs an earlier result.
- How do you decide whether to maintain per-vendor schemas or one thin shared contract?Weigh how many providers you actually route to, how often the shape changes, and what a malformed response costs. Few providers and a stable shape favour per-vendor schemas with a shared generator. Many providers, or a shape in flux, favour a thin contract plus one validator you own — accepting more repair work in exchange for one definition.
- Where does a strict schema make the answer worse rather than better?When it removes a truthful option or reorders reasoning. A required non-nullable field forces a value even when the source is silent, producing a fabrication; and because generation follows property order, declaring a conclusion before its supporting analysis means the analysis is written to justify a committed answer. Both read as model-quality problems but are schema bugs.
saying these in an interview costs you the question
- Insists every call should carry a strict schema
- Ignores compile cost for schemas seen only once
- Assumes strict dialects are identical across vendors
- Wraps prose deliverables in a single-field object
- Treats the schema as the entire validation layer