When does forcing a strict JSON schema on an LLM hurt output quality?
answer
- structure guarantees shape, not truth
- what does the schema forbid?
- required non-nullable invites fabrication
- closed enum forces the nearest wrong label
- reasoning field before the answer field
basics
~20 sWhen the schema makes truth unrepresentable or removes reasoning room. Required non-nullable fields push the model to fabricate, a closed enum forces a wrong nearest label, and constraining from the first token denies it space to think before committing.
solid answer
~50 sConstrained generation guarantees shape, and shape can fight content in three ways. First, **unrepresentable truth**: if `price_cents` is required and non-nullable, a menu item with no listed price has no honest encoding, so the model invents a plausible number — the schema converted a missing value into a confident lie. Second, **forced choice**: a closed enum with no `other` or `uncertain` member turns "none of these fit" into the nearest wrong label, and your accuracy metric will not show it. Third, **no room to think**: decoding is sequential, so a schema that opens with the answer field makes the model commit before reasoning; putting a free-text reasoning field first, or splitting reasoning and formatting into two calls, usually recovers the loss. Add token cost — a fourteen-field schema with long enums is paid on every request. The fix is schema design, not abandoning structure: make absence and uncertainty first-class, then measure per-field accuracy rather than parse rate.
code
json · 10 lines{
"type": "object",
"properties": {
"reasoning": { "type": "string" },
"price_cents": { "type": ["integer", "null"] },
"price_status": { "enum": ["listed", "market_price", "not_stated"] }
},
"required": ["reasoning", "price_cents", "price_status"],
"additionalProperties": false
}go deeper
Know that a response matching the schema can still contain made-up values, and that a field which must always be filled invites the model to guess rather than admit it does not know.
Explain the mechanics: constrained decoding masks invalid tokens, so a required non-nullable field leaves fabrication as the only legal move, and a closed enum forces the nearest label onto an out-of-space input.
Demonstrate the design response — nullable fields paired with an explicit status, an escape member in every real-world enum, a reasoning field placed first — and the measurement that replaces parse rate once constraints make it meaningless.
Own where each surface sits on the constraint spectrum and why: the token and latency cost of large schemas, the coupling between schema and validator, when to split reasoning from formatting, and what audit sampling proves the pipeline is not manufacturing values.
## The premise everyone gets half right Structured output is close to free reliability on the *syntactic* axis: constrain generation and malformed responses vanish. The mistake is inferring that tighter is therefore always better. A schema is a specification of what the model is permitted to say, and a specification that cannot express the truth will be satisfied by something untrue. At principal level, the interesting question is not "how do I get valid JSON" but "what does my schema forbid, and what does the model do when reality lands there". ## Failure one: absence is not representable Take a menu-extraction service with a fourteen-field schema. `price_cents` is marked required, integer, non-nullable — because whoever wrote the schema was thinking about the happy path where menus list prices. Then a menu says "market price", or the price column was cropped from the photo. The decoder must emit an integer. The model's options are to emit a number it has no evidence for, or to emit a number it has no evidence for. It fabricates, confidently, and the value flows into your database indistinguishable from a real one. The fix is to make absence expressible: a nullable type, or better, an explicit status field (`listed` / `market_price` / `not_stated`) so a null carries a reason. Note that nullability alone is weaker than it looks — a model will still prefer a plausible value over null unless the prompt makes clear that null is the *expected* answer under uncertainty. The schema opens the door; the instruction has to say walking through it is fine. ## Failure two: forced choice in a closed enum A closed enum is a strong tool and a sharp one. Constraining a triage tag to four values guarantees you never see a fifth, which is exactly what you wanted — until an input genuinely belongs to a fifth category. The model cannot say so; the sampler will not let it. It picks the nearest allowed label, and because the output is structurally perfect, nothing anywhere in your stack flags it. The category error is now silent and permanent. So any enum that models a real-world classification needs an escape member — `other`, `unclear`, `needs_review` — and a paired instruction about when to use it. Then monitor its rate: a rising `other` rate is a live signal that your taxonomy has drifted away from your input distribution, which is one of the few genuinely useful automatic alerts in an LLM pipeline. Removing the escape hatch does not remove the ambiguity; it removes your ability to see it. ## Failure three: constraint removes reasoning room Decoding is sequential, and every token is conditioned on the tokens before it. Under an unconstrained prompt, a model asked a hard extraction question will often work through the input before answering. Under a schema whose first key is the answer, it must produce the answer immediately — there is nowhere to put the intermediate work. On easy tasks this costs nothing; on tasks with real ambiguity it can cost meaningfully. Two remedies. Put a free-text field *first* in the schema so the reasoning is generated before the constrained fields and conditions them. Or split into two calls: one unconstrained reasoning pass, one cheap formatting pass whose input is that reasoning. The two-call version costs more and is easier to evaluate, since you can inspect the reasoning independently; the single-call version is cheaper and keeps everything in one context. There is no universal winner — it depends on how hard the extraction is and what your latency budget looks like. ## The costs that do not show up as errors A large schema is serialized into every request: fourteen fields with descriptions and long enums can be a substantial fixed token cost per call, plus context you are not spending on the input. Grammar compilation adds latency on first use. Deeply nested or recursive schemas are often unsupported or poorly supported by structured-output modes, pushing you toward flatter designs than a data modeller would choose. And a schema shared between the request and the validator is a coupling point — change it in one place only, or you get the worst outcome, where generation and validation disagree. ## What to measure Parse rate is the metric everyone has and the one that stops being informative the moment you adopt constrained decoding, because it pins to a hundred percent by construction. Replace it with per-field accuracy against a labelled set, per-field null and escape-hatch rates, and a fabrication check on the fields where invention is plausible — spot-audit a sample of non-null values against the source. If you cannot tell the difference between a schema that is working and one that is quietly manufacturing values, you do not have structured output, you have structured confidence. ## The strategic call The judgment to own is where on the constraint spectrum each surface should sit. A high-volume extraction pipeline feeding a database wants maximum structure with generous escape hatches and sampled audits. A hard analytic task with a human reviewing every output may be better served by loose structure and rich prose. Deciding that per-surface, and writing down who owns each schema and how it changes, is the part of this that does not delegate.
- If you add an "other" member to every enum, what do you now have to operate?A monitored rate. The escape hatch is only useful if someone watches how often it fires and acts on it: a rising `other` rate means the taxonomy has drifted from the input distribution, and a near-zero rate on an ambiguous domain means the model is avoiding it and forcing choices anyway. Sample and review the `other` cases periodically, and treat the rate as a product signal rather than noise.
- Why does parse rate stop being a useful metric once you adopt constrained decoding?Because it is pinned to a hundred percent by construction — the decoder cannot emit anything that fails to parse. All the remaining failure has moved into the values: fabricated numbers, forced labels, fields filled with null because the model gave up. Replace it with per-field accuracy on a labelled set, null and escape-hatch rates, and sampled audits of non-null values against the source.
- When would you split reasoning and formatting into two separate model calls?When the task is genuinely hard, when you want the reasoning inspectable and evaluable on its own, or when the formatting pass can run on a cheaper model. The costs are extra latency and an extra call, plus a handoff where the formatter can misread the reasoning. For easy extraction, a reasoning field first in a single constrained response is usually enough and much cheaper.
saying these in an interview costs you the question
- Assuming a schema-valid response is a correct response
- Marking every field required so nothing is ever missing
- Closing an enum with no option for genuinely ambiguous inputs
- Reporting parse rate as the quality metric under constrained decoding
- Placing the answer field before any reasoning field in the schema