skip to content

What can still go wrong when constrained decoding guarantees schema-valid output?

level: seniorimportance: should knowfreq 48%

answer

  1. valid is not the same as correct
  2. check the finish reason, not the parser
  3. a truncated legal prefix is not an object
  4. no unknown member means forced invention
  5. rigid format can cost reasoning quality

basics

~20 s

Plenty. The object can be perfectly typed and factually wrong, the run can hit the token ceiling and stop mid-object, the model can be forced to invent a value it had no basis for, and rigid formatting can cost output quality.

solid answer

~50 s

Constraint guarantees form, not truth, so four failures survive. First, **wrong values in valid slots**: a yield of 250 percent or a solvent the abstract never mentions passes every schema check, which is why business-rule validation stays mandatory. Second, **truncation**: if generation hits the token ceiling mid-object, what you get is a legal prefix of the grammar, not a legal object — detect it from the length finish reason, never from the parse error, and never repair it by guessing the closing braces. Third, **forced fabrication**: masking removes the model's ability to abstain, so a field with no `unknown` member converts missing data into a confident value. Fourth, **quality cost**: forcing the output straight into a rigid payload suppresses the intermediate reasoning the model would otherwise do; how large that effect is remains contested in mid-2026, but the standard mitigations are a reasoning field ahead of the payload or a two-pass design.

go deeper

for a junior

Remember the headline: passing a schema check does not mean the values are right, and a response can stop halfway and still have obeyed the constraint at every step.

for a middle

Explain the mechanics of truncation — a legal prefix, no accepting state, EOS maskable but the budget not extendable — and why the length stop reason is the detection signal.

for a senior

Show the production controls: business-rule validation above the decoder, truncation and unknown-rate metrics, headroom in the token budget, and a designed abstention path instead of forced guesses.

for a principal

Own the risk framing: which residual failure classes the platform accepts, what audit coverage is worth paying for, and how you decide the schema-rigidity versus output-quality trade on contested evidence.

## Why this question is asked Once a team turns on a strict schema mode, parse errors vanish from the dashboards and the pipeline looks solved. The interviewer wants to know whether you understand what the guarantee actually covers, because the failures that remain are quieter and more damaging than a `JSONDecodeError`. ## Failure 1 — valid and wrong A constraint is a statement about the *language* of the output. It says the record has a `solvent`, that it is one of 34 approved names, and that `yield_percent` is a number. It says nothing about whether that solvent is the one the abstract described, or whether the number belongs to the reaction in question rather than the one in the sentence before. So the defect class shifts from malformed to plausible-but-false: yields above 100, temperatures below absolute zero, a solvent silently copied from the previous document because the current one omitted it, a field filled from the paper's introduction rather than its results. None of these trip a schema check. Range checks, cross-field consistency, referential checks against your own reference tables, and a sampled human audit are the controls that catch them, and they belong above the decoder rather than inside it. ## Failure 2 — truncation at the token ceiling This is the edge case interviewers most like, because it is where people's mental model breaks. Suppose a grammar-constrained extraction over a dense abstract runs long and hits the maximum-token limit while the object is half built. The constraint has been honoured at every single step: every token emitted was a legal continuation. But the automaton never reached an accepting state, so the output is a **legal prefix**, not a legal object — `{"solvent": "THF", "temperature": "65 °C", "yield_pe` and nothing more. The important consequences: - The constraint cannot save you. It can forbid the end-of-sequence token until the object closes, but it cannot conjure the remaining budget; when the ceiling wins, generation stops regardless. - **Detect it from the finish reason, not the parser.** A length-capped stop is a distinct signal from a validation failure, and conflating them sends a truncated record into a repair path where it does not belong. - **Never synthesise the closing braces.** Appending `"}` makes the object parse and quietly invents a value — you have manufactured a data error that now looks clean. - Mitigate at design time: bound array lengths in the schema, keep the record flat and the key names short, reserve headroom between the expected output size and the ceiling, split a large extraction into several smaller constrained calls, and alert on the truncation rate as a first-class metric. A related trap is a schema so large that the *keys themselves* dominate the budget: a 200-key record spends a substantial share of its tokens on field names, leaving less for values than anyone estimated. ## Failure 3 — the abstention path is masked away Masking removes options, and "I don't know" is an option. If the abstract never states a temperature and the schema requires a pattern-constrained string, the decoder must emit something matching that pattern. The model is not lying by choice; the honest answer was unreachable. Every constrained field therefore needs a legitimate out: a nullable type, or an explicit `not_reported` / `unknown` enum member. Then monitor its rate — that rate is one of the few cheap, high-signal quality indicators an extraction pipeline has, and a sudden drop usually means someone tightened a schema. ## Failure 4 — the quality cost of rigidity Masking renormalises the distribution over allowed tokens, and forcing the model directly into a rigid payload denies it the space to work a problem out in prose first. Published work since 2024 has reported degradation on reasoning-heavy tasks under strict formatting, while other results find the effect small when the schema is well shaped; as of mid-2026 this is genuinely unsettled, and the honest interview answer says so rather than picking a side. What is not contested are the mitigations: - put a free-text reasoning or evidence field **before** the structured fields, so the model can think inside the schema — and it doubles as an audit trail, since you can check the quoted evidence against the source, - or split into two passes: unconstrained analysis, then a constrained extraction over that analysis, - and shape the schema for compliance — flat, enum-heavy, no deep nesting or exotic composition. ## The summary an interviewer wants Constrained decoding eliminates one failure class completely and leaves three: semantic error, truncation, and forced fabrication, plus a possible quality tax. Its real value is that it converts a broad, noisy problem into a narrow one you can instrument — and instrumenting it means tracking truncation rate, unknown-value rate, and business-rule violation rate, not just parse success.

  • A constrained response comes back cut off mid-object. Why is closing the braces yourself the wrong fix?
    Because it converts a detected failure into an undetected one. The missing fields get invented — as nulls, defaults, or whatever your patch implies — and the record then passes every downstream check while carrying values the model never produced. Treat a length-capped stop as a failed call: retry with more headroom, a smaller slice, or a leaner schema.
  • How would you tell truncation apart from a genuine schema violation in production?
    By the stop reason. A length-capped generation is reported as such and should route to a budget-and-retry path; a violation of your own business rules is reported by your validator and routes to review or repair. Track the two as separate metrics — merging them hides a capacity problem inside what looks like a model-quality problem.
  • You add a free-text reasoning field before the payload to recover output quality. What does it cost?
    Tokens and latency on every call, and it competes with the payload for the same ceiling — which raises truncation risk unless you raise the budget. Cap the field's length, put it first so it is complete before the record starts, and be ready to measure whether the quality gain justifies the spend on your own eval set rather than assuming it.

saying these in an interview costs you the question

  • Treating schema-valid output as verified output
  • Repairing truncated JSON by appending closing braces
  • Detecting truncation from a parse error instead of the stop reason
  • Assuming the model can decline when no unknown value exists
  • Claiming constrained decoding has no effect on output quality

context