In LLM structured output, why enforce a field's enum at decode time rather than by prompt instruction?
answer
- instruction shifts odds, constraint removes options
- non-members are unreachable, not just rare
- one canonical spelling instead of three
- exact match: casing is never corrected
- always give the enum an unknown member
basics
~10 sDecode-time enforcement makes non-members unreachable: the sampler is only offered tokens that continue an allowed value, so a variant spelling cannot be produced. A prompt instruction only shifts probabilities, so violations survive at volume.
solid answer
~50 sThey are guarantees of different kinds. Telling the model "use one of these 34 approved solvent names" raises the probability of compliance; run 12,000 chemistry abstracts through it and a small per-call violation rate still yields hundreds of unusable rows, clustered on the hardest documents. A decode-time enum compiles the allowed values into a state machine over the tokenizer's vocabulary and masks every token that cannot continue one of them, so if only `DCM` is listed, `dichloromethane` is not unlikely — it is unreachable. That kills surface-form collapse, the everyday nuisance where the same solvent arrives under three spellings and your grouping breaks. A regex constraint does the same for formats: pin a temperature field to `^\d+(\.\d+)?\s?(°C|K)$` and "room temperature" can never land there. The catch is that the constraint guarantees the value is *allowed*, not correct, so always give the enum an explicit unknown member rather than forcing a guess.
code
json · 16 lines{
"type": "object",
"properties": {
"solvent": {
"type": "string",
"enum": ["DCM", "THF", "toluene", "methanol", "not_reported"]
},
"temperature": {
"type": "string",
"pattern": "^\\d+(\\.\\d+)?\\s?(°C|K)$"
},
"yield_percent": { "type": "number" }
},
"required": ["solvent", "temperature", "yield_percent"],
"additionalProperties": false
}go deeper
Know the difference in one sentence: a prompt makes the right value likely, a decode-time enum makes every other value impossible. Mention that normalising spellings is the everyday payoff.
Explain that enum members compile into a token-level state machine, that matching is exact, and that a regex pins format only — plus why every constrained field needs a legitimate unknown value.
Judge when constraint is wrong: open or runtime-varying sets, where a stale enum manufactures errors. Describe monitoring the field's value distribution to detect an enum that no longer fits the corpus.
Own the vocabulary itself — who curates the canonical enum, how changes are versioned across pipelines, and where normalisation should live so that extraction and downstream analytics cannot disagree.
## Instruction versus impossibility A prompt instruction is evidence the model weighs against everything else in its context. It usually wins, and that is the trap: a 98 percent compliance rate looks excellent in a notebook and is a data-quality incident at 12,000 documents. Worse, the 2 percent is not random — it concentrates on the ambiguous, verbose, or unusual inputs, which are exactly the rows a reviewer most needs to be right. A decode-time constraint is a different category of statement. The set of producible outputs for that field shrinks to the enum, so the violation rate is zero by construction, not by luck, and it does not drift when the prompt is edited or the model is swapped. ## What the decoder does with an enum The enum members are compiled into a small state machine over the tokenizer's vocabulary — effectively a prefix tree of the token sequences that spell each member. When generation reaches that field, only tokens that extend some member's remaining prefix are left unmasked; everything else is driven to zero probability. Once enough tokens have been emitted to disambiguate, the rest of the member is forced, and in many implementations fast-forwarded without further model calls. The field therefore cannot hold a value outside the list. It can still hold the *wrong* value from the list, which is where validation and audit continue to matter. ## Surface-form collapse is the practical win In an extraction over organic-chemistry abstracts, the same solvent appears as `DCM`, `dichloromethane`, `CH2Cl2`, and occasionally `methylene chloride`. Each is a correct reading of the paper and each breaks a `GROUP BY solvent`. A canonical enum forces one path: whatever the abstract says, the field can only contain the canonical spelling. You have moved normalisation from a downstream cleanup script — which needs a synonym table nobody maintains — into the decoder, where it cannot be skipped. ## Regex constraints pin format, not meaning The same argument applies to shape. A temperature field constrained to `^\d+(\.\d+)?\s?(°C|K)$` accepts `25 °C`, `25.5°C`, and `298 K`, and can never accept `room temperature`, `reflux`, or `ambient`. Downstream parsing becomes total: every value in the column converts. Two cautions: not every strict-schema dialect supports string patterns (a full grammar constraint does), and a permissive regex can compile to a large automaton, so keep patterns tight and anchored. ## Tokenization details that bite Because the constraint is enforced over tokens, not characters, a few things surprise people: - Enum members rarely align to token boundaries; the compiler handles that, but it means the mask is over *prefixes*, and a member that shares a long prefix with another only diverges late. - Matching is exact. `thf` will never be produced when the enum says `THF`, but neither will it be corrected — decide the casing you want and put that in the enum. - Non-ASCII characters such as `°` may be several tokens; that is fine, but it makes hand-reasoning about the automaton harder, so verify with a real decode rather than by eye. ## Where a constraint is the wrong tool Constrain a closed, stable set. Do not constrain a set that is really open or changes at runtime — customer IDs, catalogue codes pulled from a live table, free-text titles. Pinning those to a stale list does not prevent an error, it *manufactures* one: the model is forced to emit a wrong member because the right one is unreachable. That failure has a general form worth naming: **the constraint removes the model's escape hatch**. If the abstract genuinely does not state the solvent and the enum has no way to say so, the decoder must still pick something. Every constrained field needs a legitimate out — an explicit `unknown` or `not_reported` member, or a nullable type — otherwise you have converted honest silence into a confident fabrication. ## Knowing it works Monitor the per-field value distribution. A healthy extraction shows a long-tailed but plausible spread plus a modest `not_reported` share. A spike in one member, or a collapse toward `unknown`, means the enum no longer matches the corpus — the signal that a constraint has become a straitjacket.
- The abstract genuinely never states the solvent. What does a constrained enum do, and what should you have done?With no escape hatch it must emit some listed member, turning a missing value into a fabricated one. The fix is to design the out into the constraint: add an explicit `not_reported` member, or make the field nullable, so the honest answer is reachable. Then monitor that member's share — a rising rate is real signal, not noise.
- Would you constrain a field to a list of catalogue codes read from your production database?Only if the list is small, stable, and passed in with the request. An enum baked into a schema goes stale the moment the catalogue changes, and a stale enum forces a wrong code rather than surfacing a miss. For large or volatile identifier sets, leave the field a plain string and validate it against the live table afterwards.
- How does a regex constraint interact with the schema-subset limits of strict modes?String `pattern` support varies by dialect, so a regex you rely on may be silently unsupported or rejected. Where it is unavailable, the options are to express the format as a full grammar constraint instead, or to keep the field unconstrained and enforce the pattern in the validation layer — accepting that you are back to detecting rather than preventing.
saying these in an interview costs you the question
- Saying a firmly worded prompt gives the same guarantee
- Believing a near-miss synonym still slips through masking
- Constraining volatile identifier sets to a fixed enum
- Omitting an unknown member and forcing a guess
- Expecting lowercase input to be corrected to the enum's casing