How does logit masking during decoding make invalid JSON impossible rather than unlikely?
answer
- enforcement sits in the sampler, not the parser
- allowed-token set per generation step
- automaton state maps to a token mask
- logits set to negative infinity before softmax
- a validator only sees the finished string
basics
~20 sThe decoder tracks its position in a state machine compiled from the schema and, at every step, drives the logits of all tokens that cannot legally continue to negative infinity. An illegal token has zero probability, so it is never sampled.
solid answer
~50 sConstrained decoding moves enforcement into the sampling loop. The schema or grammar is compiled once into an automaton — a finite-state machine for flat patterns, a pushdown automaton when nesting has to be balanced — whose states are positions in the target language. Before each token is sampled, the runtime looks up the set of vocabulary tokens that can legally continue from the current state and masks the rest by setting their logits to negative infinity; softmax then renormalises over what remains. Because an illegal token has probability zero, malformed output is not improbable, it is unreachable, and validity does not decay with output length the way a prompt-instructed format does. A post-hoc validator can only observe a finished string: it detects the error after you have paid for the whole generation, and the probability of at least one violation grows with every token. The trade is rigidity — the model cannot step outside the grammar even when it should.
code
python · 13 linesimport math
VOCAB = ["{", "\"", "a", "}", "1"]
ALLOWED = {"start": {"{"}, "in_object": {"\"", "}"}}
def masked_logits(logits, state):
return [
logit if token in ALLOWED[state] else -math.inf
for token, logit in zip(VOCAB, logits)
]
print(masked_logits([0.1, 0.9, 2.0, 0.5, 1.0], "start"))
print(masked_logits([0.1, 0.9, 2.0, 0.5, 1.0], "in_object"))go deeper
Recall the one-line mechanism: tokens that would break the format are blocked before sampling, so invalid output cannot be produced rather than merely being rare.
Explain the pipeline — schema compiled to an automaton, state tracked as tokens are appended, allowed-token set looked up, illegal logits set to negative infinity — and contrast it with validating a finished string.
Show where the guarantee ends: no cross-field or uniqueness enforcement, no notion of correct values, and a real quality question when the payload leaves no room for reasoning. Say how you would leave that room.
Frame it as moving an invariant from the application layer into the runtime, and reason about what that centralisation buys across many teams versus the rigidity and vendor coupling it introduces.
## Where the constraint lives An unconstrained decoder loops: run a forward pass, get a logit per vocabulary token, apply the sampling policy, append the chosen token, repeat. Constrained decoding inserts one step between logits and sampling — a mask. Everything else about the model is untouched; no weights change and no extra forward passes occur. ## Compiling the schema into an automaton The schema or grammar is first turned into a recogniser for the target language. A regex-shaped constraint compiles to a finite-state machine. JSON needs more than that, because braces and brackets must balance to arbitrary depth, so implementations use a pushdown automaton — an FSM plus a stack — or an equivalent structure that records how many objects and arrays are open and which key is being filled. A state is a position in the language: *before the first key*, *inside a string*, *after a colon, expecting a number*, *inside an array, having just closed an element*. From each state, only some continuations are legal. ## From states to token masks The subtlety is that the automaton is defined over characters while the model samples over tokens, and a token is usually several characters. So the compiler builds an index mapping each automaton state to the set of vocabulary tokens whose character expansion keeps the document inside the language — a token may be legal, illegal, or legal-and-advance-the-state-by-several-transitions. That index is the expensive artefact, which is why it is compiled once per schema and cached. At each step, the runtime looks up the allowed set for the current state and applies the mask: illegal logits become negative infinity, so after softmax their probability is exactly zero. Sampling then proceeds normally over the surviving distribution. Because the model's relative preferences among the *legal* tokens are preserved, the model still chooses the content; it just cannot leave the language. Two consequences follow. First, the end-of-sequence token is itself maskable: while a required field is still unfilled, EOS can be forbidden, which is how required keys get enforced. Second, when only one continuation is legal for several steps — the fixed characters of a known key name, say — some implementations fast-forward those tokens without a forward pass at all, which makes constrained generation *faster* than free generation on schema-heavy output. ## Why a validator cannot give the same guarantee A validator sees a finished string. Three differences matter: - **Timing.** The violation is detected after the full generation is paid for, in tokens and in latency. - **Compounding.** Under a prompt instruction, each token carries a small chance of leaving the format; over hundreds of tokens those chances accumulate, so long outputs fail more often than short ones. A mask removes the per-token chance entirely, so validity stops depending on length. - **Reachability.** The validator's verdict is a filter over what the model produced; the mask changes what the model *can* produce. Only the second makes a class of output impossible. The honest framing in an interview: a validator narrows the outcome distribution after sampling, a mask narrows the support before sampling. ## What masking cannot enforce The automaton is a streaming recogniser with local knowledge, so it enforces exactly what the language can express positionally: - it cannot compare two fields ("end date must be after start date"), - it cannot enforce uniqueness across array elements or global counts beyond what the grammar encodes, - it cannot know whether a value is *true*, only whether it is *allowed*, and - it cannot repair a state it should never have entered — once a wrong-but-legal token is sampled, the constraint dutifully keeps the output legal around the mistake. There is also a quality argument to be honest about: masking renormalises the distribution, and forcing the model straight into a rigid payload can suppress the intermediate reasoning it would otherwise produce. The usual mitigation is to make room for reasoning inside the schema — a free-text field emitted before the structured fields — or to run reasoning unconstrained and extract in a second constrained pass. ## Implementation vocabulary worth having Grammar-constrained decoding is standard in open serving stacks — GBNF grammars in llama.cpp, regex-to-FSM indexing in Outlines, adaptive token-mask caching in XGrammar, jump-forward decoding in SGLang — and is what sits behind providers' strict schema modes. Knowing that the mechanism is a mask over logits, not a retry loop or a post-processor, is the point interviewers check.
- How does masking enforce that a required key is present, given the model could just stop early?By masking the end-of-sequence token. While the automaton is in a state where the document is incomplete — a required key unwritten, an object unclosed — EOS is not a legal continuation, so it is driven to negative infinity along with every other illegal token. The model cannot terminate until the automaton reaches an accepting state.
- Why can constrained decoding sometimes be faster than unconstrained decoding, not slower?Two reasons. Structured output is shorter — no preamble, no explanation, no markdown fence — so fewer tokens are generated. And when the automaton has exactly one legal continuation for a stretch (the fixed characters of a known key, a closing brace), implementations can fast-forward those tokens without a forward pass. On schema-heavy output these can outweigh the per-step mask cost.
- Does masking change the model's preferences among the legal tokens?Not their relative ordering. Illegal tokens go to zero and the remaining probabilities are renormalised, so the model's ranking among allowed continuations is preserved. What does change is the overall distribution: probability mass the model wanted to spend outside the grammar is redistributed inside it, which is the mechanism behind reported quality effects on reasoning-heavy tasks.
saying these in an interview costs you the question
- Describing it as generate-then-retry until it parses
- Claiming the model is fine-tuned to obey the schema
- Saying it post-processes or rewrites sampled tokens
- Believing a mask can enforce cross-field rules
- Assuming validity degrades with output length under a mask