skip to content

How does TGI's grammar parameter force a response to match a JSON schema?

level: middleimportance: should knowfreq 45%

answer

  1. Enforced during decoding, not after
  2. A mask over the vocabulary each step
  3. Structure guaranteed, meaning is not
  4. Regex type exists for narrow formats
  5. Required fields invite invented values

basics

~20 s

TGI accepts a grammar object in a /generate request's parameters, either type json with a JSON Schema or type regex with a pattern. At each decoding step the server masks out any token that could not continue a valid match, so the output is guaranteed to parse.

solid answer

~50 s

TGI's guidance feature takes a `grammar` field inside a request's `parameters`. Two shapes exist: `{"type": "json", "value": <JSON Schema>}` and `{"type": "regex", "value": "<pattern>"}`. The schema or pattern is compiled into a state machine, and at every decoding step the server computes which tokens could legally extend the partial output and masks the rest of the vocabulary before sampling. That gives a *structural* guarantee, not a semantic one. The response will parse and conform to the schema's types and required fields; whether the field values are correct for the prompt is still ordinary model quality, and an over-constrained schema can actively push the model toward filling required fields with junk. Costs to mention: compiling a new schema is CPU work, deeply nested or heavily-alternated schemas are more expensive to mask against, and constrained decoding removes the model's option to answer with an error or a refusal.

code

bash · 17 lines
bash
curl 127.0.0.1:8080/generate -X POST -H 'Content-Type: application/json' -d '{
  "inputs": "Extract the person from: Ada Lovelace, age 36, London.",
  "parameters": {
    "max_new_tokens": 64,
    "grammar": {
      "type": "json",
      "value": {
        "type": "object",
        "properties": {
          "name": {"type": "string"},
          "age": {"type": "integer"}
        },
        "required": ["name", "age"]
      }
    }
  }
}'

go deeper

for a junior

Know that TGI can force a response into a JSON shape by passing a grammar in the request parameters, and that this beats asking for JSON in the prompt and hoping.

for a middle

Explain the mechanism as per-step masking of illegal tokens before sampling, and state clearly that the guarantee is structural - the response parses, but the values can still be wrong.

for a senior

Discuss schema design as a hallucination control: nullable fields and explicit unknown values so the model can decline, plus the compile and per-step masking costs of deep or heavily-alternated schemas.

for a principal

Weigh constrained decoding against alternatives at the system level - retry loops, a smaller extractor model behind a cascade, or validation in the client - and account for tokens spent on structure in cost per request.

## What the feature is Most applications that call an LLM want a data structure back, not prose. The naive approach - asking politely in the prompt and parsing with a retry loop - fails a few percent of the time, and those failures cluster exactly where you least want them: long outputs, unusual inputs, models under quantization. TGI's guidance feature removes the failure mode at the decoding layer instead. You pass a `grammar` object in the request's `parameters`: - `{"type": "json", "value": { ...JSON Schema... }}` constrains output to that schema. - `{"type": "regex", "value": "..."}` constrains output to a regular expression, which is the right tool for narrow formats - a date, an enum, a phone number, a single classification label. ## The mechanism: masking, not retrying The important mental model is that this happens *inside* generation, once per token, not after generation as validation. 1. The schema or regex is compiled into a finite state machine over the model's tokenizer vocabulary. 2. At each step the server knows the current state - say, 'we have emitted `{"name": "` and are inside a string value'. 3. It computes the set of vocabulary tokens that could legally extend from that state and sets the logits of every other token to negative infinity. 4. Sampling proceeds normally over what remains. Because an illegal token can never be selected, there is no such thing as a malformed response to recover from. Contrast the two alternatives candidates usually propose: prompt-and-hope (fails stochastically), and generate-then-validate-then-retry (doubles latency and cost on every failure, and can loop). Masking pays a small constant cost per step and cannot fail structurally. ## What it does *not* guarantee This distinction separates a middle answer from a strong one. - **Types and shape: guaranteed.** Required fields will be present, an integer field will hold an integer, an enum field will hold one of its members. - **Truth: not guaranteed.** A schema demanding `"invoice_total": number` will get a number even when the document had no total. The constraint has removed the model's ability to say so. - **Quality can drop.** If the schema fights the model's natural output order or forces fields it has no evidence for, you are pushing probability mass toward tokens the model considers unlikely. The mitigation is designing schemas that let the model opt out: nullable fields, an explicit `"unknown"` enum member, a `confidence` field, or a top-level union between a result object and an error object. ## Costs and operational notes Compiling a schema into its state machine is CPU work on the router side. A workload that sends a handful of stable schemas amortises this quickly; one that generates a fresh schema per request pays it repeatedly, and pathological schemas - deep nesting, many alternations, unbounded string patterns - make both compilation and the per-step legality check more expensive. Keep schemas as flat and as bounded as the task allows. Also remember the constraint is per-request, not per-server: two clients hitting the same TGI deployment can use different grammars, and requests without a `grammar` field are unaffected. And a grammar interacts with your token budget - a verbose schema with long field names spends output tokens on structure, which is real money at scale and real latency for the user. ## What interviewers listen for The strong answer names the layer: *token masking during decoding*, not post-hoc validation. Then it draws the structural-versus-semantic line without prompting, and closes with schema design as the mitigation for forced-field hallucination. Candidates who describe it as 'the server retries until the JSON is valid' have the wrong mechanism entirely, and that misconception predicts bad decisions about latency and cost.

  • Why can a strict required-fields schema make hallucination worse rather than better?
    Because the mask removes every token that would let the model decline. If the document has no age and the schema requires an integer age, the model cannot emit null, an apology, or an error - it must emit digits, so it invents them. The fix is schema design: make uncertain fields nullable, add an explicit unknown enum member, or model the response as a union of a result object and an error object.
  • When would you use the regex grammar type instead of the json type?
    When the output is a single narrow-format value rather than a structure: a classification label from a fixed set, an ISO date, an identifier pattern. A regex is cheaper to compile and check than a schema, and it avoids spending output tokens on braces and field names for a value that has no structure worth expressing.
  • What performance cost does constrained decoding add per request?
    Two things: a one-off CPU cost to compile the schema or pattern into its state machine, and a small per-step cost to compute the legal-token mask before sampling. Both are usually small next to the GPU forward pass, but deeply nested schemas with many alternations make the per-step check meaningfully more expensive, and a workload that sends a unique schema on every request never amortises the compile.

saying these in an interview costs you the question

  • Believing the server retries until JSON parses
  • Claiming grammar guarantees the values are correct
  • Treating it as a prompt-engineering trick
  • Assuming it must be enabled with a launcher flag for every request
  • Ignoring that required fields force invented values

context