skip to content

Output Formatting

Getting output your code can parse: format instructions, an explicit schema, delimiters around fields, and keeping the shape stable as the context grows long. This comes up whenever an LLM sits inside a pipeline rather than in front of a human.

part ofPrompt engineeringoverview, primer and where to startread it →
on this pageshow

questions

5

In an LLM prompt, why is "respond only with JSON" unreliable, and what beats it?

level: middleimportance: must knowfreq 78%

answer

  1. instruction versus constraint
  2. the model can still say hello
  3. three levels of control
  4. masking tokens at decode time
  5. shape guaranteed, values not

basics

~20 s

A format instruction only biases the model; every token stays available, so it can still emit a preamble, a ```json fence, or a truncated object. Constraining decoding — a schema-based structured-output mode or a grammar — makes invalid output impossible rather than unlikely.

solid answer

~50 s

Format instructions are a bias, not a constraint. "Reply with JSON only" leaves every token available, so you still get a ```json fence, a "Sure, here you go" preamble, a trailing comma, or an object cut off by the output-token limit — rare enough to pass a demo, frequent enough to page you at volume. Think in three levels of control. Weakest is **wording**: state the format, show the exact shape, forbid prose. Stronger is a **schema-based structured-output mode**, where you hand the provider a JSON Schema and generation is constrained to it. Strongest is **grammar-constrained decoding**, which masks any token that cannot continue a valid parse, so malformed output cannot be produced. Use the strongest level your stack supports — and still parse and validate, because none of them makes the *values* right, only the shape.

code

json · 13 lines
json
{
  "type": "object",
  "properties": {
    "dish_name": { "type": "string" },
    "price_cents": { "type": ["integer", "null"] },
    "allergens": {
      "type": "array",
      "items": { "enum": ["milk", "eggs", "peanuts", "tree_nuts", "soy", "wheat", "fish", "shellfish"] }
    }
  },
  "required": ["dish_name", "price_cents", "allergens"],
  "additionalProperties": false
}

go deeper

for a junior

Know that asking for JSON is a request, not a guarantee, and that real responses often arrive wrapped in a markdown fence or preceded by a sentence. Always parse defensively instead of assuming the string is clean.

for a middle

Be able to lay out the three levels — wording, schema-based structured output, grammar-constrained decoding — and explain that the last one masks invalid tokens at decode time, so malformed output becomes unreachable rather than improbable.

for a senior

Show the production judgment: pick the strongest level the serving stack supports, share one schema between request and validator, and instrument parse and validation failure rates so a model or prompt change surfaces as a metric rather than a support ticket.

for a principal

Own the tradeoff between constraint and quality. Argue when the schema token cost and the reasoning room it removes outweigh the reliability gain, and set the org-wide contract: where structure is enforced, who owns the schema, and how model upgrades are gated against it.

## What a format instruction actually is When you write "Respond with a single JSON object and nothing else", you have added tokens to the prompt. That shifts the model's probability distribution toward JSON-shaped output; it removes nothing from the vocabulary. Every production failure follows from that fact. The model remains free at each decoding step to emit a friendly preamble, to wrap the object in a ```json fence because that is how JSON appears throughout its training data, to append "Let me know if you'd like me to adjust anything", to emit a trailing comma, or to hit the output-token cap halfway through the object. Under easy inputs it complies almost always — which is exactly the trap. A one-percent failure rate is invisible when you eyeball ten outputs and constant when a menu-ingestion pipeline processes thousands of restaurants a night. ## Level 1 — instruction wording This is where most people start, and it does help. Effective wording is concrete rather than emphatic: give the literal shape you want (an example object with the exact keys), name the field types, say explicitly "no markdown fences, no commentary before or after", and specify what to emit when a value is absent (`null`, not omission). Wording alone is the only lever available when you have no control over the serving stack, and it is worth doing well even when you do — the model still has to choose values, and a clear shape spec reduces semantic mistakes as well as syntactic ones. But it is probabilistic. Two things reliably degrade it: long inputs that push the instruction far from the generation point, and unusual inputs that make the model want to explain itself. ## Level 2 — schema-based structured output Most providers expose a mode where you supply a JSON Schema (sometimes via the same mechanism used for tool arguments) and the response is constrained to conform. Practically this eliminates the fence, the preamble and the malformed-object class of failures in one move, because the surrounding prose is simply not part of what the model is allowed to produce. It also documents the contract in a machine-readable artifact your validator can reuse, so the prompt and the parser cannot drift apart. Constraints are real: schemas typically support a subset of JSON Schema (deeply recursive or exotic constructs are often unsupported), the schema costs tokens on every request, and enumerating a large enum inflates it further. ## Level 3 — grammar-constrained decoding The strongest level operates on the sampler. At every step the decoder computes which tokens could still continue a string in the target language — described by a grammar, a regular expression or a compiled schema — and masks the logits of all the others to negative infinity. Invalid output is then not unlikely, it is unreachable. This is how local runtimes and inference servers implement structured generation, and it generalizes beyond JSON: you can constrain to a CSV row shape, a specific enum, a date format, or a tiny DSL. The costs are that you need control over decoding (so it is available on self-hosted or feature-supporting stacks, not through every hosted endpoint), grammar compilation adds startup overhead, and an over-tight grammar can force the model down a path it would not have chosen, which shows up as worse content rather than worse syntax. ## What none of the levels buy you All three control *shape*. None controls *truth*. A grammar-constrained extractor will happily emit `{"dish_name": "Pad Thai", "price_cents": 1450, "allergens": ["peanuts"]}` for a menu that never listed a price and lists shellfish too. Worse, forcing a required non-nullable field pressures the model into inventing a value, because the alternative — saying nothing — is not representable. So structured output raises two obligations rather than removing them: make absence expressible in the schema (nullable fields, an explicit `not_stated` status), and keep validating semantically after parsing, since only your code knows that a price of 1450 cents is implausible for that dish. ## The practical recipe Pick the strongest level your stack supports. Define the schema once and share it between the request and the validator. Make every genuinely optional value nullable, and provide an explicit escape hatch for "unknown" so the model never has to fabricate. Parse defensively anyway — a tolerant pre-parse that strips a stray fence is cheap insurance when you are on wording-only control. Validate, fail closed on violations, and track the parse-failure and validation-failure rates as first-class metrics: they are your early warning that a model upgrade, a prompt edit or an unusual input distribution has broken the contract.

  • If grammar-constrained decoding makes malformed JSON impossible, why keep a validator at all?
    Because the grammar only guarantees the response parses and matches the declared types. It cannot tell you that an allergen list is incomplete, that a price is implausible, that a required identifier is a hallucinated string of the right shape, or that two fields contradict each other. The validator carries your business invariants — cross-field checks, ranges, referential lookups — which no decoder-level constraint can express.
  • What does a tight schema cost you on every single request?
    Tokens and, sometimes, quality. The schema is serialized into the request, so a fourteen-field spec with long enums is paid for on every call and counts against the context budget. Compilation of a grammar adds startup latency. And an over-tight schema removes room the model would have used to reason, which is why a free-text reasoning field placed before the constrained fields often recovers accuracy.
  • How would you detect that a model upgrade broke your output contract?
    Instrument the boundary: emit metrics for parse failures, schema-validation failures, repair-prompt invocations and per-field null rates, tagged by model version. A silent regression usually shows as a jump in one of those, not as an outage. Pin the model version, run a held-out fixture set against the new version before switching, and diff per-field results rather than only checking that responses parse.

A format instruction is a sign asking drivers to slow down; grammar-constrained decoding is a speed bump. One appeals to intent, the other removes the option.

saying these in an interview costs you the question

  • Claiming a strongly worded instruction guarantees valid JSON
  • Believing constrained decoding also guarantees the values are correct
  • Treating a stray ```json fence as a rare edge case not worth handling
  • Regex-scraping fields out of prose instead of constraining output
  • Assuming low temperature alone prevents format violations

context

open as a page

When an LLM's JSON fails schema validation, how should the calling code respond?

level: middleimportance: must knowfreq 63%

basics

~20 s

Fail closed: reject the response, never coerce or scrape it. Then optionally send one targeted repair prompt quoting the exact validator error, cap retries at one or two, and fall back to a deterministic path or human review while logging the failure rate.

open as a page

When should an LLM return Markdown instead of JSON, and why not mix both?

level: juniorimportance: should knowfreq 52%

basics

~20 s

Ask for Markdown when a human reads the output directly, and JSON when code consumes it. Mixing them breaks parsing, since prose around an object stops the whole response from parsing. Put human-facing text inside a JSON field instead.

open as a page

Why does an LLM's long structured output drift, like row 60 of a 100-row table losing a field?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Adherence decays over a long generation: the format spec recedes into distant context, each new row is copied from nearby rows rather than the schema, and one omission propagates. Fix by batching, per-row validation with row-count postconditions, and constrained decoding.

open as a page

When does forcing a strict JSON schema on an LLM hurt output quality?

level: principalimportance: should knowfreq 36%

basics

~20 s

When the schema makes truth unrepresentable or removes reasoning room. Required non-nullable fields push the model to fabricate, a closed enum forces a wrong nearest label, and constraining from the first token denies it space to think before committing.

open as a page