skip to content

Structured Output & Tool-Call Reliability

You learn how to make model output machine-consumable and its tool calls dependable enough that downstream code can trust them. This is where an LLM stops being a chat box and starts being a component in a program.

on this pageshow

questions

14

Why validate an LLM's JSON output against a strict schema before your code uses it?

level: juniorimportance: must knowfreq 72%

answer

  1. parse is not the same as validate
  2. the model is an untrusted producer
  3. types, enums, required fields
  4. fail loudly at the boundary
  5. schema-valid can still be semantically wrong

basics

~20 s

Parsing only proves the text is syntactically JSON. Validation proves the fields your code reads actually exist, with the right types and allowed values — so bad output fails loudly at the boundary instead of corrupting logic downstream.

solid answer

~50 s

Parsing and validating are two different checks. A JSON parser answers one question: is this text well-formed? It happily accepts an object with a misspelled key, a quantity returned as the string `"12"`, a missing required field, or an invented field nobody asked for. Schema validation answers the question your program actually cares about: does this value have the shape the rest of the code assumes? So you validate once, at the boundary where model output enters the system, and pass a typed object inward — never a raw dictionary. Treat the model as an untrusted producer, because it is non-deterministic: the same prompt can yield a different key spelling or a truncated object on the next call. Validation failure should be an explicit, logged, handled outcome — repair, escalate, or reject — not a `KeyError` three layers deep in business code.

go deeper

for a junior

Be able to say plainly that parsing checks syntax while validation checks shape, and that you validate model output before using it. Naming a few things a schema enforces — required fields, types, enum values — is enough at this level.

for a middle

Explain the two layers: structural validation against the schema, then business-rule checks a schema cannot express, such as cross-field constraints and lookups against your own reference data. Be ready to say where in the code the boundary sits.

for a senior

Show the operational side. Talk about logging the raw output with the validator error, measuring per-field failure rates, and treating validation failure as expected traffic with an explicit handling path rather than an exception nobody catches.

for a principal

Own the position that the model is an untrusted producer inside your own system, and that validation is a contract you keep regardless of provider capability. Be ready to argue what belongs in the schema versus a separate rules layer, and who owns each when they drift.

## Parsing is not validating When a model returns text you intend to consume programmatically, two independent checks stand between that text and your business logic. The first is **parsing**: turning bytes into an in-memory value. A JSON parser answers exactly one question — is this text syntactically well-formed JSON? Balanced braces, quoted keys, no trailing comma. That is all. A parser will cheerfully hand you `{"quantitiy": "12", "mode": "aeroplane"}` because that text is perfectly valid JSON. It is also useless to a program expecting `quantity` as an integer and `mode` drawn from a fixed set. The second is **validation**: checking the parsed value against a declared schema. This is the check that answers the question your code actually depends on — does this value have the shape the rest of the program assumes? ## What a schema check actually buys A schema language (JSON Schema being the common one in this space) lets you assert, mechanically: - **Presence** — which fields are `required`, so a silently omitted field is caught rather than read as null. - **Types** — `"integer"` versus a stringified number, `"boolean"` versus the word `"yes"`. - **Allowed values** — `enum` membership, so a status field cannot arrive as a synonym the model preferred. - **Format** — `pattern` for identifiers with a fixed shape, numeric `minimum`/`maximum` for ranges. - **Closed shape** — `additionalProperties: false` rejects invented keys instead of letting them pass unnoticed. Each of these turns a class of silent corruption into a loud, localized failure with a precise message naming the offending path. ## Why the model boundary needs this more than an ordinary API boundary You validate input from any external system, but a language model is a harsher producer than most. It is non-deterministic: the same prompt at the same settings can return a different key spelling, a different enum synonym, or a differently nested object across calls. Output can be **truncated mid-object** when it hits a token limit, producing text that is not even parseable. And a model asked to fill a field it has no grounds for will often invent a plausible-looking value rather than omit it — the failure mode that most needs a mechanical check, because it looks entirely reasonable. Even where a provider offers a strict schema mode that constrains generation, validation on your side remains the check you own: it survives a provider without that mode, a self-hosted model, a schema-mode fallback path, and a truncated response. It is the safety net, not the duplicate. ## Schema-valid is still not the same as correct This is the point that separates a rehearsed answer from a real one. A schema constrains **structure**; it cannot constrain **meaning**. Consider a freight forwarder turning a broker's email into a customs declaration. A schema can guarantee that `commodity[2].hs_code` is a six-digit string, that `incoterm` is one of the eleven allowed codes, and that all fourteen required fields are present. It cannot tell you that the HS code is not a real tariff line, that gross weight came back lower than net weight, that the declared value is off by three orders of magnitude, or that the incoterm is incompatible with the transport mode. So production pipelines run two layers: **structural validation** against the schema, then **semantic validation** — cross-field constraints, referential checks against your own reference data, and plausibility bounds. Both layers produce the same kind of artifact: a precise, machine-readable error naming what failed and why. That artifact is what makes an automated repair attempt possible at all. ## Fail closed, at the edge The operational discipline is: validate once, immediately after parsing, and convert to a typed domain object right there. Downstream code should never see an unvalidated map. If validation fails, that is a first-class outcome to handle explicitly — log the raw output alongside the validator error and a request identifier, then choose between a repair attempt, escalation to a human, or a clean rejection. What you must not do is let the raw parsed value flow inward and discover the problem as a type error deep in business logic, where the original output is no longer available to diagnose. ## The cost argument Validation is cheap in the only currency that matters here. Checking an object against a schema takes microseconds against a generation that took seconds and cost real money. There is no throughput argument for skipping it, and the debugging asymmetry is enormous: a validator error names the exact failing path, while an unvalidated bad field surfaces later as a corrupted record with no trace of where it came from.

  • If a provider already enforces the schema during generation, why keep your own validator?
    Because the check you own is the one that survives everything: a model or deployment without strict mode, a fallback path, a truncated response, and — most importantly — the business rules no schema can express. Provider enforcement narrows the structural failures; it does not make a wrong HS code, an impossible weight, or an inconsistent field pair go away. Keeping validation also means your failure handling is identical regardless of which model served the request.
  • What do you log when validation fails, and why does it matter?
    The raw model output verbatim, the full validator error including the failing path and expected-versus-actual, the schema version, the model and settings used, and a correlation id. Without the raw output you cannot tell a schema-design problem from a one-off sampling artifact. Aggregated over time these logs give you a per-field failure rate, which is the signal that tells you to fix the schema rather than lean harder on retries.
  • Should validation errors be raised as exceptions or returned as values?
    Either works, but the boundary must be explicit and total: every call site handles the failure path. Returning a result value tends to be safer here because the failure is expected traffic, not an exceptional condition — a small percentage of calls will fail validation as a matter of course. What matters is that no code path can accidentally proceed with an unvalidated object, which usually means converting to a typed domain object at the boundary and never passing the raw map inward.

Parsing is checking that a form was filled in with a pen rather than smeared; validation is checking that the boxes contain a date, a country code, and a number.

saying these in an interview costs you the question

  • Says a successful JSON parse means the output is valid
  • Assumes a strict provider mode removes the need to validate
  • Reads model output fields directly without any check
  • Thinks schema validation catches wrong or implausible values
  • Lets validation failures surface as type errors deep in business code

context

open as a page

How does logit masking during decoding make invalid JSON impossible rather than unlikely?

level: middleimportance: must knowfreq 64%

basics

~20 s

The decoder tracks its position in a state machine compiled from the schema and, at every step, drives the logits of all tokens that cannot legally continue to negative infinity. An illegal token has zero probability, so it is never sampled.

open as a page

What is the difference between JSON mode, a strict schema mode, and grammar-constrained decoding?

level: middleimportance: must knowfreq 74%

basics

~20 s

JSON mode guarantees only that the text parses as JSON. A strict schema mode additionally guarantees the object conforms to your schema — required keys, types, enums. A custom grammar constrains any formal language, JSON or otherwise.

open as a page

How do you shape a JSON schema so a model actually complies with it?

level: middleimportance: must knowfreq 54%

basics

~20 s

Prefer flat objects over deep nesting, enums over free text, and a single shape with a discriminator field over unions of alternative shapes. Give fields self-describing names and descriptions that state units and format. Compliance is a property of the schema's design, not only of the model.

open as a page

Why do LLM tool calls arrive with invented ids or out-of-enum values, and how do you catch them?

level: middleimportance: must knowfreq 68%

basics

~20 s

Tool arguments are generated as text, so a model fills in a plausible order id, category or amount rather than leaving a field blank. Validate every call at the tool boundary — types, enum membership, units, and whether the referenced record actually exists — before executing anything.

open as a page

When schema validation fails, what should the repair turn send back to the model?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Send the exact validator error — the failing path, what was expected, what was received — together with the offending output and an instruction to return only the corrected object. Generic messages like "invalid output, try again" waste a full generation and rarely fix anything.

open as a page

A refund tool times out after the payment succeeded — how do you stop a double refund?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Give every side-effecting tool call an idempotency key minted by the harness, not by the model, and pass it downstream so a repeat is a no-op returning the original result. Then return an honest unknown-outcome result telling the model to check status before retrying.

open as a page

In LLM structured output, why enforce a field's enum at decode time rather than by prompt instruction?

level: juniorimportance: should knowfreq 58%

basics

~10 s

Decode-time enforcement makes non-members unreachable: the sampler is only offered tokens that continue an allowed value, so a variant spelling cannot be produced. A prompt instruction only shifts probabilities, so violations survive at volume.

open as a page

When should you force an LLM to call a tool instead of letting it decide?

level: juniorimportance: should knowfreq 45%

basics

~20 s

Force a call on any turn where the answer must come from a system of record. Left free to choose, a model will often answer an order-status question from context or invention; requiring the lookup tool on that turn removes the option and grounds the reply in real data.

open as a page

Where does grammar-constrained decoding's runtime overhead actually come from?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Two places: compiling the schema into an automaton plus its token index, which happens once and should be cached, and computing or looking up a token mask at every generation step for every sequence in the batch.

open as a page

What can still go wrong when constrained decoding guarantees schema-valid output?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Plenty. The object can be perfectly typed and factually wrong, the run can hit the token ceiling and stop mid-object, the model can be forced to invent a value it had no basis for, and rigid formatting can cost output quality.

open as a page

How do you get a usable partial object out of a JSON response that is still streaming?

level: seniorimportance: should knowfreq 32%

basics

~20 s

Buffer the tokens and run an incremental parser that tolerates an unterminated document — typically by closing any open string and brackets on a copy of the buffer, then parsing that. Only values whose closing delimiter has already arrived are trustworthy, and schema validation still runs once on the complete object.

open as a page

An agent repeats a failing tool call every turn — what should the error result say?

level: seniorimportance: should knowfreq 58%

basics

~20 s

An error fed back to a model is a prompt. It must say what failed, what the current state actually is, whether a side effect happened, and what to do instead. A bare status code or stack trace gives the model nothing to change, so it repeats the identical call.

open as a page

How do you set a repair-attempt budget and escalation policy for structured-output failures?

level: principalimportance: should knowfreq 28%

basics

~20 s

Derive the cap from measured marginal success per attempt — usually one or two, because failures are correlated and later attempts rarely convert. Then decide what happens when the budget runs out: a human queue, a partial record, or a hard failure, chosen by the cost of a wrong record versus the cost of no record.

open as a page