skip to content

Guardrails & Output Constraining

How to constrain what a model emits: grammar-constrained structured output, repair loops, policy rails, and guardrails inside agent SDKs. Interviewers check that you treat output as untrusted.

on this pageshow

questions

5

Why doesn't a strict JSON schema on an LLM response make its content safe?

level: juniorimportance: must knowfreq 62%

answer

  1. shape is not meaning
  2. valid parse, unchecked contents
  3. free-text fields escape the schema
  4. enums are checkable, prose is not
  5. danger appears at the sink

basics

~20 s

A schema constrains shape, not meaning. It guarantees the reply parses and carries the required fields and types. It says nothing about whether a field holds an invented price, a competitor's name, an unsafe instruction, or text that is dangerous where you paste it.

solid answer

~50 s

Structured output and a strict schema solve one problem: the response deserializes into your type, with the required keys, the declared types, and only the enum values you allowed. That removes a whole class of parsing bugs, and it is worth doing. But every free-text field inside that valid envelope is still model-generated prose that no schema inspects. A `summary` field can name a competitor, quote a policy you never wrote, or contain markup; a `refund_amount` of 9,999,999 is a perfectly valid number and a terrible business decision; a `url` field can point anywhere. So structural validation is the first of several layers, not the whole stack: after it you still want business-rule checks on the values, content rails on the free text, and correct escaping or authorization at whatever sink finally consumes the field.

go deeper

for a junior

Be ready to say plainly that a schema checks fields and types, not truth or tone, and that free-text fields inside a valid response are still unchecked model output.

for a middle

Explain the layering: shape, then business rules on the values, then content checks on the free text, then escaping or authorization wherever the field is finally used. Give an example of a valid-but-wrong document.

for a senior

Show the design instinct of moving meaning out of prose into enums and IDs so it becomes mechanically checkable, and name the sinks in your own system where a string field turns into an action.

for a principal

Own the position that model output is untrusted data across the whole platform, and decide where structural validation, policy rails and sink-side defenses each live so teams are not each inventing their own answer.

## The two different jobs "Constraining output" is one phrase covering two jobs that fail independently. The first is **structural**: making the model emit something your program can consume without guessing — valid JSON, the keys you asked for, the types you declared, values drawn from your enum. The second is **semantic and policy**: making sure what those fields *say* is true, permitted, and safe. Providers sell the first as strict structured output; a lot of teams quietly assume they bought the second too. ## What a schema promises A JSON Schema describes shape: required keys, value types, enum membership, nesting, array bounds. When the provider enforces it during generation rather than checking afterwards, you get a strong guarantee that the response parses and conforms. Practically, that means you can drop the defensive re-prompting people used to write around a bare `json.loads`, and you can type the result. It also narrows the model's freedom in a useful way — an `action` field restricted to `refund`, `escalate`, `reject` cannot come back as `refund_partially_and_apologise`. ## What a schema cannot promise Four things a valid document can still be: - **False.** `{"policy_number": "AZ-4471", "covered": true}` is schema-perfect and may be entirely invented. Structure has no truth condition. - **Out of policy.** A schema does not know your product may not compare itself to a named rival, may not give legal advice, and may not swear. Those are content rules, and only a content check applies them. - **Business-invalid.** Types are not constraints on meaning. `refund_amount: 9999999` satisfies `number`; `valid_until` in the year 1200 satisfies `string`. Range, referential and cross-field rules ("refund cannot exceed the original charge") live in your own validation code. - **Unsafe at the sink.** The dangerous moment is usually not generation but use. A `summary` string rendered as raw HTML, a `path` field concatenated into a filesystem call, a `query` field passed to a database, a `url` field auto-opened — the schema said `string`, and `string` is exactly what an attacker wants a field to be. ## Free text is the hole Most real schemas are mostly free text. Enum fields and numbers are genuinely constrained; `explanation`, `message_to_customer`, `notes` are not — they are ordinary generation wearing a field name. If the interesting content of your response lives in a free-text field, the schema has constrained the packaging and left the payload untouched. This is why output guardrails run *on the field contents*, not on the document. ## Why the confusion happens The word "validation" is overloaded. Schema validation answers "does this fit the type?" Content validation answers "is this allowed?" Business validation answers "is this correct for this customer?" They are three different checks with three different owners, and passing the cheapest one is not evidence about the other two. A related trap: because the schema *feels* like a security control, teams stop treating the output as untrusted. It is untrusted. It was produced by a probabilistic system that may have read attacker-controlled text somewhere upstream. ## The layered picture worth drawing In an interview, sketch four layers and say which failure each one catches: 1. **Shape** — strict schema / structured output. Catches unparseable and malformed responses. 2. **Values** — your own business-rule checks. Catches impossible amounts, unknown IDs, contradictory field combinations. 3. **Content** — output rails over the free-text fields. Catches policy violations, forbidden topics, unwanted disclosures. 4. **Sink** — escaping, parameterisation, authorization at the point of use. Catches the damage the first three missed. A good candidate adds the design consequence: **push meaning into structure wherever you can**. Every fact you can move out of prose into an enum, an ID, or a numeric field is a fact you can check mechanically instead of judging with another model. If the response must choose one of nine escalation reasons, an enum makes it checkable; a sentence explaining the reason does not. That is the practical way schemas contribute to safety — not by validating the text, but by shrinking how much text there is to validate.

  • So what does a strict schema actually buy you on the safety side?
    It shrinks the surface. Every decision you move from prose into an enum, a boolean or an ID becomes mechanically checkable, so downstream rails and business rules have less free text to judge. It also removes parse-failure retries, which quietly cause their own outages. The benefit is reduced ambiguity, not content assurance.
  • Where would you put the check that a refund amount is affordable?
    In your own code, after deserialization and before any side effect — it needs data the model never had, such as the original charge and the customer's balance. Treat the model's number as a proposal, not a decision. If it fails, either clamp with an explanation or return a refusal path rather than silently executing a smaller refund.
  • A field is typed as a URL string. What still worries you?
    That the value is attacker-influenced and will be followed. Type says nothing about destination, so an exfiltration link or an internal address both validate. Constrain it to an allowlist of hosts, never auto-fetch it, and render it as inert text unless it passes that check.

saying these in an interview costs you the question

  • Says structured output means the answer cannot be wrong
  • Treats a strict schema as a security control
  • Assumes free-text fields are covered by validation
  • Renders model text into HTML because the schema said string
  • Thinks schema validation replaces business-rule checks

context

open as a page

Where do you place input, output and tool-call guardrails around an LLM agent?

level: middleimportance: must knowfreq 72%

basics

~20 s

Three placements, each catching a different failure. Input rails screen the incoming turn before or alongside generation. Output rails screen the finished response before the user sees it. Tool rails wrap each function call, checking arguments before the side effect and treating the returned result as fresh untrusted input.

open as a page

When is a declarative guardrail framework worth it over hand-rolled validators?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A framework pays when policy churns, when non-engineers must read or edit it, and when you need many off-the-shelf checks composed the same way. Hand-rolled validators win when you have three rules, tight latency budgets, and no appetite for a second runtime and its DSL.

open as a page

How do you design an LLM refusal so it isn't a dead end for the user?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Treat a blocked turn as a product state, not an error. Name the category of what you cannot do, offer the sanctioned path for it, and carry the user's context into that path. A refusal that ends the conversation converts a policy win into an abandoned session.

open as a page

How do you resolve conflicts when several guardrails judge one LLM response?

level: principalimportance: should knowfreq 30%

basics

~20 s

Give rails distinct verbs and a fixed precedence. Blocks outrank redactions, which outrank rewrites; a block short-circuits. Decide per rail whether it fails open or closed when it errors, run mutating rails in a defined order, and cap any repair attempt so rails cannot fight each other indefinitely.

open as a page