skip to content

With OpenAI strict Structured Outputs, when can a response still fail to parse against your schema?

level: seniorimportance: should knowfreq 47%

answer

  1. the guarantee has a precondition
  2. completed generations only
  3. two fields decide before you parse
  4. a prefix is not a document
  5. conformant is not the same as correct

basics

~20 s

The guarantee covers completed generations only. A refusal returns a refusal string instead of content, a generation cut off by the token limit returns a valid prefix that is not valid JSON, and content filtering or transport errors return no usable body at all.

solid answer

~50 s

Strict mode guarantees that *if the model finishes generating*, the text conforms to the schema. Three things break that precondition and each needs explicit handling. A **refusal**: the model declines, and the message carries a populated `refusal` string with `content` null — parsing that as JSON throws, so branch on `refusal` before parsing. **Truncation**: generation hit the token budget, `finish_reason` comes back as `"length"`, and you hold a syntactically incomplete prefix; raise the completion-token budget or narrow the schema, and never repair the fragment. **Content filtering or a transport failure**: `finish_reason` is `"content_filter"`, or the call errored before a body arrived. There is a fourth, quieter failure that is not a parse failure at all: shape conformance says nothing about correctness, so a schema-valid object can still carry fabricated values that only your own invariant checks will catch.

code

python · 36 lines
python
import json
from openai import OpenAI

client = OpenAI()

completion = client.chat.completions.create(
    model="gpt-4o-2024-08-06",
    messages=[{"role": "user", "content": "Extract the invoice fields."}],
    max_completion_tokens=800,
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "invoice",
            "strict": True,
            "schema": {
                "type": "object",
                "properties": {
                    "vendor": {"type": "string"},
                    "total": {"type": "number"},
                },
                "required": ["vendor", "total"],
                "additionalProperties": False,
            },
        },
    },
)

choice = completion.choices[0]

if choice.message.refusal:
    print("refused:", choice.message.refusal)
elif choice.finish_reason != "stop":
    print("incomplete, finish_reason =", choice.finish_reason)
else:
    invoice = json.loads(choice.message.content)
    print(invoice["vendor"], invoice["total"])

go deeper

for a junior

Know that you must look at the response before parsing it: if a refusal is present, or the generation stopped early, there is no complete object to parse and your code should say so rather than crash.

for a middle

Explain the precondition — the guarantee covers completed generations — and name refusal, truncation via finish_reason length, and content filtering as the ways completion fails.

for a senior

Show the production handling: one ordered branch covering transport error, refusal, incomplete finish_reason and parse, with separate metrics per branch, plus semantic validation because conformance is not correctness.

for a principal

Frame it as a reliability budget — decide acceptable refusal and truncation rates, where retries and dead-letter handling live, and how schema width trades against truncation risk and per-call cost across the fleet.

## The exact shape of the guarantee It is easy to hear "guaranteed schema adherence" as "you will always get your object." The precise claim is narrower and much more defensible: *a generation that runs to completion under a compiled grammar cannot contain a token sequence that violates the schema.* Everything that prevents the generation from running to completion sits outside the guarantee, and production code must handle each case. ## Failure 1 — refusal The model may decline a request on safety grounds. When it does, the assistant message carries a populated `refusal` field containing a natural-language explanation, and `content` is null. This is deliberate: rather than smuggling "I can't help with that" into a `summary` string field where it would sail through your schema and poison downstream data, the API puts it in a separate channel your code must look at. The practical rule: check `refusal` first, `content` second. In a batch extraction job, count refusals as their own outcome class — they usually cluster on a specific kind of input and tell you something about your prompt or your data. ## Failure 2 — truncation If generation stops because it exhausted the completion-token budget, `finish_reason` is `"length"` and the text is a *prefix*: correct as far as it goes, unterminated as JSON. Strict mode cannot help here — the grammar constrains which tokens may be emitted, not how many are available. This is the failure that surprises teams, because it is load-dependent. The schema that fits comfortably on typical inputs blows the budget on the one document with two hundred line items. Mitigations, in order of preference: raise the completion-token limit with headroom for the worst realistic case; shrink the schema so fewer tokens are spent on structure; paginate or chunk the work so a single call never has to emit an unbounded array; and bound array sizes in the prompt. What you must not do is close the braces yourself. A repaired fragment is silently missing data, which is worse than a loud failure — the whole point of adopting Structured Outputs was to delete that repair code. ## Failure 3 — filtering and transport `finish_reason` may come back as `"content_filter"`, and ordinary operational failures — a 429, a 500, a timeout, a dropped stream — mean there is no body to parse at all. These are retry-and-alert concerns rather than parsing concerns, but they belong in the same defensive branch as the other two so that one code path decides whether a usable object exists. ## Failure 4 — conformant but wrong The response satisfies the schema and your code parses it happily, and the values are invented. Constrained decoding removes a class of *format* errors; it has no opinion about truth. Two habits follow. First, keep semantic validation after parsing: totals that must reconcile, identifiers that must resolve, dates that must fall in range. Second, prefer enums over free-form strings for anything categorical — the grammar makes an out-of-vocabulary label impossible, which converts a silent data-quality problem into a constrained choice. A related trap: over-constraining can *cause* wrong values. If the schema offers no way to say "not present in the source," the model must still fill the field, and it will invent something plausible. Give extraction schemas a nullable value or an explicit `"unknown"` enum member so honesty is expressible. ## Streaming When you stream a structured response, intermediate chunks are partial JSON by construction. Accumulate the deltas and parse once at the end, or use a streaming-tolerant parser that understands it is looking at a prefix. Treating an arbitrary mid-stream buffer as a complete document reintroduces exactly the fragility the feature exists to remove — and a stream that ends early is the truncation case again, so check the terminal `finish_reason` before trusting the accumulated buffer. ## What good handling looks like One branch, evaluated in order: transport error → retry with backoff; `refusal` populated → record as a refusal, do not parse; `finish_reason` not `"stop"` → treat as incomplete, do not parse; otherwise parse, then run business validation. Instrument each branch separately, because their rates tell you different things — refusals point at prompt or input problems, truncations at budget or schema-size problems, and validation failures at model quality. ## Interview framing Say the guarantee is conditional on completion, then name refusal, truncation and filtering as the ways completion fails, and close with the observation that shape is not truth. That sequence shows you have run this in production rather than read the feature announcement.

  • How do you tell a refusal apart from a normal response in the API payload?
    The assistant message carries a populated `refusal` string while `content` is null. Branch on `refusal` before attempting to parse. Keeping refusals in their own field is deliberate — it stops "I can't help with that" from landing inside a schema-valid string field where downstream code would treat it as extracted data.
  • Your job truncates on roughly one percent of documents. What do you change?
    Raise the completion-token budget with headroom for the worst realistic input, and reduce how many tokens the structure itself costs — drop rarely-used fields, shorten enum members, and split unbounded arrays across paginated calls. Never repair the fragment: a closed-off prefix is silently missing records, which is worse than a visible failure.
  • Does strict mode help when you stream the response?
    The final accumulated text still conforms, but every intermediate buffer is partial JSON by construction, so parse at the end or use a parser that understands prefixes. Check the terminal finish_reason too — a stream that ends on length is the truncation case, and the accumulated buffer is incomplete regardless of how well-formed it looks.
  • How do you stop the model inventing values for fields absent from the source document?
    Make honesty expressible in the schema. Allow null for values that may genuinely be missing, or add an explicit unknown member to the enum, and say in the prompt that those are the correct answers when the source is silent. If the only legal outputs are populated values, a required field forces a fabrication.

saying these in an interview costs you the question

  • Believes strict mode makes parse failures impossible
  • Parses content without checking the refusal field
  • Repairs a truncated fragment by closing the braces
  • Treats schema conformance as evidence the values are correct
  • Parses mid-stream buffers as complete documents

context