skip to content

In the OpenAI API, how does response_format json_object differ from json_schema strict mode?

level: middleimportance: must knowfreq 72%

answer

  1. two different strengths of promise
  2. syntax valid versus shape valid
  3. one mode needs the word JSON in the prompt
  4. strict true compiles the schema into a grammar
  5. strict false quietly downgrades the guarantee

basics

~20 s

json_object only promises the reply parses as JSON — which keys appear and what types they hold is left to the prompt. json_schema with strict true constrains decoding to a schema you supply, so the returned object provably matches that shape.

solid answer

~40 s

Both are values of the `response_format` field on a Chat Completions request, but they promise very different things. `{"type": "json_object"}` is JSON mode: the model is nudged to emit syntactically valid JSON and the API rejects the request unless some message mentions JSON, but nothing stops it from renaming a key, dropping a field, or returning a number as a string — you still write defensive parsing. `{"type": "json_schema", "json_schema": {"name": ..., "schema": {...}, "strict": true}}` is Structured Outputs: the service compiles your schema into a grammar and masks tokens during sampling, so a completed response conforms to the schema by construction. The price is that `strict: true` accepts only a restricted JSON Schema subset. If you set `strict: false`, the schema is treated as a hint, which behaves much closer to JSON mode.

code

python · 38 lines
python
from openai import OpenAI

client = OpenAI()

# JSON mode: valid JSON, no promise about keys or types.
loose = client.chat.completions.create(
    model="gpt-4o-2024-08-06",
    messages=[
        {"role": "system", "content": "Reply in JSON with keys title and year."},
        {"role": "user", "content": "Describe the movie Alien."},
    ],
    response_format={"type": "json_object"},
)

# Structured Outputs: decoding is constrained to the schema.
strict = client.chat.completions.create(
    model="gpt-4o-2024-08-06",
    messages=[{"role": "user", "content": "Describe the movie Alien."}],
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "movie",
            "strict": True,
            "schema": {
                "type": "object",
                "properties": {
                    "title": {"type": "string"},
                    "year": {"type": "integer"},
                },
                "required": ["title", "year"],
                "additionalProperties": False,
            },
        },
    },
)

print(loose.choices[0].message.content)
print(strict.choices[0].message.content)

go deeper

for a junior

Be able to name both settings and say plainly that one only promises parseable JSON while the other promises your exact shape. Knowing where response_format goes in the request is enough at this level.

for a middle

Explain the mechanism: strict mode compiles the schema into a grammar that constrains token sampling, which is why it is a guarantee. Mention the prompt-must-say-JSON rule for json_object and the strict false trap.

for a senior

Show production judgment — that shape conformance is not correctness, that truncation and refusals still break parsing, and that a rejected schema at submit time is preferable to silent drift you discover in a downstream service.

for a principal

Own the tradeoff: constrained decoding buys reliability but couples you to one vendor's schema dialect and adds compile latency for dynamic schemas. Be ready to argue where the validation boundary should sit across a multi-provider estate.

## The problem both modes address An LLM emits tokens, not data structures. If you want to feed a model's answer into code, you need the text to be machine-readable, and historically teams did that with prompt begging ("respond with JSON only") plus repair heuristics: strip Markdown fences, find the first `{`, balance the braces, retry on failure. OpenAI shipped two increasingly strong server-side answers to this, both selected through the `response_format` field of a Chat Completions request. ## json_object — JSON mode `response_format={"type": "json_object"}` makes exactly one promise: if the generation runs to completion, the text parses as JSON. It does not know about your data model, because you never gave it one. The keys, nesting, and types come entirely from your prompt, and the model can and does deviate — inventing an extra field, spelling `total_price` as `totalPrice`, or emitting `"42"` where you wanted `42`. Two operational details matter. First, the API requires that the conversation itself instruct the model to produce JSON; if no message contains the word, the request fails with a 400 rather than silently ignoring the setting. Second, JSON mode does not prevent truncation: if the generation stops because it hit the token budget, you get a valid *prefix* of JSON, which is not valid JSON. ## json_schema with strict: true — Structured Outputs Here you pass a named schema: - `name` — an identifier for the schema (also useful in logs) - `schema` — the JSON Schema document - `strict` — set `true` to turn on constrained decoding With `strict: true` the schema is converted, server-side, into a formal grammar. At each decoding step the sampler is restricted to tokens that can still lead to a schema-valid document. That is why the guarantee is stronger than prompting: it is not that the model was persuaded to comply, it is that non-compliant tokens were never available. Compilation of a *new* schema costs a little extra latency on first use; the compiled artefact is cached, so repeat calls with the same schema do not pay it again. `strict: false` is legal and means something quite different: the schema is passed along as guidance, no grammar is built, and you are back in json_object territory in terms of guarantees. Reviewers love this distinction because it looks like a typo but changes the contract. ## What strict mode costs you Constrained decoding only works on schemas the grammar compiler can express, so `strict: true` accepts a subset of JSON Schema rather than the whole specification, and it imposes structural rules (notably that objects close themselves off to unknown keys and enumerate their properties as required). Deeply nested or very large schemas run into hard caps on nesting depth, property count and total schema size. If your schema does not fit the subset, the request is rejected at submit time — a fast, loud failure rather than a silent quality regression. ## What neither mode guarantees Shape is not truth. A schema-conformant `{"invoice_total": 0}` is still wrong if the invoice was 412.90. Structured Outputs removes parsing failures from your incident list; it does not remove hallucination, so you still validate business invariants after parsing. Neither mode survives an interrupted generation. If the response stops early because of the token limit, strict mode gives you a valid-so-far prefix, and if the model declines the request you get a refusal instead of content. Both cases must be handled explicitly. ## Choosing between them Use strict `json_schema` whenever the shape is known ahead of time and downstream code depends on it — extraction pipelines, classification into a fixed label set, UI-driving payloads. Use `json_object` when the shape is genuinely dynamic or when you are targeting an older model generation that does not support constrained decoding: strict `json_schema` arrived with the `gpt-4o-2024-08-06` snapshot and is supported by later generations, while earlier models offer JSON mode only. Use plain text when the deliverable is prose; wrapping prose in a one-field object buys nothing and can cost quality. ## Interview framing The sentence that lands is: *json_object is a promise about syntax, json_schema with strict true is a promise about shape, enforced during decoding rather than requested in a prompt.* Follow it with the cost — a restricted schema dialect and a first-call compile — and you have said everything the question is testing.

  • What actually happens server-side that makes strict mode a guarantee rather than a strong suggestion?
    The schema is compiled into a grammar, and during sampling the token mask excludes any token that could not continue a schema-valid document. Non-conforming output is unreachable rather than merely discouraged. Compiling a previously unseen schema adds latency to the first request; the compiled form is then cached, so subsequent calls with the identical schema are not penalised.
  • You set json_schema but left strict as false. What changed?
    You lost constrained decoding. The schema is forwarded as guidance for the model, but no grammar is built and nothing prevents extra keys, missing fields, or wrong types — practically the same guarantee level as json_object. It is a common review catch precisely because the request still looks like Structured Outputs at a glance.
  • Does either mode reduce hallucination?
    No. Both constrain form, not content. A response can satisfy every schema rule and still contain fabricated values, so post-parse validation of business invariants — ranges, referential checks, cross-field consistency — stays your responsibility. Structured Outputs removes parse failures and repair heuristics from the pipeline, not factual review.

saying these in an interview costs you the question

  • Claims json_object validates the schema written in the prompt
  • Thinks strict mode also guarantees the values are correct
  • Forgets json_object requires an instruction mentioning JSON
  • Treats strict false as equivalent to strict true
  • Says structured outputs are just a better system prompt

context