In OpenAI strict json_schema mode, what schema rules apply and how do you mark a field optional?
answer
- a subset, not all of JSON Schema
- objects must close themselves off
- every declared key is mandatory
- optional means nullable, not absent
- field order is generation order
basics
~20 sStrict mode accepts a restricted JSON Schema subset: every object must set additionalProperties to false and list every one of its properties in required. Optionality is expressed by making the type a union with null, never by leaving the key out of required.
solid answer
~50 sConstrained decoding needs a schema it can turn into a grammar, so `strict: true` narrows JSON Schema to a subset with two structural rules that catch everyone. First, every object — the root and every nested one — must declare `"additionalProperties": false`, so the model cannot invent keys. Second, every key in `properties` must appear in `required`; a partially-required object is rejected at submit time. That seems to forbid optional fields, and the sanctioned workaround is a nullable union: keep the key in `required` and give it `"type": ["string", "null"]`, so the model must emit the key but may emit `null`. Beyond that, the root must be an object, only a subset of type-specific validation keywords is supported, `$defs` and `$ref` let you express recursion, and there are hard caps on nesting depth, property count and total schema size. Violations fail the request rather than degrading quality silently.
code
json · 26 lines{
"name": "contact",
"strict": true,
"schema": {
"type": "object",
"properties": {
"full_name": { "type": "string" },
"middle_name": { "type": ["string", "null"] },
"role": { "type": "string", "enum": ["engineer", "manager", "other"] },
"reports": {
"type": "array",
"items": { "$ref": "#/$defs/person" }
}
},
"required": ["full_name", "middle_name", "role", "reports"],
"additionalProperties": false,
"$defs": {
"person": {
"type": "object",
"properties": { "full_name": { "type": "string" } },
"required": ["full_name"],
"additionalProperties": false
}
}
}
}go deeper
Remember the two rules you will hit on your first attempt: additionalProperties false on every object, and every property listed in required. Know that optional means a nullable type, not a missing key.
Explain why the subset exists — the schema becomes a decoding grammar — and show the nullable-union pattern from memory. Mention that violations are rejected at request time, not silently ignored.
Demonstrate schema design judgment: narrow per-call schemas over reused domain models, enums for classification, declaration order placing reasoning before conclusions, and application-level validation for invariants the dialect cannot express.
Own the policy question of where schemas come from — generated from typed models in one place, versioned alongside consumers — and how a vendor-specific dialect constrains portability if you later route the same call to another provider.
## Why there is a dialect at all Strict Structured Outputs works by compiling your schema into a formal grammar and masking the sampler to tokens that can still complete a valid document. A grammar compiler cannot handle arbitrary JSON Schema — validation keywords that require looking at a finished value (arbitrary cross-field conditionals, unconstrained open objects) do not translate into a decision the decoder can make one token at a time. So `strict: true` defines a subset, and anything outside it is rejected when you submit the request. That rejection is a feature. You learn about an unsupported schema in development, as a 400, rather than in production as a slow quality drift. ## Rule 1 — closed objects Every object in the schema, root and nested alike, must carry `"additionalProperties": false`. Without it the grammar would have to admit arbitrary keys, which defeats the purpose. In practice this is the most common cause of a rejected schema, because most hand-written and most tool-generated schemas leave `additionalProperties` unset (which in plain JSON Schema means "anything goes"). ## Rule 2 — everything is required Every key listed under `properties` must also appear in the `required` array. A schema with five properties and three required entries is rejected. Newcomers read this as "strict mode cannot express optional fields," which is not quite right — it cannot express *absent* fields. ## Expressing optionality: the nullable union The supported pattern is to keep the key required and widen its type: - `"middle_name": {"type": ["string", "null"]}` Now the model must emit `middle_name`, but `null` is an allowed value. Your application then treats `null` exactly as it would have treated an absent key. If you generate schemas from typed models — Pydantic in Python, Zod in TypeScript — the SDK helpers do this translation for you when a field is declared optional, which is one strong argument for generating schemas rather than writing them by hand. A subtlety worth knowing: because the field is always present, the model spends tokens emitting it. On a wide schema with many rarely-populated fields, that is real cost and real latency, and it is a reason to keep extraction schemas narrow. ## Other shape rules - **Root must be an object.** You cannot make the root a bare array or a top-level `anyOf`; wrap it, e.g. `{"items": [...]}`. - **`anyOf` is allowed for non-root positions**, which is how you model a discriminated union of result shapes. - **Recursion is supported through `$defs` and `$ref`**, including a reference to the root, so tree-shaped outputs like an expression AST or a nested outline are expressible. - **Enums are supported** and are usually a better classification tool than free-form strings, because the grammar makes an out-of-vocabulary label impossible rather than merely unlikely. - **Only a subset of type-specific validation keywords is honoured**, and the supported set has expanded over time. Treat any keyword beyond structure — patterns, numeric bounds, array cardinality — as something to confirm against current documentation and, regardless, to re-check in your own code. - **Hard limits exist** on total properties, nesting depth and total schema size. A schema that models an entire domain will hit them; one that models a single call's output rarely does. ## Design consequences The dialect pushes you toward small, purpose-built schemas rather than reusing your internal domain model. That is generally the right instinct anyway: the model should be asked for the fields the call needs, not for every field a database row happens to have. Property order matters more than in ordinary JSON Schema. Because decoding is left-to-right, fields are produced in the order the schema declares them, so a field intended to hold the model's working — an `analysis` or `evidence` string — must be declared *before* the conclusion it supports, or the conclusion is generated first and the analysis becomes post-hoc narration. Finally, the restricted dialect means the schema is not your whole validator. Business rules — a total that must equal the sum of the line items, a date that must be in the past, an identifier that must exist in your database — live in application code no matter how expressive the schema subset becomes. ## Interview framing Name the two structural rules, give the nullable-union answer for optionality, and add one consequence — usually that you should generate the schema from a typed model rather than hand-maintain it, since the helper applies both rules automatically.
- Can the root of a strict schema be an array of results?No — the root must be an object. Wrap the collection in a named property, such as `{"type": "object", "properties": {"results": {"type": "array", "items": {...}}}, "required": ["results"], "additionalProperties": false}`. The wrapper also gives you a place to add per-call metadata later without restructuring consumers.
- How would you express a recursive structure, like a nested outline?Define the node shape once under `$defs` and have its children array `$ref` back to that definition; a reference to the root schema itself is also allowed. The grammar compiler handles the recursion, though deeply recursive outputs still consume tokens and can hit the nesting and size caps, so bound the depth in your prompt.
- Why does the order of properties in the schema matter?Generation is left-to-right, and fields are emitted in declaration order. If a `reasoning` field is declared after `answer`, the answer is committed first and the reasoning is written to justify it. Declaring the working before the conclusion lets the earlier tokens actually condition the later ones.
- Does satisfying the schema mean you can skip validation in your own code?No. The dialect cannot express most business invariants — cross-field sums, referential existence, domain-specific ranges — and even supported constraints say nothing about whether a value is true. Parse with confidence, then validate semantics as you would with any untrusted input.
saying these in an interview costs you the question
- Marks a field optional by omitting it from required
- Leaves additionalProperties unset on nested objects
- Assumes any valid JSON Schema is accepted in strict mode
- Makes the root a bare array of results
- Declares the conclusion field before the reasoning field