skip to content

In the OpenAI Python SDK, what does the chat completions parse() helper give you over create()?

level: juniorimportance: should knowfreq 62%

answer

  1. a typed convenience layer
  2. the class is the schema
  3. no hand-written schema dictionary
  4. an object comes back, not a string
  5. one field still has to be checked first

basics

~20 s

parse() accepts a Pydantic model as response_format, converts it into a strict JSON Schema for you, and hands back the reply already deserialised into an instance of that model on message.parsed — plus the separate refusal field to check first.

solid answer

~50 s

With `create()` you hand-write the `response_format` dictionary, including the `strict` flag and the schema, and you get back a JSON string in `message.content` that you deserialise and validate yourself. `parse()` takes a typed model instead — a Pydantic class in Python — and does three things: it generates the strict-mode schema from that class, applying the dialect rules such as closing objects to unknown keys and marking optional fields nullable; it sends the request; and it returns a parsed completion where `choices[0].message.parsed` is an instance of your class, fully typed for your editor and type checker. The raw JSON is still available in `content`, and `refusal` is still the field to branch on first. The practical win is that one class is the single definition of the output shape, so the schema cannot drift from the type your code consumes.

code

python · 28 lines
python
from openai import OpenAI
from pydantic import BaseModel, Field


class Step(BaseModel):
    explanation: str = Field(description="One reasoning step, shown before the answer")
    output: str


class Solution(BaseModel):
    steps: list[Step]
    final_answer: str


client = OpenAI()

completion = client.chat.completions.parse(
    model="gpt-4o-2024-08-06",
    messages=[{"role": "user", "content": "Solve 8x + 7 = -23 step by step."}],
    response_format=Solution,
)

message = completion.choices[0].message
if message.refusal:
    print("refused:", message.refusal)
else:
    solution = message.parsed
    print(len(solution.steps), solution.final_answer)

go deeper

for a junior

Be able to write the call: define a Pydantic class, pass it as response_format to parse(), read choices[0].message.parsed, and check refusal first. That is the everyday shape of this feature.

for a middle

Explain that the class is converted into a strict-mode schema — closed objects, all fields required, optionals as nullable — so the helper removes the most common cause of rejected schemas.

for a senior

Argue why deriving the schema from the consumed type matters: no drift between contract and code, type errors at every call site when the shape changes, and field descriptions doubling as model instructions.

for a principal

Own where these response classes live and how they version — shared package versus per-service, how a shape change rolls out across consumers, and what the equivalent looks like if part of the fleet calls a different provider.

## What the helper is The OpenAI Python SDK ships a typed convenience layer over Chat Completions. Alongside `client.chat.completions.create(...)` there is `client.chat.completions.parse(...)`, which accepts a Pydantic model as the `response_format` argument. (In older SDK releases this lived under the `client.beta.chat.completions` namespace before being promoted; if you are reading legacy code, that is the same helper.) The TypeScript SDK exposes an equivalent helper with Zod schemas. ## The three jobs it does **Schema generation.** Your Pydantic class is converted to a JSON Schema that satisfies the strict-mode dialect: objects closed to unknown keys, every property enumerated as required, optional fields expressed as nullable unions. Doing this by hand is where most first-attempt schema rejections come from, so letting the helper do it removes a whole class of 400s. **The call.** It sends an ordinary Chat Completions request with `strict` enabled, so you get real constrained decoding, not a prompt-level suggestion. **Deserialisation.** The response is a parsed completion object. `choices[0].message.parsed` is an instance of your class rather than a string, so attribute access is checked by your type checker and your IDE completes field names. The raw text remains in `choices[0].message.content` if you want to log or store it. ## What it does not do It does not remove the checks you owe the response. `choices[0].message.refusal` is still populated when the model declines, and `parsed` is not usable in that case, so the refusal branch comes first. A generation cut short by the token budget is still incomplete — the helper will not hand you a silently half-filled object, so treat an error or an empty `parsed` from a non-`stop` finish reason as a failure to retry rather than something to patch up. It also does not validate business rules. Pydantic will enforce the types you declared, because your class is the deserialisation target, but it knows nothing about whether the invoice total matches the line items. Semantic validation is still application code. ## Why interviewers ask this at the junior tier It is the everyday-use question for structured outputs: the candidate has either written this call or has not. The good answer names the Pydantic-to-schema conversion and `message.parsed`, and mentions that `refusal` still needs checking. A weaker answer describes `parse()` as "the same thing but shorter," which misses that the schema is derived from the type — the property that keeps the contract from drifting. ## Design consequences Because the class is the schema, the class is also your documentation and your migration unit. When the output shape changes, you edit one definition and every consumer's type errors point at the places that need updating. Compare this with a hand-written schema dictionary living next to a separate `TypedDict`: nothing keeps the two in step, and the failure mode is a runtime `KeyError` in a code path nobody exercised. A few practical habits follow. Keep the response class narrow — model what this specific call needs to return, not your whole domain object; every declared field costs output tokens because strict mode requires them all to be emitted. Use `Enum` or `Literal` types for categorical fields so the grammar makes an invalid label impossible. Add docstrings and field descriptions: they are carried into the generated schema and act as inline instructions to the model, which is often more effective than adding another paragraph to the system prompt. And declare fields in the order you want them generated, because decoding is left-to-right — a field that holds the model's reasoning belongs before the field that holds its conclusion. ## Interview framing "`parse()` takes my Pydantic class, turns it into a strict schema, and gives me back an instance instead of a string — but I still check `refusal` and the finish reason before I touch `parsed`." That single sentence covers the mechanism, the benefit, and the caveat.

  • What in your Pydantic class travels into the schema besides field names and types?
    Field descriptions and the class docstring are carried into the generated JSON Schema, where the model reads them as inline guidance. That makes them a compact place to say what a field means — often more effective than another paragraph of system prompt, because the hint sits next to the field it governs.
  • You still get an exception when the response is refused. How should the call site be structured?
    Check `choices[0].message.refusal` before touching `parsed`. A refusal leaves no object to deserialise, so treat it as its own outcome — log it, count it, and route the input for review — rather than as a parse error to retry. Retrying an unchanged refused prompt usually just refuses again.
  • Does using parse() change what the model actually does compared with a hand-written strict schema?
    No. The helper produces the same strict-mode request; constrained decoding behaves identically. What changes is on your side: the schema is derived from the type your code consumes, so the two cannot drift apart, and the response arrives typed instead of as a string you must deserialise.

saying these in an interview costs you the question

  • Calls parse() a shorthand with no schema behind it
  • Reads message.content and ignores message.parsed
  • Skips the refusal check because the call is typed
  • Assumes Pydantic validation covers business rules
  • Reuses a wide domain model as the response class

context