In LangChain, when should you use with_structured_output() instead of PydanticOutputParser?
answer
- one asks nicely, one constrains the API
- schema goes to the provider, not the prompt
- format instructions are just text
- include_raw keeps the usage metadata
- method= picks the enforcement mechanism
basics
~20 sPrefer with_structured_output() whenever the provider supports tool calling or JSON-schema mode: the schema is enforced at the API layer. Fall back to PydanticOutputParser only for models with no native structured output, since it merely asks nicely in the prompt.
solid answer
~40 s`with_structured_output(Schema)` wraps a chat model so the schema is sent to the provider as a tool definition or a JSON-schema response format, and the call returns a validated Pydantic object (or dict) instead of an `AIMessage`. The provider constrains generation, so malformed output is rare. `PydanticOutputParser` works differently: you inject `get_format_instructions()` into the prompt, the model emits free text, and you parse afterwards — a plain request that a model can ignore, raising `OutputParserException`. Use the parser only when the model has no structured-output support. Two things to know about the wrapper: `method=` selects `"function_calling"`, `"json_mode"` or `"json_schema"` (strict), and by default you lose the `AIMessage`, so pass `include_raw=True` when you still need token usage or the raw text alongside `parsed` and `parsing_error`.
code
python · 14 linesfrom pydantic import BaseModel, Field
from langchain_openai import ChatOpenAI
class Ticket(BaseModel):
title: str = Field(description="short summary of the issue")
priority: int = Field(description="1 (low) to 5 (urgent)")
model = ChatOpenAI(model="gpt-4o-mini", temperature=0)
structured = model.with_structured_output(Ticket, include_raw=True)
result = structured.invoke("Printer on floor 3 is jammed again.")
print(result["parsed"]) # Ticket instance, or None
print(result["parsing_error"]) # exception, or None
print(result["raw"].usage_metadata)go deeper
Know that with_structured_output(Schema) hands you a validated object while a parser only reads text the model produced, and that you must handle the case where parsing fails.
Explain the mechanism: the schema is sent as a tool definition or JSON-schema response format, the method= choices differ in strictness, and include_raw=True preserves the AIMessage.
Show the operational side — schema tokens on every call, provider-by-provider support gaps, retry and fallback on parsing_error, and why field descriptions are your cheapest accuracy lever.
Own the portability tradeoff: leaning on strict schema decoding couples you to providers that offer it, so decide deliberately whether structured output is a hard dependency or a capability you degrade gracefully without.
## Two ways to get JSON out of a model Getting a typed object out of an LLM is a daily production problem, and LangChain offers two mechanisms that look interchangeable and are not. **Prompt-and-parse.** `PydanticOutputParser(pydantic_object=Ticket)` derives a JSON schema from your Pydantic model and renders it as prompt text via `get_format_instructions()`. You insert that text into the prompt, the model produces free-form output, and the parser extracts the JSON and validates it. If the model wraps the JSON in prose, emits a trailing comma, or invents a field, the parser raises `OutputParserException`. **Provider-native structured output.** `model.with_structured_output(Ticket)` returns a runnable that sends the schema *to the API* — as a tool/function definition, or as a JSON-schema response format — and returns a validated `Ticket` instance directly. The constraint is applied by the provider during decoding, not by hope. ## What with_structured_output actually does The `method` parameter selects the mechanism: - `"function_calling"` — the schema becomes a tool definition and the model is asked to call it. Broadest support, and the historical default for most providers. - `"json_mode"` — the provider guarantees syntactically valid JSON but not that it matches your schema; your prompt must still describe the fields. - `"json_schema"` — strict schema-constrained decoding where the provider supports it. Strongest guarantee, but the supported subset of JSON Schema is narrower: deeply nested unions, some `anyOf` shapes and certain constraints may be rejected. The schema argument can be a Pydantic model (you get an instance back), a `TypedDict` or a raw JSON-schema dict (you get a dict back). `include_raw=False` is the default and returns just the parsed object. That is convenient and quietly costly: the `AIMessage` is discarded, and with it `usage_metadata`, `response_metadata` and the raw text. Passing `include_raw=True` returns a dict with `raw`, `parsed` and `parsing_error` keys — and it also changes failure behaviour: instead of raising, a parse failure comes back with `parsed=None` and the exception in `parsing_error`, which you can inspect and route. For any service that meters tokens or logs model output, `include_raw=True` is usually the right default. ## Costs of the native path It is not free. The schema is serialized into the request, so a large nested model adds input tokens on every call — sometimes hundreds. Structured output usually occupies the same channel as tool calling, so a model asked to both call tools and emit a schema in one turn can behave awkwardly; the common design is to separate those steps. Strict `json_schema` mode may force you to simplify a schema you would rather keep rich. And support varies by provider and by model within a provider, which is a real portability constraint if you swap models by config. ## When the parser is still right - A local or older model with no tool-calling or JSON mode at all. - Output that is not JSON: a delimited list, an enum token, XML. - Cases where you want the model's reasoning text *and* a structure in the same free-form response and you are willing to parse both. A sensible middle ground is `JsonOutputParser`, which returns dicts and tolerates partial output while streaming, at the price of no validation. ## Failure handling either way Neither approach removes the need for a failure path. With the parser, catch `OutputParserException` and decide: retry with the error fed back, fall back to a stricter model, or degrade. With the wrapper, use `include_raw=True` and branch on `parsing_error`. Validation errors from Pydantic — a wrong type, a missing required field, an out-of-range value — can still occur under `"function_calling"` because the provider is not enforcing your validators, only the shape it was given. Field descriptions matter more than people expect: `Field(description=...)` text is carried into the schema the provider sees and is often the cheapest accuracy fix available. ## Deciding in an interview State the rule first — native structured output when the provider supports it, parser as the fallback — then show you know the edges: token cost of the schema, `include_raw=True` to keep usage metadata and turn exceptions into inspectable errors, and the narrower schema subset under strict mode. Candidates who only know `PydanticOutputParser` are usually reciting pre-1.0 tutorials.
- What do you lose by leaving include_raw at its default of False?You get only the parsed object, so the underlying `AIMessage` is gone — no `usage_metadata` for cost tracking, no `response_metadata` for finish reason, no raw text for debugging or auditing. Failures also raise instead of returning inspectable state. With `include_raw=True` you receive a dict of `raw`, `parsed` and `parsing_error` and can log all three.
- Under method="function_calling", can Pydantic validation still fail?Yes. The provider is constrained to the JSON shape it was given, not to your Python validators. A field with a range check, a custom validator, or a semantic constraint the schema cannot express can still come back invalid, and Pydantic raises when the object is constructed. Treat a validation failure path as mandatory, not optional.
- Your structured-output call started failing after you nested a union three levels deep. What is the likely cause?Strict `json_schema` mode supports a narrower subset of JSON Schema than Pydantic can express — deep unions, some `anyOf` shapes and certain constraints are rejected by the provider. Either flatten the schema, or switch `method` to `"function_calling"`, which is more permissive at the cost of a weaker guarantee.
- How would you make a schema produce better results without changing the model?Write real `Field(description=...)` text on every field and use precise types and enums rather than free strings. Those descriptions are serialized into the schema the provider receives, so they act as targeted instructions exactly where ambiguity occurs — usually a bigger accuracy win than rewording the system prompt.
saying these in an interview costs you the question
- Thinking format instructions in the prompt enforce the schema
- Assuming with_structured_output works on every provider and model
- Not handling parse or validation failure at all
- Believing include_raw=False still exposes token usage
- Using the parser path by default because a pre-1.0 tutorial did