Why can LangChain's JsonOutputParser stream partial results while PydanticOutputParser cannot?
answer
- it is about the output contract
- JSON prefixes can be closed
- validation is all-or-nothing
- each emission is whole, not a delta
- one parser can borrow the other's schema
basics
~20 sJsonOutputParser parses incomplete JSON on every chunk and emits progressively fuller dicts. PydanticOutputParser must construct and validate a complete model instance, and a half-finished object cannot satisfy required fields, so it only parses once the text is complete.
solid answer
~50 s`JsonOutputParser` implements incremental parsing: on each streamed chunk it attempts to parse the accumulated text as *partial* JSON, closing unterminated strings and brackets, and yields whatever dict it can build so far. Downstream you see `{}`, then `{"colours": []}`, then `{"colours": ["red"]}` — useful for rendering a form or a list as it arrives. `PydanticOutputParser` cannot do that because its output contract is a validated model instance: required fields must be present and types must check, and a truncated object satisfies neither. It therefore buffers and parses once, at the end. `StrOutputParser` sits at the other extreme — it just extracts `.content`, so it passes chunks straight through. The tradeoff is validation versus latency: choose the JSON parser when you want to paint the screen early, the Pydantic parser when the object must be trustworthy before anything downstream sees it.
code
python · 12 linesfrom langchain_core.output_parsers import JsonOutputParser
from langchain_openai import ChatOpenAI
chain = ChatOpenAI(model="gpt-4o-mini", temperature=0) | JsonOutputParser()
for partial in chain.stream('Reply with JSON: {"colours": [three primary colours]}'):
print(partial)
# {}
# {'colours': []}
# {'colours': ['red']}
# {'colours': ['red', 'yellow']}
# {'colours': ['red', 'yellow', 'blue']}go deeper
Know the three core parsers and what each returns: text, a dict, or a validated Pydantic object — and that only the dict one produces useful partial output while streaming.
Explain the mechanism: JSON prefixes can be closed into a valid document on each chunk, while Pydantic validation requires every required field, so it can only parse the completed text.
Argue the latency-versus-trust tradeoff for a real surface, and show the consumer discipline partial dicts demand — render from the latest state, defer side effects, validate once at the end.
Frame it as an interface contract decision: what your system promises downstream determines the parser, and a streaming user experience should not quietly weaken the validation guarantee other services depend on.
## Three parsers, three contracts LangChain's core output parsers live in `langchain_core.output_parsers` and differ mainly in what they promise about their output type. - `StrOutputParser` — takes an `AIMessage` (or chunk) and returns its text content. No structure, no validation. - `JsonOutputParser` — returns a `dict` parsed from the model's text. No validation, but tolerant of incomplete input. - `PydanticOutputParser` — returns an instance of the Pydantic model you configured. Full validation, all or nothing. That last column is what decides streaming behaviour. ## Why partial JSON is possible JSON is a bracketed grammar with a useful property: a prefix of a valid document is *almost* a valid document. Given `{"colours": ["re`, you can close the open string and the open array and the open object and get `{"colours": ["re"]}` — a legal dict that reflects everything known so far. `JsonOutputParser` does exactly this on each chunk, over the accumulated buffer, and emits the best-effort dict. It also strips markdown code fences, because models habitually wrap JSON in triple backticks. The consumer therefore sees a sequence of dicts that grows monotonically in information, not a sequence of string fragments. Keys appear one at a time; a partially generated string value may be shown truncated on one chunk and complete on the next. Application code must be written for that: render from the latest dict rather than accumulating, and don't fire side effects on a value until the stream ends. ## Why Pydantic cannot join in `PydanticOutputParser`'s contract is *an instance of your model*. Constructing it runs Pydantic validation, which enforces required fields, coerces or rejects types, and applies any custom validators. A truncated payload usually has missing required fields — so construction fails outright. Emitting a half-populated object would break the guarantee that made you choose the parser. The parser therefore waits for the complete text, extracts the JSON, validates once and returns. There is a middle option that people miss: `JsonOutputParser(pydantic_object=Ticket)` uses your model only to generate the format instructions, while still yielding plain dicts and still supporting partial output. You get schema-shaped prompting and streaming, and you give up validation — you can always validate the final dict yourself with `Ticket.model_validate(...)` once the stream completes. ## Format instructions are shared machinery Both JSON-shaped parsers expose `get_format_instructions()`, which renders the expected schema as prompt text. That string is the parser's entire influence on the model: the parser sits *after* the call and can only read what arrived. Nothing about a parser constrains generation. If the model answers with prose, the parser raises `OutputParserException`, and the fix lives in the prompt or in switching to provider-enforced structured output. ## Choosing in practice Use `StrOutputParser` when the consumer wants text — the overwhelmingly common case for chat UIs, and the cheapest thing in the list. Use `JsonOutputParser` when a user-visible surface benefits from filling in progressively: a form, a table, a checklist, a plan the user reads while it is written. Perceived latency improvements here are large, because the first partial dict arrives in a few hundred milliseconds rather than after the full generation. Use `PydanticOutputParser` when the output feeds code rather than eyes: anything that will be persisted, used to dispatch a decision, or handed to another service. In those cases a partially-formed object is worse than a slow one. Even then, prefer the provider-native structured-output path when the model supports it, and keep the parser as the fallback for models that do not. ## Failure modes to name - Trailing prose after the JSON — the JSON parser is fairly robust, but a model that narrates around its answer is a prompt problem. - Very large arrays — partial parsing runs on each chunk over the accumulated buffer, so extremely long outputs do repeated work; it is fine for typical response sizes, worth measuring for very large ones. - Consumers written as if chunks were deltas — with `JsonOutputParser` each emission is the whole known object, not a diff. - Assuming the streamed dicts are validated — they are not, and the last one is not automatically validated either. ## Interview framing Lead with the contract difference — dict versus validated instance — then explain prefix-parsing of JSON in one sentence, then name the middle path (`pydantic_object` on the JSON parser) and the consumer discipline partial dicts require. That progression shows mechanism, not memorized API names.
- How do you keep schema-driven prompting but still stream?Pass your model to the JSON parser as `JsonOutputParser(pydantic_object=Ticket)`. It uses the model to render `get_format_instructions()` for the prompt while still yielding dicts and still parsing partial input. Validate the final dict yourself with `Ticket.model_validate(...)` after the stream ends if you need the guarantee.
- What must consumer code do differently for a streamed JsonOutputParser?Treat every emission as the complete current state, not a delta — re-render from the latest dict rather than appending. Defer side effects such as writes, API calls or user-visible confirmations until the stream finishes, because intermediate values can be truncated strings or short arrays that later grow.
- If the model answers with prose instead of JSON, what happens?The parser raises `OutputParserException` once it cannot extract JSON from the text. Parsers run after the call and cannot constrain generation, so the fix is upstream: strengthen the format instructions, lower temperature, or move to provider-enforced structured output where the schema is sent with the request rather than described in the prompt.
Partial JSON is like watching a form fill itself in field by field; a validated model is a signed document — you cannot hand over half of it and call it signed.
saying these in an interview costs you the question
- Saying parsers constrain what the model generates
- Treating streamed dicts as deltas to accumulate
- Believing PydanticOutputParser can emit half-built objects
- Assuming the last streamed dict was validated
- Thinking StrOutputParser does JSON cleanup