In LangGraph, when is a Pydantic state schema better than a TypedDict?
answer
- typing vs checking at runtime
- fails at the boundary, not downstream
- defaults remove seeding boilerplate
- validation costs per node entry
- merge rules are unchanged either way
basics
~20 sChoose a Pydantic BaseModel when you want runtime validation and real defaults on the state passed into nodes — useful at untrusted or model-generated boundaries. A TypedDict is lighter and has no runtime cost. Neither changes merge semantics: reducers still come from Annotated.
solid answer
~50 s`StateGraph` accepts a `TypedDict`, a dataclass or a Pydantic `BaseModel` as its state schema. A `TypedDict` is the default choice: it is a plain dict at runtime, costs nothing, and gives static typing only — a wrong type or a missing key is not caught until something reads it. A Pydantic model buys you **runtime** validation and real defaults: LangGraph validates the state against the model when it is passed into a node, so a malformed value fails at the graph boundary with a clear error instead of surfacing three nodes later inside a prompt. That is worth paying for when state is seeded from user input, from another service, or from model-generated structured output. The costs are validation overhead on every node entry and stricter serialization requirements. Nothing else changes: nodes still return partial dicts, and a field that must accumulate still needs a reducer declared via `Annotated`.
code
python · 23 linesimport operator
from typing import Annotated
from pydantic import BaseModel, Field
from langgraph.graph import StateGraph, START, END
class State(BaseModel):
topic: str
attempts: int = 0
notes: Annotated[list[str], operator.add] = Field(default_factory=list)
def work(state: State) -> dict:
return {"attempts": state.attempts + 1, "notes": [f"saw {state.topic}"]}
builder = StateGraph(State)
builder.add_node("work", work)
builder.add_edge(START, "work")
builder.add_edge("work", END)
# notes and attempts need no seeding; the node still returns a partial dict
print(builder.compile().invoke({"topic": "schemas"}))go deeper
Know that TypedDict is the common default and that Pydantic is the option when you want values actually checked at runtime rather than only by a type checker.
Explain that Pydantic validates state on the way into a node and supplies defaults, while nodes still return partial dicts and reducers still come from Annotated.
Argue the tradeoff for a real system: validate at untrusted and model-generated boundaries, keep internal scratch state cheap, and account for serialization strictness before adopting it graph-wide.
Own schema evolution: a state schema is a contract across nodes and across persisted snapshots, so tightening a field is a migration, and the boundary where validation lives is an architectural decision, not a style preference.
## Three ways to type state `StateGraph(State)` accepts a `TypedDict`, a Python dataclass, or a Pydantic v2 `BaseModel`. All three do the same structural job — declare the keys, which become the graph's channels. They differ in what happens at runtime. - **TypedDict** — a dict at runtime. Annotations are erased; nothing is checked. Access is `state["key"]`. - **dataclass** — a real object with attribute access and defaults, but no validation beyond what you write yourself. - **Pydantic BaseModel** — attribute access, defaults, coercion and validation, at a per-instantiation cost. ## What Pydantic actually buys The payoff is a **failure that happens at the boundary**. In a `TypedDict` graph, a node that puts an `int` where the schema says `str` produces no error; the bad value flows onward and surfaces as a confusing failure somewhere downstream — often inside a prompt template or a tool call, far from the node that caused it. With a Pydantic schema, state is validated when it is passed into a node, so the run fails with a validation error naming the field. That matters most at three boundaries: input seeded from an external caller, values produced by an LLM (structured output is best-effort, not guaranteed), and long-lived graphs where a bad value would otherwise be persisted and replayed. Defaults are the second win — `Field(default_factory=list)` means a node can read a field on the very first superstep without seeding it, where a `TypedDict` read of an unset key raises `KeyError`. Attribute access (`state.messages`) is a smaller but real ergonomic gain: typos become attribute errors rather than silently returning nothing from a `.get()`. ## What it does not change Candidates routinely over-claim here. A Pydantic schema does **not**: - change the update protocol — nodes still receive the state and return a **partial dict** of changed keys; - give you merge semantics — a list field that must accumulate still needs `Annotated[list[str], operator.add]`, exactly as with a `TypedDict`, and without it the field is still last-value with one write per superstep; - validate deeply on every mutation — validation applies where state is constructed and passed in, so mutating a nested model in place does not re-run validators; - make anything persistent — persistence is a separate, explicitly configured concern. ## The costs Validation runs on state passed into nodes, so a graph with many small nodes pays it many times per run. For a big state object with nested models this is measurable, though usually small next to an LLM call — which is the honest framing: the cost is negligible in an LLM-bound graph and can matter in a tight non-LLM loop. The subtler cost is **serialization strictness**. State that will be snapshotted has to be representable; Pydantic models with custom types, arbitrary objects, or `arbitrary_types_allowed` push you toward custom serializers. A `TypedDict` of plain JSON-ish values sidesteps that entirely. A third cost is coercion surprise. Pydantic v2 will happily coerce some inputs, so a value can pass validation with a different type than the caller intended. If you want that to fail, say so with strict typing rather than assuming validation means identity. ## Choosing A reasonable default: start with a `TypedDict`. Reach for Pydantic when the graph has a real external input surface, when state carries model-generated structured data you do not trust, when defaults would otherwise force every caller to seed half the schema, or when the team benefits from one enforced contract across many nodes owned by different people. A hybrid is often best: a Pydantic model for the *input* the graph accepts, and a lighter internal schema for the scratch keys nodes pass among themselves — validate at the edge, stay cheap inside. ## Migration note Switching a live graph from `TypedDict` to `BaseModel` is not a pure refactor. Every `state["key"]` becomes `state.key`, every node that returned an unexpected extra key now has to be honest, and any previously persisted snapshot must still parse under the new model — a field you made required will reject old data that lacks it.
- If Pydantic validates state, why do nodes still return plain dicts?Because the update protocol is per key, not per object. A node returns only the keys it changed, and LangGraph applies each to its channel; the merged result is what gets constructed and validated as a model on the way into the next node. Returning a whole model would force every node to know the entire schema, which is exactly the coupling partial updates avoid.
- Does a Pydantic schema remove the need for reducers on list fields?No. Merge semantics come from the channel, and the channel's rule comes from the Annotated metadata on the field — identical to a TypedDict. A plain list field is still last-value, still accepts one write per superstep, and still loses history across loop iterations. Pydantic validates the value; it does not decide how two values combine.
- Where does a Pydantic state schema cause trouble with persistence?When fields hold types that do not serialize cleanly. A snapshot has to be written and read back, so custom classes, clients, or arbitrary_types_allowed fields need custom serialization or they break the round trip. Fields you later make required also reject snapshots written before the change, so schema evolution on a graph with saved state needs the same care as a database migration.
saying these in an interview costs you the question
- Claims Pydantic gives automatic merging of fields
- Thinks TypedDict annotations are checked at runtime
- Believes nodes must return a full model instance
- Says validation runs on every nested mutation
- Assumes Pydantic state implies persistence