In CrewAI Flows, what does typed Pydantic state give you over dict state?
answer
- two modes, one shared object
- dict versus declared model
- misspelled key fails silently in one
- attribute access and coercion in the other
- both carry an auto-generated run id
basics
~20 sDeclaring the flow as Flow of a Pydantic model makes self.state a validated model with attribute access, declared defaults and type checking. Unstructured state is a plain dict you index by key, where a misspelled key silently creates a new one.
solid answer
~50 sCrewAI supports two state modes. **Unstructured**: subclass `Flow` and `self.state` is a dict — write any key, read any key, no schema. **Structured**: define a Pydantic `BaseModel` and subclass `Flow[MyState]`; `self.state` is then an instance of that model, so you write `self.state.report = ...`, defaults come from the model, and Pydantic validates and coerces values on assignment. The practical difference is when a mistake surfaces. With a dict, `self.state["reprot"] = x` succeeds and the bug appears three steps later as a missing key on a different line; with a model it fails where it happened. You also get autocomplete and a readable contract for what the flow actually carries. Both modes get an auto-generated `id` on the state identifying the run. Values passed as `kickoff(inputs={...})` are applied to the state before the start methods run.
code
python · 24 linesfrom crewai.flow.flow import Flow, listen, start
from pydantic import BaseModel
class ReportState(BaseModel):
topic: str = ""
draft: str = ""
revisions: int = 0
class ReportFlow(Flow[ReportState]):
@start()
def outline(self):
self.state.draft = f"Outline for {self.state.topic}"
@listen(outline)
def revise(self):
self.state.revisions += 1
return self.state.draft
flow = ReportFlow()
flow.kickoff(inputs={"topic": "solid-state batteries"})
print(flow.state.revisions, flow.state.id)go deeper
Know there are two modes: a plain dict you index by key, or a Pydantic model you access by attribute, chosen by parameterizing the Flow base class with your model. Both are reachable as self.state from any step.
Explain the real difference: the model validates, coerces and declares defaults so a typo fails at the assignment, while a dict accepts any key and defers the failure to a later consumer. Mention the auto-generated run id and kickoff inputs.
Talk about state as an interface between steps that outlive their author: one writer per field for parallel branches, references rather than raw payloads to keep prompts from growing, and the run id as the thread you log everything against.
Own state as a schema decision with cost consequences. What the flow carries determines what every downstream crew pays in prompt tokens, what can be persisted and resumed, and what a future contributor can safely change without reading every step.
## The two modes A Flow always has state, reachable as `self.state` from any step. What differs is its type. **Unstructured (dict) state** is the default: subclass `Flow`, and `self.state` behaves as a dictionary. Any step can write any key at any time. It is the fastest thing to prototype with and the shape most examples start from. **Structured (Pydantic) state** is opt-in: declare a model and parameterize the base class with it. ``` class ReportState(BaseModel): topic: str = "" draft: str = "" revisions: int = 0 class ReportFlow(Flow[ReportState]): ... ``` Now `self.state` is a `ReportState` instance. Access is by attribute, defaults live in one declaration, and Pydantic validates and coerces on assignment. ## Why the difference matters in a flow specifically State is the *only* channel between steps that are not directly wired. A listener receives its trigger's return value; everything else — the original request, each crew's output, flags a router will branch on — travels on state. That makes state a de facto interface between methods written at different times, and interfaces without schemas rot. The concrete failure with dict state is the silent key. Writing `self.state["summry"]` is legal, so the producer looks fine; the consumer reading `self.state["summary"]` blows up several steps and one crew invocation later, having already spent the tokens. With a model, the assignment itself is the error, at the line that caused it. In a system where each step may cost a dollar and thirty seconds, moving a failure earlier is worth real money. The second benefit is documentation. A reviewer can read the model and know exactly what the flow carries. With dict state they have to grep every step for string keys — and the set is open, so they can never be sure they found them all. ## The run id Both modes carry an auto-generated `id` on the state, a UUID identifying this run. It is what you log alongside crew outputs so a support ticket can be tied back to a specific execution, and it is the handle any state-persistence layer keys on. ## Passing values in `kickoff(inputs={"topic": "batteries"})` applies those values to the state before the start methods run, so a start method can read `self.state.topic` rather than taking parameters. With structured state the values go through the model, so types are validated and coerced instead of being stored blindly; with dict state whatever you pass lands as-is. Prefer inputs over module-level globals: it keeps the flow instantiable in a test with a synthetic state. ## Concurrency on state Parallel branches — two start methods, or two arms both running crews — share one state object. There is no locking and no copy-on-write. The discipline that works is *one writer per field*: give each branch its own field and let the join step read both. Two branches appending to the same list, or incrementing the same counter, is a race with no ordering guarantee, and it will not reproduce reliably. ## Size and cost State is not free context. Whatever you put on state and then interpolate into a crew's inputs becomes prompt tokens. A common anti-pattern is stuffing every intermediate crew output onto state and passing the whole object into the next crew "just in case" — the prompt grows monotonically across the flow, and cost with it. Keep raw payloads out of state, or keep them behind a reference (a file path, a row id) that only the step needing them dereferences. Structured state helps here too, because a model with three fields is a design decision you can review, while a dict with eleven keys accreted without anyone deciding. ## Persistence Because state is a well-defined object with an id, CrewAI can persist it between runs: the `@persist` decorator, backed by `SQLiteFlowPersistence`, saves flow state so a run can be restored rather than restarted. Structured state makes this materially safer — a schema round-trips predictably, whereas an open dict may deserialize into something whose shape a later version of the code no longer expects. ## When dict state is the right call Proof of concept, a three-step flow, a spike you will throw away. The cost of the model is one class; the cost of the dict is deferred to whoever maintains it. The moment a flow has branches, more than a handful of steps, or more than one author, structured state pays for itself.
- What is the auto-generated id on flow state used for?It identifies the run. Both dict and model state carry an auto-generated UUID id, which is what you log next to crew outputs so a result can be traced back to one execution, and what a persistence layer keys on when saving or restoring the state of a flow.
- Two parallel branches both write to state. What can go wrong?They share one object with no locking and no copy-on-write, so two branches touching the same field race — appends interleave, increments are lost, and it will not reproduce on demand. The workable discipline is one writer per field: each branch owns its own field, and the join step reads them all.
- How does state relate to what your crews actually cost?Directly, whenever state values are interpolated into a crew's inputs — they become prompt tokens. Flows that accumulate every intermediate output on state and pass the whole thing forward grow their prompts monotonically. Keep bulky payloads out of state, or store a reference such as a file path that only the step needing it dereferences.
- Can flow state survive between runs?Yes — CrewAI's @persist decorator, backed by SQLiteFlowPersistence, saves flow state so a run can be restored instead of restarted from scratch, keyed on the state's id. Structured state makes this safer, since a declared schema round-trips predictably while an open dict may come back in a shape the current code no longer expects.
saying these in an interview costs you the question
- Claiming dict state validates keys or types
- Thinking each step gets its own copy of state
- Assuming parallel branches are serialized on state writes
- Saying only structured state has an id
- Treating state as free because it is not a prompt