skip to content

How does Haystack's Agent State pass data between steps without going through the LLM?

level: seniorimportance: should knowfreq 42%

answer

  1. a side channel next to the messages
  2. typed slots declared up front
  3. tools write in, tools read out
  4. the model never has to supply it
  5. it becomes an agent output socket

basics

~20 s

Declare typed slots with the Agent's state_schema; tools write into them with outputs_to_state and read from them with inputs_from_state. Values stay as real Python objects across steps and are returned as agent outputs, instead of being stringified into the prompt.

solid answer

~50 s

By default everything a tool produces reaches the next step only as text inside a tool message, which means real objects — `Document` lists, dataframes, ids — get serialized into the prompt and must be re-parsed later. Haystack's `State` gives the agent a side channel. You declare it with `state_schema={"documents": {"type": list[Document]}}`, optionally with a `handler` that says how repeated writes merge (append versus replace). A tool declares `outputs_to_state={"documents": {"source": "documents"}}` to route part of its output into that slot, and `inputs_from_state={"documents": "documents"}` to have a value injected as an argument the model never has to supply — and therefore cannot hallucinate. Each declared key also becomes an output of the agent, so a downstream pipeline component receives the actual objects. Practically this cuts token cost (you can send a lean rendering to the model while keeping the full objects in state) and removes a whole class of parsing failures.

code

python · 20 lines
python
from haystack import Document
from haystack.components.agents import Agent
from haystack.tools import ComponentTool

search_tool = ComponentTool(
    component=retriever,
    name="search_docs",
    description="Search the handbook for passages answering a question.",
    outputs_to_state={"documents": {"source": "documents"}},
    outputs_to_string={"source": "documents", "handler": render_titles},
)

agent = Agent(
    chat_generator=chat_generator,
    tools=[search_tool],
    state_schema={"documents": {"type": list[Document]}},
)

result = agent.run(messages=[user_message])
cited = result["documents"]  # real Document objects, never re-parsed

go deeper

for a junior

Know that an agent can carry typed data between steps in a state object declared by state_schema, separate from the chat messages.

for a middle

Explain outputs_to_state and inputs_from_state concretely, and that each declared key also becomes an output of the agent for downstream components.

for a senior

Show the production payoff: real objects for citations and downstream ranking, injected values the model cannot invent, and lean tool renderings that cut tokens re-sent every step. Name the last-write-wins hazard.

for a principal

Own the boundary: decide what belongs in the prompt versus the side channel, how request-scoped identity is injected rather than prompted, and how run-scoped state relates to whatever durable memory the system keeps.

## The problem: the prompt is a lossy bus In a plain tool loop, the only channel between steps is the conversation. A retrieval tool returns `Document` objects; those get stringified into a tool message; the model reads them; three steps later something downstream wants the documents back and has to reconstruct them from prose. Two costs follow. The obvious one is tokens: the serialized blob is re-sent with every subsequent model call, so a fat rendering is paid for repeatedly. The subtle one is fidelity: ids, scores and metadata either bloat the prompt or are lost. `State` is Haystack's answer — a typed key-value store that lives for the duration of an agent run, alongside the messages rather than inside them. ## Declaring the schema `state_schema` is a dict of slot name to descriptor: ``` state_schema={ "documents": {"type": list[Document]}, "customer_id": {"type": str}, } ``` The `type` is what the agent uses to expose a correctly typed output socket, so a surrounding pipeline can connect to `agent.documents` and type checking works. A descriptor may also carry a `handler`: a callable taking the current value and the incoming one and returning the merged result. That is how you choose between "the latest search replaces the previous documents" and "accumulate documents across searches" — for list-typed slots, accumulation is the common intent and a handler makes it explicit rather than accidental. ## Writing: outputs_to_state A `Tool` or `ComponentTool` declares where its output goes: ``` outputs_to_state={"documents": {"source": "documents"}} ``` The key is the state slot, `source` names the field of the tool's output dict. An optional `handler` lets you transform on the way in. This is independent of `outputs_to_string`, which controls what the *model* sees. The pairing is the point: send the model a compact rendering — titles and snippets — while the full objects go to state. ## Reading: inputs_from_state `inputs_from_state={"customer_id": "customer_id"}` binds a state slot to a tool argument. The framework supplies it at invocation time, so it is not something the model is asked for. Two things follow. First, correctness: a tenant id, an auth scope, or a filter that must be exact is no longer subject to model invention. Second, prompt economy: the parameter does not have to be advertised and argued about. This is the standard way to pin per-request context — set the slot when you start the run, and every tool that needs it gets it. ## Seeding and reading state around a run State keys can be provided when you run the agent (as additional keyword inputs alongside `messages`) and are returned in the result dict alongside `messages` and `last_message`. So a run is: seed the request-scoped facts, let the loop enrich them, read the structured result out the other end. Inside a pipeline, those same keys are output sockets, which is what lets an agent feed a ranker or a citation renderer with real objects. ## What State is not It is **run-scoped**, not durable memory: it lives for the invocation and is not automatically persisted across separate `run()` calls or user sessions. Long-lived conversation memory is a different concern, handled by whatever store you keep messages in. Nor is it shared across concurrent runs — each run gets its own, which is exactly what you want for request isolation, but it means it is not a cache. ## Failure modes worth knowing - **Undeclared slot.** Writing to a key that is not in `state_schema` is a wiring bug; declare every slot a tool targets. - **Silent overwrite.** Two tools writing the same slot without a merge handler means last-write-wins and quietly lost documents. If a slot can be written more than once, specify the handler. - **Unbounded growth.** An accumulating handler on a slot fed by a search tool the model calls twenty times holds twenty result sets in memory. Cap it in the handler. - **Assuming the model sees state.** It does not, unless something renders it into a message. State that must influence the model's reasoning has to be surfaced deliberately. ## Why this is a senior question It is the difference between an agent that emits a paragraph and an agent that is a composable component. Citations, structured outputs, per-tenant scoping and token control all run through this one mechanism.

  • Two different tools write the same state slot. What decides the result?
    The slot's handler in `state_schema`. Without one you effectively get last-write-wins, which silently discards the earlier value — a classic cause of "the agent found the documents but they vanished". Declare a merge handler (append for lists, or a custom reconcile) for any slot more than one tool can write, and bound growth inside it.
  • How is state different from the conversation history?
    History is text the model reads and re-reads on every step, so it costs tokens repeatedly and only carries strings. State holds real Python objects, is invisible to the model unless you render it, and is exposed as typed agent outputs. Use history for what the model must reason over, state for what the system must keep.
  • Does state survive between separate runs of the agent?
    No — it is scoped to a single run. Nothing persists it across invocations or sessions automatically, and concurrent runs each get their own, which gives you request isolation. Durable memory across turns is a separate design: persist the messages and any needed slots yourself, and seed the next run's state from that store.
  • How do you keep a tenant id out of the model's hands entirely?
    Seed it into a state slot at the start of the run and bind it with `inputs_from_state` on every tool that needs it. The framework injects the value at invocation, so it never appears as a model-supplied argument and cannot be hallucinated or overridden by prompt content — which also makes it the right place to enforce per-request scoping.

saying these in an interview costs you the question

  • Believing the model can read state directly without it being rendered
  • Treating state as durable memory across sessions
  • Letting two tools write one slot with no merge handler
  • Round-tripping documents through the prompt instead of state
  • Assuming inputs_from_state values are still requested from the model

context