What does Haystack's Agent component add over calling a ChatGenerator yourself?
answer
- one component, not one model call
- who executes the tool calls
- the loop and its stopping rule
- ToolInvoker runs inside it
- messages plus last_message come back
basics
~20 sHaystack's Agent packages the whole tool-calling loop into one component: it sends messages to the chat generator, executes requested tools through an internal ToolInvoker, feeds the results back, and stops on its exit conditions or step limit.
solid answer
~40 sA `ChatGenerator` is single-shot: you give it messages, it returns replies, and if those replies contain tool calls you have to execute them and call the generator again yourself. `Agent(chat_generator=..., tools=[...], system_prompt=...)` wraps that cycle. On `run(messages=[...])` it appends the system prompt, calls the generator, hands any tool calls to an internal `ToolInvoker`, appends the resulting tool messages, and loops until an exit condition matches (default: the model returns a plain text reply) or `max_agent_steps` is reached. It returns a dict with `messages` (the full conversation, including tool calls and tool results) and `last_message`. Crucially, `Agent` is itself a Haystack component, so the same object can be run standalone or dropped into a `Pipeline` — you do not choose between the agent model and the pipeline model.
code
python · 22 linesfrom haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.dataclasses import ChatMessage
from haystack.tools import tool
@tool
def get_weather(city: str) -> str:
"""Return the current weather for a city."""
return f"Sunny in {city}"
agent = Agent(
chat_generator=OpenAIChatGenerator(model="gpt-4o-mini"),
tools=[get_weather],
system_prompt="You are a concise weather assistant.",
)
agent.warm_up()
result = agent.run(messages=[ChatMessage.from_user("Weather in Berlin?")])
print(result["last_message"].text)
print(len(result["messages"]))go deeper
Be able to say that a chat generator only returns tool calls, and that the Agent is the component that actually executes them and loops until the model answers in plain text.
Explain the loop concretely: generator, internal ToolInvoker, tool messages appended, repeat, bounded by exit conditions and a step limit; and name the messages and last_message outputs.
Talk about cost and debuggability — the transcript grows every step, tool-heavy runs are expensive, and the messages list is your primary trace when an agent picks the wrong tool in production.
Frame the design bet: the Agent confines model-driven control flow to one node of an otherwise explicit graph. Be ready to argue when that confinement is worth its nondeterminism and when a deterministic wiring is the better engineering choice.
## The problem the Agent solves A Haystack chat generator such as `OpenAIChatGenerator` is a one-shot component. You pass it `messages` (a list of `ChatMessage`) plus a list of tools, and it returns `replies`. If the model decided to call a tool, those replies contain `ToolCall` objects — a name and an argument dict — and nothing has actually run yet. Executing the call, wrapping the result as a tool message, appending it to the conversation and calling the generator again is entirely on you. Doing that by hand means writing a `while` loop, deciding when to stop, and bounding it so a confused model cannot spin forever. `Agent` is the component that owns that loop. ## What the constructor takes ``` Agent( chat_generator=..., # any Haystack ChatGenerator tools=[...], # list[Tool] or a Toolset system_prompt="...", # prepended as a system message exit_conditions=["text"], state_schema={...}, max_agent_steps=100, streaming_callback=None, raise_on_tool_invocation_failure=False, ) ``` Only `chat_generator` is required. `tools` may also be left to the generator if it was already configured with them, but passing them to the `Agent` is the normal path because the `Agent` needs them for the invoker as well as the schema. ## The loop, step by step 1. The incoming `messages` are taken as the starting conversation, with `system_prompt` prepended as a system message if you set one. 2. The chat generator is called with the conversation and the tool schemas. 3. If the reply is a plain assistant text message and `"text"` is in `exit_conditions`, the loop ends. 4. If the reply contains tool calls, they are handed to an internal `ToolInvoker`, which looks each name up in the tool list, validates and passes the arguments, runs the underlying callable or component, and produces one tool message per call. 5. Those tool messages are appended and the loop returns to step 2. A step counter is incremented each time; when it reaches `max_agent_steps` the agent stops and returns what it has. ## What you get back `run()` returns a dict. `messages` is the full transcript — user, system, assistant messages including their tool calls, and the tool result messages — which is what you inspect when debugging why the agent did what it did. `last_message` is the final message, normally the model's answer, and is what you usually render. If you declared a `state_schema`, the state values appear as additional keys in the same output dict, so data a tool produced can leave the agent without ever having been serialized into the prompt. `Agent` also has `run_async()`, and a `warm_up()` that warms the chat generator — a `Pipeline` calls that for you. ## What it deliberately does not do The `Agent` does not plan, does not decompose tasks, and does not summarize or trim history for you: the conversation grows monotonically until the loop ends, so a long tool-heavy run is also an expensive one. It does not choose tools — the model does — and it does not validate that a tool result is correct. It also does not spawn other agents; multi-agent behaviour, when you want it, is expressed by making one agent a tool or a component that another pipeline drives. ## Why the component framing matters Because `Agent` implements the component interface, its loop is invisible from outside: to a surrounding `Pipeline` it is one node that consumes `messages` and produces `messages`/`last_message`. The internal iteration is not pipeline iteration, so you do not model it with edges or run-limits at the pipeline level. This is the design bet Haystack makes — the agent is a bounded region of model-driven control flow embedded in an otherwise explicit, inspectable graph, rather than an alternative to it. ## When to hand-roll instead If the control flow is actually known — retrieve, then build a prompt, then generate — an agent adds nondeterminism, latency and cost for nothing. Reach for `Agent` when the number and order of tool calls genuinely depends on the input.
- How would you stream an agent's output to a user while it is still working?Pass `streaming_callback` to `Agent` (either in the constructor or per `run()` call). The callback receives streaming chunks from the chat generator as they arrive, including tool-call deltas, so you can render partial assistant text and show which tool is being invoked instead of waiting for the whole loop to finish. The final `messages`/`last_message` are still returned when `run()` completes.
- Does the Agent need the same tools to be configured on the chat generator too?No. Pass the tools to the `Agent` and it supplies the schemas to the generator on each call and routes the resulting calls to its internal invoker. Configuring tools in two places is a common source of drift — a tool the generator advertises but the agent cannot invoke produces a call the invoker rejects.
- What is in the messages list that is not in last_message?Everything the loop produced: the system prompt, the user turns, every assistant message including its `ToolCall` objects, and every tool result message. `last_message` is just the final element. When an agent misbehaves, `messages` is the trace you read — it shows which tool it picked, with which arguments, and what came back.
saying these in an interview costs you the question
- Thinking a ChatGenerator executes tool calls by itself
- Believing the Agent replaces the Pipeline rather than sitting inside one
- Assuming the agent plans or decomposes the task for you
- Expecting only the final answer back and ignoring the messages transcript
- Thinking the agent trims or summarizes conversation history automatically