skip to content

In Haystack's Agent, what do exit_conditions and max_agent_steps control?

level: middleimportance: must knowfreq 58%

answer

  1. intended stop vs emergency stop
  2. default is a plain text reply
  3. a tool name can be a stop signal
  4. the ceiling on generate-invoke passes
  5. hitting the cap warns, does not raise

basics

~20 s

They bound the Agent's tool loop. exit_conditions lists what ends it — by default ["text"], meaning any plain text reply; adding a tool name ends the run right after that tool executes. max_agent_steps (default 100) is the hard ceiling that stops a runaway loop.

solid answer

~50 s

Haystack's `Agent` keeps calling its chat generator and invoking tools until something tells it to stop, and there are exactly two mechanisms. `exit_conditions` is a list of strings: the default `["text"]` means the loop ends as soon as the model produces an assistant message with no tool calls. Put a **tool name** in the list instead — e.g. `exit_conditions=["final_answer"]` for a tool of that name — and the agent returns immediately after that tool has been invoked, which is how you build a deterministic hand-off out of the loop instead of hoping the model stops talking. You can list several conditions. `max_agent_steps` (default 100) is the safety net: when the step counter reaches it the agent logs a warning, breaks out, and returns the messages collected so far rather than raising. Set it low in production — a runaway loop is paid for in tokens and latency, not just wall time.

code

python · 9 lines
python
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIChatGenerator

agent = Agent(
    chat_generator=OpenAIChatGenerator(model="gpt-4o-mini"),
    tools=[search_tool, final_answer_tool],
    exit_conditions=["text", "final_answer"],
    max_agent_steps=8,
)

go deeper

for a junior

Know that the default exit condition is the model replying with plain text, and that max_agent_steps is a safety limit with a default of 100.

for a middle

Explain both mechanisms and their difference, including that a tool name in exit_conditions ends the run right after that tool executes, and that the cap returns rather than raises.

for a senior

Show production judgment: lower the cap deliberately, detect cap-terminated runs from the returned messages, and recognize that failing tools consume steps because errors are fed back to the model.

for a principal

Own the budget argument — each step re-sends the growing transcript, so step limits are a cost control as much as a liveness one. Be ready to specify per-run token and step budgets, and a typed exit tool for contract stability.

## Two different kinds of stop A tool-calling loop needs an *intended* stop and an *emergency* stop, and Haystack's `Agent` separates them cleanly. `exit_conditions` describes the outcome you are waiting for. `max_agent_steps` describes how much you are willing to spend before giving up. Conflating them is the classic mistake: people set `max_agent_steps=3` to "make it stop after the answer", and then wonder why complex inputs return half-finished transcripts. ## exit_conditions The parameter is a `list[str]`, defaulting to `["text"]`. - `"text"` matches whenever the chat generator returns an assistant message that contains no tool calls — a plain text answer. That is the ordinary end of a ReAct-shaped run: the model has stopped asking for tools and has said something. - Any other entry is interpreted as a **tool name**. If the model calls a tool with that name, the tool runs and the agent then returns instead of feeding the result back to the model. The final message is that tool's result message. The tool-name form is how you make an agent's exit deterministic. A common pattern is a `final_answer` tool with a typed argument schema: the model must call it to finish, so the shape of the answer is enforced by the tool's parameters rather than by prompt pleading. It is also how you hand control back to a surrounding pipeline — the agent stops the moment the interesting tool has produced its output, and the pipeline continues with that value. A subtlety: with a tool name as the only exit condition, a model that simply answers in prose never satisfies the condition, so the loop keeps going until the step cap fires. Most setups therefore keep `"text"` in the list alongside the tool name unless the strictness is exactly what you want. ## max_agent_steps Each pass through generate-then-invoke increments a counter. When it reaches `max_agent_steps` (default 100), the agent stops the loop, logs a warning, and returns the conversation as it stands. It does **not** raise, and it does not synthesize an answer for you. That means a run that hit the cap looks superficially like a normal result — the caller must be prepared for `last_message` to be a tool result or a half-finished assistant turn rather than an answer. Why the default is high: the framework has no idea whether your task needs two tool calls or forty. Why you should lower it: every step is at least one model call whose prompt contains the entire transcript so far, so cost grows super-linearly with steps. A loop that ping-pongs between two tools for 100 steps is an expensive way to produce nothing. ## Reading a run that ended badly Because both stops return normally, the way to tell them apart is the returned `messages`. If the last element is a plain assistant text message, the `"text"` condition fired. If it is a tool result message, either a tool-name condition fired or you hit the cap. Counting assistant turns against `max_agent_steps` distinguishes those two. In production the pragmatic move is to check `last_message` for a tool-call-free assistant message and treat anything else as a failed run, plus alert on the warning the agent logs when the cap is hit. ## Interaction with tool errors A failing tool does not end the loop by default: the error comes back as a tool message and the model gets another chance, consuming a step. So a persistently broken tool is precisely the situation where `max_agent_steps` earns its keep — without it, a model that keeps retrying a 500-ing endpoint burns your budget until something else times out. ## Practical settings For a retrieval-and-answer agent, two or three tool calls is typical, so `max_agent_steps` in the 6-10 range gives generous headroom and still fails fast. For an open-ended research agent, higher — but pair it with per-run token accounting rather than trusting the step count alone, since one step with a 100k-token transcript costs more than ten small ones.

  • What does the agent return if it hits max_agent_steps?
    It logs a warning, breaks the loop and returns normally — `messages` holds everything collected so far and `last_message` is whatever the final turn happened to be, often a tool result. No exception is raised, so callers must detect the case themselves, typically by checking that `last_message` is a tool-call-free assistant message before treating the run as successful.
  • Why would you add a tool name to exit_conditions instead of relying on "text"?
    To make the exit deterministic and typed. A `final_answer`-style tool forces the model to emit its result through a schema you control, and the agent returns the moment that tool runs. It also gives a surrounding pipeline a predictable value to consume, instead of free-form prose that a downstream component has to parse.
  • If exit_conditions contains only a tool name, what is the risk?
    A model that answers in prose never triggers the condition, so the loop continues until `max_agent_steps` fires and the run ends without a satisfied exit. Unless you deliberately want that strictness — and have prompt and tool-choice settings that make the tool call near-certain — keep `"text"` in the list as well.

saying these in an interview costs you the question

  • Using max_agent_steps as the normal way to end a run
  • Assuming hitting the step cap raises an exception
  • Thinking exit_conditions accepts arbitrary regexes or prompts
  • Believing a tool error ends the loop by default
  • Leaving max_agent_steps at 100 in a cost-sensitive service

context