skip to content

Haystack

You will learn deepset's pipeline framework for production RAG and search: explicitly wired components, generators, document stores, retrievers and rankers, and agents with tools. Interviewers ask because Haystack's explicit component graph makes a pipeline inspectable and testable, which is the property production teams value over convenience.

part ofAI agent & RAG frameworksoverview, primer and where to startread it →
on this pageshow

explore

questions

page 1 of 2

What does Haystack's Agent component add over calling a ChatGenerator yourself?

level: juniorimportance: must knowfreq 66%

answer

  1. one component, not one model call
  2. who executes the tool calls
  3. the loop and its stopping rule
  4. ToolInvoker runs inside it
  5. messages plus last_message come back

basics

~20 s

Haystack's Agent packages the whole tool-calling loop into one component: it sends messages to the chat generator, executes requested tools through an internal ToolInvoker, feeds the results back, and stops on its exit conditions or step limit.

solid answer

~40 s

A `ChatGenerator` is single-shot: you give it messages, it returns replies, and if those replies contain tool calls you have to execute them and call the generator again yourself. `Agent(chat_generator=..., tools=[...], system_prompt=...)` wraps that cycle. On `run(messages=[...])` it appends the system prompt, calls the generator, hands any tool calls to an internal `ToolInvoker`, appends the resulting tool messages, and loops until an exit condition matches (default: the model returns a plain text reply) or `max_agent_steps` is reached. It returns a dict with `messages` (the full conversation, including tool calls and tool results) and `last_message`. Crucially, `Agent` is itself a Haystack component, so the same object can be run standalone or dropped into a `Pipeline` — you do not choose between the agent model and the pipeline model.

code

python · 22 lines
python
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.dataclasses import ChatMessage
from haystack.tools import tool


@tool
def get_weather(city: str) -> str:
    """Return the current weather for a city."""
    return f"Sunny in {city}"


agent = Agent(
    chat_generator=OpenAIChatGenerator(model="gpt-4o-mini"),
    tools=[get_weather],
    system_prompt="You are a concise weather assistant.",
)
agent.warm_up()

result = agent.run(messages=[ChatMessage.from_user("Weather in Berlin?")])
print(result["last_message"].text)
print(len(result["messages"]))

go deeper

for a junior

Be able to say that a chat generator only returns tool calls, and that the Agent is the component that actually executes them and loops until the model answers in plain text.

for a middle

Explain the loop concretely: generator, internal ToolInvoker, tool messages appended, repeat, bounded by exit conditions and a step limit; and name the messages and last_message outputs.

for a senior

Talk about cost and debuggability — the transcript grows every step, tool-heavy runs are expensive, and the messages list is your primary trace when an agent picks the wrong tool in production.

for a principal

Frame the design bet: the Agent confines model-driven control flow to one node of an otherwise explicit graph. Be ready to argue when that confinement is worth its nondeterminism and when a deterministic wiring is the better engineering choice.

## The problem the Agent solves A Haystack chat generator such as `OpenAIChatGenerator` is a one-shot component. You pass it `messages` (a list of `ChatMessage`) plus a list of tools, and it returns `replies`. If the model decided to call a tool, those replies contain `ToolCall` objects — a name and an argument dict — and nothing has actually run yet. Executing the call, wrapping the result as a tool message, appending it to the conversation and calling the generator again is entirely on you. Doing that by hand means writing a `while` loop, deciding when to stop, and bounding it so a confused model cannot spin forever. `Agent` is the component that owns that loop. ## What the constructor takes ``` Agent( chat_generator=..., # any Haystack ChatGenerator tools=[...], # list[Tool] or a Toolset system_prompt="...", # prepended as a system message exit_conditions=["text"], state_schema={...}, max_agent_steps=100, streaming_callback=None, raise_on_tool_invocation_failure=False, ) ``` Only `chat_generator` is required. `tools` may also be left to the generator if it was already configured with them, but passing them to the `Agent` is the normal path because the `Agent` needs them for the invoker as well as the schema. ## The loop, step by step 1. The incoming `messages` are taken as the starting conversation, with `system_prompt` prepended as a system message if you set one. 2. The chat generator is called with the conversation and the tool schemas. 3. If the reply is a plain assistant text message and `"text"` is in `exit_conditions`, the loop ends. 4. If the reply contains tool calls, they are handed to an internal `ToolInvoker`, which looks each name up in the tool list, validates and passes the arguments, runs the underlying callable or component, and produces one tool message per call. 5. Those tool messages are appended and the loop returns to step 2. A step counter is incremented each time; when it reaches `max_agent_steps` the agent stops and returns what it has. ## What you get back `run()` returns a dict. `messages` is the full transcript — user, system, assistant messages including their tool calls, and the tool result messages — which is what you inspect when debugging why the agent did what it did. `last_message` is the final message, normally the model's answer, and is what you usually render. If you declared a `state_schema`, the state values appear as additional keys in the same output dict, so data a tool produced can leave the agent without ever having been serialized into the prompt. `Agent` also has `run_async()`, and a `warm_up()` that warms the chat generator — a `Pipeline` calls that for you. ## What it deliberately does not do The `Agent` does not plan, does not decompose tasks, and does not summarize or trim history for you: the conversation grows monotonically until the loop ends, so a long tool-heavy run is also an expensive one. It does not choose tools — the model does — and it does not validate that a tool result is correct. It also does not spawn other agents; multi-agent behaviour, when you want it, is expressed by making one agent a tool or a component that another pipeline drives. ## Why the component framing matters Because `Agent` implements the component interface, its loop is invisible from outside: to a surrounding `Pipeline` it is one node that consumes `messages` and produces `messages`/`last_message`. The internal iteration is not pipeline iteration, so you do not model it with edges or run-limits at the pipeline level. This is the design bet Haystack makes — the agent is a bounded region of model-driven control flow embedded in an otherwise explicit, inspectable graph, rather than an alternative to it. ## When to hand-roll instead If the control flow is actually known — retrieve, then build a prompt, then generate — an agent adds nondeterminism, latency and cost for nothing. Reach for `Agent` when the number and order of tool calls genuinely depends on the input.

  • How would you stream an agent's output to a user while it is still working?
    Pass `streaming_callback` to `Agent` (either in the constructor or per `run()` call). The callback receives streaming chunks from the chat generator as they arrive, including tool-call deltas, so you can render partial assistant text and show which tool is being invoked instead of waiting for the whole loop to finish. The final `messages`/`last_message` are still returned when `run()` completes.
  • Does the Agent need the same tools to be configured on the chat generator too?
    No. Pass the tools to the `Agent` and it supplies the schemas to the generator on each call and routes the resulting calls to its internal invoker. Configuring tools in two places is a common source of drift — a tool the generator advertises but the agent cannot invoke produces a call the invoker rejects.
  • What is in the messages list that is not in last_message?
    Everything the loop produced: the system prompt, the user turns, every assistant message including its `ToolCall` objects, and every tool result message. `last_message` is just the final element. When an agent misbehaves, `messages` is the trace you read — it shows which tool it picked, with which arguments, and what came back.

saying these in an interview costs you the question

  • Thinking a ChatGenerator executes tool calls by itself
  • Believing the Agent replaces the Pipeline rather than sitting inside one
  • Assuming the agent plans or decomposes the task for you
  • Expecting only the final answer back and ignoring the messages transcript
  • Thinking the agent trims or summarizes conversation history automatically

context

open as a page

In Haystack, what does ChatPromptBuilder do inside a RAG pipeline?

level: juniorimportance: must knowfreq 62%

basics

~20 s

ChatPromptBuilder renders a Jinja2 template into a list of ChatMessage objects. Every placeholder in the template becomes an input it expects at run time, and the rendered messages leave on its prompt output, ready for a chat generator.

open as a page

Which components make up a Haystack indexing pipeline, and what does each do?

level: juniorimportance: must knowfreq 72%

basics

~10 s

A Haystack indexing pipeline runs a converter (file bytes to Document), then DocumentCleaner (strip whitespace and boilerplate), DocumentSplitter (cut into smaller Documents), a document embedder (attach vectors), and DocumentWriter (persist to a DocumentStore).

open as a page

How do you shape the input dict for Haystack's Pipeline.run(), and what does it return?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Pipeline.run() takes a dict keyed by component name, whose value is that component's run() keyword arguments. It returns a dict, also keyed by component name, holding the outputs no other component consumed. Use include_outputs_from to also get intermediate outputs.

open as a page

In Haystack, how does wiring InMemoryBM25Retriever differ from InMemoryEmbeddingRetriever?

level: juniorimportance: must knowfreq 74%

basics

~20 s

InMemoryBM25Retriever takes the raw query string on its query input and scores lexically over document text. InMemoryEmbeddingRetriever takes a query_embedding, so a text embedder must run first in the pipeline and feed its vector into that socket.

open as a page

How do you embed a Haystack Agent inside a Pipeline, and what sockets does it expose?

level: middleimportance: must knowfreq 50%

basics

~20 s

Agent implements the component interface, so you add it with Pipeline.add_component and connect into its messages input and out of its messages or last_message outputs. Its tool loop runs entirely inside one component execution — the pipeline never sees the iterations.

open as a page

What does Haystack's ComponentTool wrap, and where does its schema come from?

level: middleimportance: must knowfreq 54%

basics

~20 s

ComponentTool turns any Haystack component into a tool the model can call. It derives the JSON parameter schema from the component's run() signature and type hints, uses the docstring or an explicit description as the tool description, and converts the component's output to a string for the model.

open as a page

In Haystack's Agent, what do exit_conditions and max_agent_steps control?

level: middleimportance: must knowfreq 58%

basics

~20 s

They bound the Agent's tool loop. exit_conditions lists what ends it — by default ["text"], meaning any plain text reply; adding a tool name ends the run right after that tool executes. max_agent_steps (default 100) is the hard ceiling that stops a runaway loop.

open as a page

In Haystack 3.0, what replaced OpenAIGenerator and what must you change?

level: middleimportance: must knowfreq 55%

basics

~20 s

Haystack 3.0 removed the legacy string-in generators, OpenAIGenerator among them. OpenAIChatGenerator replaces it: it accepts ChatMessage objects, or a plain string, and returns replies as ChatMessage objects, so downstream code must read reply.text instead of a raw string.

open as a page

In Haystack, what does the @component decorator require of your class?

level: middleimportance: must knowfreq 78%

basics

~20 s

A Haystack component is a class decorated with @component that exposes a run() method. The typed parameters of run() become its inputs, and run() must return a dict whose keys match the names declared in @component.output_types on that method.

open as a page

In Haystack, what does DocumentWriter's DuplicatePolicy control?

level: middleimportance: must knowfreq 62%

basics

~20 s

DuplicatePolicy tells a Haystack DocumentWriter what to do when an incoming document's id already exists in the store: OVERWRITE replaces the stored copy, SKIP keeps it, FAIL raises DuplicateDocumentError, and NONE defers to the store's own default.

open as a page

In Haystack, how do you express AND/OR conditions in a metadata filter?

level: middleimportance: must knowfreq 58%

basics

~10 s

Haystack filters are nested dicts of two node shapes. A comparison node is {"field", "operator", "value"}; a logical node is {"operator": "AND"|"OR"|"NOT", "conditions": [...]} holding further nodes. Metadata fields must be written as meta.<key>.

open as a page

In Haystack, what does Pipeline.connect() validate, and when does it fail?

level: middleimportance: must knowfreq 75%

basics

~20 s

Pipeline.connect() checks at wiring time that both named sockets exist and that the sender's declared output type is compatible with the receiver's input type. A mismatch raises immediately, so a broken graph fails before any model is called.

open as a page

In Haystack, how does DocumentJoiner's reciprocal_rank_fusion mode build hybrid retrieval?

level: middleimportance: must knowfreq 66%

basics

~20 s

You wire a BM25 retriever and an embedding retriever into the same DocumentJoiner, whose documents input is variadic and accepts several connections. With join_mode set to reciprocal_rank_fusion the joiner deduplicates by document id and replaces each score with one derived from the document's rank in every incoming list.

open as a page

Why does a Haystack component load its model in warm_up(), not __init__?

level: middleimportance: should knowfreq 42%

basics

~20 s

warm_up() separates cheap construction from expensive resource loading. A component can be built, wired, type-checked, serialized and drawn without downloading weights or claiming GPU memory; the model loads only when the component is about to actually run.

open as a page

What does meta_fields_to_embed do on a Haystack document embedder?

level: middleimportance: should knowfreq 38%

basics

~20 s

It prepends the named metadata values to the document text before the embedding model sees it, joined by embedding_separator. The vector then encodes that metadata, while the stored Document's content and meta are left unchanged.

open as a page

How do you serialize a Haystack Pipeline to YAML, and what fails to round-trip?

level: middleimportance: should knowfreq 38%

basics

~20 s

Pipeline.dumps() emits YAML describing every component's import path, init parameters and the connections; Pipeline.loads() rebuilds it. Round-trips break when a component's module is not importable in the loading process, when init parameters are not serializable, or when secrets were passed as raw literals.

open as a page

In Haystack, why must TransformersSimilarityRanker be warmed up before it ranks?

level: middleimportance: should knowfreq 52%

basics

~20 s

Its constructor only records the model name and device so components stay cheap to build and serialise. The model is downloaded and loaded onto the device in warm_up(), which a pipeline calls for you before the first run; calling run() standalone without it raises an error.

open as a page

In Haystack retrievers, how do init-time top_k and filters differ from run-time ones?

level: middleimportance: should knowfreq 44%

basics

~20 s

Constructor values are the component's defaults and are captured in its serialised form. Values passed in Pipeline.run's per-component input dict apply to that run only. top_k always overrides, while filters are combined or replaced according to the retriever's filter_policy setting.

open as a page

How does Haystack's Agent State pass data between steps without going through the LLM?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Declare typed slots with the Agent's state_schema; tools write into them with outputs_to_state and read from them with inputs_from_state. Values stay as real Python objects across steps and are returned as agent outputs, instead of being stringified into the prompt.

open as a page

When a tool raises during a Haystack Agent run, what happens by default?

level: seniorimportance: should knowfreq 38%

basics

~20 s

By default the Agent does not propagate the exception. The error is captured and returned to the model as a tool result message, so the model can correct its arguments or pick another tool. Setting raise_on_tool_invocation_failure=True makes the run abort instead.

open as a page

Why does Haystack take an API key as Secret.from_env_var, not a string?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Secret.from_env_var stores only the variable's name, so serializing a component or exporting a pipeline to YAML writes a reference rather than the key. Secret.from_token holds the literal value and refuses to serialize at all, which is deliberate.

open as a page

When must a custom Haystack component implement to_dict and from_dict?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Only when default serialization cannot round-trip your constructor arguments. Haystack's default maps each init parameter to a same-named attribute and needs JSON-friendly values, so sets, callables, custom objects and secrets require explicit conversion in to_dict and reconstruction in from_dict.

open as a page

In Haystack, what changes when you swap InMemoryDocumentStore for a database store?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The DocumentStore protocol keeps write_documents, filter_documents, count_documents and delete_documents identical, so the indexing pipeline is unchanged. What changes is the store class and its connection config, the fixed vector dimension, the store-specific retriever, and who owns durability.

open as a page

Why does a nightly Haystack indexing pipeline keep adding duplicate chunks?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Because a Haystack Document's auto-generated id is a hash of its content and meta. Any change to the extracted text, splitter settings or metadata produces new ids, so OVERWRITE finds nothing to replace and writes the chunks as new rows.

open as a page

How does a Haystack pipeline branch on a runtime value and rejoin the branches?

level: seniorimportance: should knowfreq 52%

basics

~20 s

ConditionalRouter evaluates Jinja conditions over its inputs and emits on only the matching route's output socket, so the untaken branch never runs. To merge branches back, route them through a joiner with a variadic input, because a plain input socket accepts one connection.

open as a page

What makes a looping Haystack pipeline terminate, and what caps a runaway loop?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A loop terminates because a router inside it has an exit route that eventually fires. As a backstop, Pipeline(max_runs_per_component=...) caps how many times any single component may run in one call — default 100 — and exceeding it aborts the run with an error instead of spinning forever.

open as a page

How do you design document metadata in Haystack for tenant-filtered retrieval?

level: principalimportance: should knowfreq 30%

basics

~20 s

Attach the scoping keys at conversion time so DocumentSplitter copies them into every chunk, keep them low-cardinality scalars the store can index, and make the tenant condition a mandatory outermost AND assembled by shared code rather than by each caller.

open as a page

When would you split one Haystack pipeline into several rather than adding more branches?

level: principalimportance: should knowfreq 30%

basics

~20 s

Split when parts of the graph have different lifecycles, owners or failure budgets — indexing versus query being the standard case. They share a document store rather than an edge. Keep one pipeline when the whole flow is a single request's dataflow.

open as a page

In Haystack's SentenceTransformersDiversityRanker, how do its two strategies differ?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

With strategy="greedy_diversity_order" it repeatedly picks the candidate least similar to what it has already selected, so the output spreads out. With strategy="maximum_margin_relevance" each pick balances query relevance against similarity to the selection so far, tuned by lambda_threshold.

open as a page

showing 1–30 of 31