skip to content

Pipelines & Wiring

You will learn the Pipeline object at Haystack's core: components added by name, connections declared between typed output and input sockets, and validation that fails at wiring time rather than mid-run. Interviewers ask about branching, joining, and loops, because that is where pipelines stop being linear.

part ofAI agent & RAG frameworksoverview, primer and where to startread it →
on this pageshow

questions

6

How do you shape the input dict for Haystack's Pipeline.run(), and what does it return?

level: juniorimportance: must knowfreq 70%

answer

  1. Keys are the names you registered
  2. Connections fill the rest for you
  3. Result carries only unconsumed outputs
  4. Ask explicitly for intermediates
  5. One flag widens what comes back

basics

~20 s

Pipeline.run() takes a dict keyed by component name, whose value is that component's run() keyword arguments. It returns a dict, also keyed by component name, holding the outputs no other component consumed. Use include_outputs_from to also get intermediate outputs.

solid answer

~40 s

The input is a nested dict: `pipe.run({"retriever": {"query": q}, "prompt": {"question": q}})`. Each top-level key is the name you passed to `add_component`, and the inner dict is that component's `run()` keyword arguments. You only supply values that do **not** arrive over a connection — anything wired from an upstream socket is filled in by the pipeline, and supplying it yourself for a connected socket is an error. The return value mirrors that shape: a dict keyed by component name, containing the outputs that no downstream component consumed — typically just your terminal component. Intermediate outputs are dropped by default so results stay small; pass `include_outputs_from={"retriever"}` to get a specific component's output back as well. A flat dict of inputs is also accepted and distributed to every component that declares a parameter of that name.

code

python · 19 lines
python
from haystack import Document, Pipeline
from haystack.components.builders import PromptBuilder
from haystack.components.retrievers.in_memory import InMemoryBM25Retriever
from haystack.document_stores.in_memory import InMemoryDocumentStore

store = InMemoryDocumentStore()
store.write_documents([Document(content="Haystack is a pipeline framework.")])

pipe = Pipeline()
pipe.add_component("retriever", InMemoryBM25Retriever(document_store=store))
pipe.add_component("prompt", PromptBuilder(template="{{ documents }}\n{{ question }}"))
pipe.connect("retriever.documents", "prompt.documents")

result = pipe.run(
    {"retriever": {"query": "Haystack"}, "prompt": {"question": "What is it?"}},
    include_outputs_from={"retriever"},
)
print(result["prompt"]["prompt"])
print(len(result["retriever"]["documents"]))

go deeper

for a junior

Memorize the two shapes: input is {component_name: {param: value}}, output is {component_name: {socket: value}}. Say that you only supply inputs nothing upstream provides.

for a middle

Explain why the result holds only unconsumed outputs and how include_outputs_from widens it, and why the same query often appears under two component keys.

for a senior

Show the service shape: build and validate the pipeline once at startup, keep per-request data in the input dict, and pull only the intermediates you need so responses stay small.

for a principal

Frame the input and output dicts as your pipeline's public contract — renaming a component is a breaking API change for every caller, which argues for stable component names and a thin adapter over raw run().

## The input dict `Pipeline.run()` takes one required argument: a mapping from **component name** to that component's `run()` keyword arguments. ``` pipe.run({"retriever": {"query": "what is RAG?"}, "prompt": {"question": "what is RAG?"}}) ``` The keys are exactly the names you passed to `add_component`, which is the first reason those names matter: they are the public API of your pipeline. A typo'd key is not silently ignored — the pipeline tells you it has no such component. ## Only unconnected inputs go in the dict An input socket is filled either by a connection or by you. `prompt.documents` wired from `retriever.documents` is the pipeline's job; `prompt.question` that nothing produces is yours. Passing a value for a socket that is already connected is rejected, because it would be ambiguous which value wins. This is also why the same user query often appears twice in a real RAG input dict — once for the retriever's `query` and once for the prompt builder's `question`. That duplication looks redundant and surprises newcomers, but it is the honest consequence of two components independently needing the same external value. If a **mandatory** input is neither connected nor supplied, the component never becomes runnable, and the pipeline reports that it could not execute. Inputs with defaults are optional and simply take their default. ## The flat form A flat dict — `pipe.run({"query": q})` — is also accepted. Haystack distributes each key to every component that declares a `run()` parameter of that name. It is convenient when one value feeds several components, and it is a footgun when two components happen to share a parameter name and you only meant one of them. Prefer the nested form in code you will maintain; the flat form is a demo convenience. ## Execution order You do not schedule anything. Haystack works out the order from the graph: a component runs once all of its mandatory connected inputs have arrived. Independent branches run in whatever order the scheduler picks — treat them as unordered rather than depending on a sequence. Components not reachable with the inputs you supplied (for example, the branch a router did not choose) simply do not run. ## The return value `run()` returns a dict keyed by component name, mapping to that component's output dict. By default it contains only the outputs **no other component consumed** — the graph's leaves. In a linear RAG pipeline that means you get back just the generator's output: ``` result["llm"]["replies"] ``` The filtering is deliberate: retrieved documents, embeddings and rendered prompts can be megabytes, and returning them from every run would make the result unusable as an API response. ## Getting intermediates back When you need an intermediate — to log which documents grounded an answer, to build a citations payload, to debug a bad answer — pass `include_outputs_from` with a set of component names: ``` pipe.run(data, include_outputs_from={"retriever"}) ``` Those components' outputs appear in the result alongside the leaves. Ask for what you need rather than everything: it is the difference between a compact response and shipping the whole retrieved corpus to your caller. Note this is an *output* control only — it does not change execution. ## Practical shape of a service A typical HTTP handler builds the pipeline once at startup (paying wiring validation and any model warm-up there), then per request calls `run()` with a small nested dict and reads one or two keys out of the result. Because the same pipeline object is reused across requests, keep per-request data in the input dict, never on the component instances. ## Common mistakes - Keying the input dict by socket or connection string instead of component name. - Supplying a value for a socket that is already fed by a connection. - Expecting every component's output in the result and concluding the pipeline "lost" the documents, when they were simply consumed downstream. - Assuming parallel branches execute in declaration order.

  • Why does the retrieved-documents list usually not appear in the result dict?
    Because the prompt builder consumed it. `run()` returns only outputs that no downstream component took, so anything mid-graph is dropped to keep the result small. If you need those documents — for citations or logging — name the retriever in `include_outputs_from` and its output is returned alongside the terminal component's.
  • What happens if you pass a value for an input socket that is already connected?
    The pipeline rejects it. A socket has one source of truth: either an upstream connection or your input dict, never both, because there would be no defined rule for which value wins. The fix is to remove it from the dict, or to remove the connection if you genuinely want to inject that value per run.
  • When is the flat input form a bad idea?
    When two components declare a `run()` parameter with the same name. The flat dict is distributed to every component that accepts a key of that name, so a value meant for one silently lands in both. In maintained code use the nested per-component form, which says exactly where each value goes.

saying these in an interview costs you the question

  • Keys the input dict by socket or connection string
  • Expects every component's output in the result
  • Supplies values for sockets already fed by a connection
  • Thinks components run in the order they were added
  • Believes include_outputs_from changes which components execute

context

open as a page

In Haystack, what does Pipeline.connect() validate, and when does it fail?

level: middleimportance: must knowfreq 75%

basics

~20 s

Pipeline.connect() checks at wiring time that both named sockets exist and that the sender's declared output type is compatible with the receiver's input type. A mismatch raises immediately, so a broken graph fails before any model is called.

open as a page

How do you serialize a Haystack Pipeline to YAML, and what fails to round-trip?

level: middleimportance: should knowfreq 38%

basics

~20 s

Pipeline.dumps() emits YAML describing every component's import path, init parameters and the connections; Pipeline.loads() rebuilds it. Round-trips break when a component's module is not importable in the loading process, when init parameters are not serializable, or when secrets were passed as raw literals.

open as a page

How does a Haystack pipeline branch on a runtime value and rejoin the branches?

level: seniorimportance: should knowfreq 52%

basics

~20 s

ConditionalRouter evaluates Jinja conditions over its inputs and emits on only the matching route's output socket, so the untaken branch never runs. To merge branches back, route them through a joiner with a variadic input, because a plain input socket accepts one connection.

open as a page

What makes a looping Haystack pipeline terminate, and what caps a runaway loop?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A loop terminates because a router inside it has an exit route that eventually fires. As a backstop, Pipeline(max_runs_per_component=...) caps how many times any single component may run in one call — default 100 — and exceeding it aborts the run with an error instead of spinning forever.

open as a page

When would you split one Haystack pipeline into several rather than adding more branches?

level: principalimportance: should knowfreq 30%

basics

~20 s

Split when parts of the graph have different lifecycles, owners or failure budgets — indexing versus query being the standard case. They share a document store rather than an edge. Keep one pipeline when the whole flow is a single request's dataflow.

open as a page