skip to content

When would you split one Haystack pipeline into several rather than adding more branches?

level: principalimportance: should knowfreq 30%

answer

  1. Ask what differs, not what fits
  2. Ingest and serve are different jobs
  3. The store is the seam
  4. Boundaries lose the type check
  5. Look at the drawn diagram first

basics

~20 s

Split when parts of the graph have different lifecycles, owners or failure budgets — indexing versus query being the standard case. They share a document store rather than an edge. Keep one pipeline when the whole flow is a single request's dataflow.

solid answer

~50 s

The first split is almost always indexing versus query: indexing runs on ingestion, is slow and batchy, and is owned by whoever manages content; the query pipeline runs per request under a latency budget. They share the **document store**, not a connection, and that shared store is the real seam. Beyond that, split when a segment has its own deploy cadence, its own on-call owner, or its own scaling profile — a reranking service you want to version independently, an offline evaluation flow. Keep things in one pipeline while the whole thing is one request's dataflow, because that is what buys you wiring-time type checks, one `draw()` diagram, one YAML artifact and one place to trace. The cost of splitting is exactly those: across a process boundary the type check disappears, and you now own the contract by hand. My heuristic: routers that encode *product* decisions belong in the graph; routers that exist because two unrelated jobs were stapled together are a split waiting to happen.

go deeper

for a junior

Know the standard arrangement: a separate indexing pipeline that writes documents and a query pipeline that retrieves from the same store, rather than one graph doing both.

for a middle

Explain that the two pipelines share a document store, not a connection, and that each pipeline is its own unit of validation, serialization and drawing.

for a senior

Argue from lifecycle, ownership and scaling: what runs per request versus per ingestion, what needs to deploy independently, and what you give up at a boundary — the type-checked edge and the single trace.

for a principal

Own the seam design: the metadata contract between write and read sides, correlation across two runs, CI that loads every serialized pipeline, and a stated rule for when in-graph branching becomes a split.

## Why this is a judgment question Haystack does not stop you putting everything in one `Pipeline`. Nothing prevents an indexing branch, a query branch, an evaluation branch and three routers in a single graph. The framework's guardrails — type-checked edges, a run cap, serialization — all still work. So the question is not "can you" but "what do you gain and lose at the boundary", which is precisely why it is asked of people expected to own an architecture. ## What one pipeline buys 1. **Wiring-time validation across the whole flow.** Every edge is type-checked when the pipeline is built. Split the flow across two pipelines and the handoff between them is ordinary Python — checked by nothing. 2. **One diagram.** `pipe.draw()` renders the graph; a reviewer sees the whole dataflow in one image. Two pipelines mean two images and a mental join. 3. **One configuration artifact.** One YAML file to diff, deploy and version. 4. **One trace.** A single `run()` is one unit of observability, and `include_outputs_from` reaches any component in it. Those are real and they argue for keeping related work together longer than instinct suggests. ## What forces a split **Different lifecycles.** Indexing runs when documents arrive — minutes, batchy, retry-tolerant. Query runs per request under a latency SLO. Nothing about running them in the same graph helps, and the run cap, timeouts and resource sizing you would want differ. This is the split everyone makes, and the important part of the answer is *why they still compose*: both reference the same document store, so the write side and the read side meet in data, not in an edge. **Different owners.** If the content team tunes the splitter and cleaner while the product team tunes retrieval and prompting, two artifacts mean two review paths instead of one file everyone fights over. **Different scaling profiles.** A GPU-backed embedding or reranking stage that you want to scale and deploy independently is a service boundary, not a component boundary. Keeping it as an in-process component ties its capacity to your web tier. **Different failure budgets.** An enrichment step whose failure should degrade the answer, not fail the request, is easier to reason about outside the main graph, where you control the exception boundary explicitly. **Testability under cycles.** A pipeline containing a correction loop is harder to test end-to-end than the same work as a linear pipeline plus explicit retry logic in the caller. If you cannot write a fast unit test for a segment, that is evidence for pulling it out. ## What does not justify a split - **Size alone.** A twelve-component linear pipeline is fine; it draws as a straight line and reads well. - **A couple of routers.** Product branching — grounded answer versus fallback, chat versus search — is genuine per-request dataflow and belongs in the graph where it is visible and serialized. - **"It feels complex."** Run `draw()` first. If the diagram is legible, the graph is fine; if it is a hairball of crossing edges, that is the signal, and it is a signal you can put in a design doc rather than an opinion you have to defend. ## Composition options short of a split Before going to two pipelines, consider whether a *sub-flow* can be packaged as a single component: a component whose `run()` internally does several steps presents one node to the graph. That collapses visual complexity while keeping one process, one trace and one artifact. The cost is that the inner steps are no longer individually wired, typed or drawn — you have traded inspectability for tidiness, so do it for stable, well-tested internals, not for the part you are still tuning. ## Making a split safe When you do split, put the effort into the seam: - **Name the shared state explicitly.** Usually a document store with a defined schema and metadata contract. Write down which fields the query side relies on; the indexing side breaking them is the top cause of silent retrieval regressions. - **Type the handoff yourself.** A small typed object between pipelines replaces the check you just gave up. - **Keep both loadable in CI.** Load every serialized pipeline in a test so wiring and imports stay valid even though no single graph spans them. - **Correlate traces.** One request now touches two runs; without a shared correlation id, debugging costs double. ## The answer to give State the default (indexing and query are separate, joined by the store), give two or three concrete forces that justify further splits — lifecycle, ownership, scaling, failure budget — name what you lose at the boundary (type-checked edges, one diagram, one trace, one artifact), and describe how you would make the seam safe. Avoid absolutes in either direction; an interviewer asking this wants to hear you weigh, not rule.

  • If indexing and query are separate pipelines, what actually connects them?
    The document store, plus a metadata contract. The indexing pipeline writes documents with a known set of fields; the query pipeline filters and retrieves on those fields. Nothing type-checks that agreement, so write it down and test it — an indexing change that renames or drops a metadata field is the most common cause of a retrieval regression that looks like a model problem.
  • What do you lose the moment a flow spans two pipelines?
    The wiring-time type check across the seam, the single diagram, the single serialized artifact, and the single trace. The handoff becomes ordinary code that nothing validates. You compensate deliberately: a typed object at the boundary, a correlation id shared by both runs, and a CI test that loads both pipelines so neither drifts out of validity unnoticed.
  • When is packaging a sub-flow as one component better than splitting the pipeline?
    When the internals are stable and well tested and the only problem is visual complexity. A component whose run() performs several steps shows as one node, keeping one process, one trace and one artifact. The price is that the inner steps are no longer separately wired, typed or drawn — so do not hide the part of the flow you are still actively tuning.

saying these in an interview costs you the question

  • Says one pipeline per application, always
  • Splits by component count rather than lifecycle
  • Forgets the document store is the shared seam
  • Assumes cross-pipeline handoffs are still type-checked
  • Treats product branching as a reason to split

context