skip to content

In Haystack, how does DocumentJoiner's reciprocal_rank_fusion mode build hybrid retrieval?

level: middleimportance: must knowfreq 66%

answer

  1. one component does the merging
  2. its documents input is not a single connection
  3. ranks travel, raw scores do not
  4. order of connect calls matters for weights
  5. two entry points at run time

basics

~20 s

You wire a BM25 retriever and an embedding retriever into the same DocumentJoiner, whose documents input is variadic and accepts several connections. With join_mode set to reciprocal_rank_fusion the joiner deduplicates by document id and replaces each score with one derived from the document's rank in every incoming list.

solid answer

~50 s

`DocumentJoiner` is the component that makes hybrid retrieval a wiring exercise rather than custom code. Its `documents` input is variadic, so you call `Pipeline.connect` once per retriever — `bm25.documents -> joiner.documents` and `dense.documents -> joiner.documents` — and the joiner receives one list per upstream branch. With `join_mode="reciprocal_rank_fusion"` it deduplicates documents by id and assigns each a fused score built from its position in each list, which is why the incompatible BM25 and cosine scales never have to be normalised. `weights` lets one leg count for more than the other, aligned to the order the connections were made, and the joiner's own `top_k` truncates the fused list before it reaches a ranker or prompt builder. The catch to remember at run time: both legs are entry points, so `pipeline.run()` must supply inputs for both — the query string for BM25 and the text for the query embedder — or one branch silently contributes nothing.

code

python · 25 lines
python
from haystack import Pipeline
from haystack.components.joiners import DocumentJoiner
from haystack.components.embedders import SentenceTransformersTextEmbedder
from haystack.components.retrievers.in_memory import (
    InMemoryBM25Retriever,
    InMemoryEmbeddingRetriever,
)
from haystack.document_stores.in_memory import InMemoryDocumentStore

store = InMemoryDocumentStore()
pipe = Pipeline()
pipe.add_component("bm25", InMemoryBM25Retriever(document_store=store, top_k=20))
pipe.add_component("text_embedder", SentenceTransformersTextEmbedder())
pipe.add_component("dense", InMemoryEmbeddingRetriever(document_store=store, top_k=20))
pipe.add_component(
    "joiner",
    DocumentJoiner(join_mode="reciprocal_rank_fusion", weights=[0.5, 0.5], top_k=10),
)

pipe.connect("text_embedder.embedding", "dense.query_embedding")
pipe.connect("bm25.documents", "joiner.documents")
pipe.connect("dense.documents", "joiner.documents")

query = "error code E4021 on startup"
pipe.run({"bm25": {"query": query}, "text_embedder": {"text": query}})

go deeper

for a junior

Know that hybrid retrieval in Haystack means two retrievers feeding one DocumentJoiner, and that join_mode chooses how the lists are merged. Be able to name reciprocal_rank_fusion as the usual mode.

for a middle

Explain the variadic documents input, why ranks are fused rather than raw scores, and that the fused score replaces the original one. Mention that run() must feed both branches.

for a senior

Show the operational picture: joiner widens recall, a ranker restores precision, weights are tuned on a labelled set, and dedup depends on stable document ids. Name the silent failure of a half-fed hybrid pipeline.

for a principal

Own the evaluation story — hybrid is only justified if a query set shows recall gains the reranker can convert into answer quality, and the extra leg's latency and cost are budgeted. Decide when a single well-tuned leg is the better system.

## What the joiner is for Hybrid retrieval means running two retrievers over the same corpus and merging their candidate lists. In Haystack that merge is an ordinary component, `DocumentJoiner`, so the whole arrangement stays visible in the pipeline graph and in the serialised YAML. ```python from haystack.components.joiners import DocumentJoiner pipe.add_component("joiner", DocumentJoiner(join_mode="reciprocal_rank_fusion", top_k=10)) pipe.connect("bm25.documents", "joiner.documents") pipe.connect("dense.documents", "joiner.documents") ``` ## The variadic input The joiner's `documents` input is *variadic*: unlike a normal socket, which accepts exactly one connection, it accepts many and delivers them to `run()` as a list of lists. This is the single most-missed mechanical detail. You do not create numbered sockets, and you do not need an intermediate component — you simply connect each producer to the same socket name. The order of those `connect` calls is not cosmetic: it defines the order of the incoming lists, and therefore which entry of `weights` applies to which leg. ## The join modes `join_mode` selects the merge policy: - **`concatenate`** — union of the lists, deduplicated by document id, keeping the highest score seen for a duplicate. Fine when both legs produce comparable scores, which in practice they rarely do. - **`merge`** — a weighted average of the scores of documents present in several lists. - **`reciprocal_rank_fusion`** — scores are discarded and each document is scored from its *rank* in each list. - **`distribution_based_rank_fusion`** — normalises each list using the distribution of its own scores before combining. Rank fusion is the default choice for BM25-plus-dense because BM25 scores are unbounded and corpus-dependent while cosine similarities sit in a narrow band; there is no honest constant that makes them comparable. Working from ranks sidesteps the problem entirely. The practical consequence, and the thing to say out loud in an interview, is that **after RRF the `score` on each Document is a fusion score, not a similarity** — do not threshold it with a cutoff you tuned against cosine values, and do not show it to users as a confidence. ## Deduplication Merging identical documents depends on `Document.id`. Haystack derives that id by hashing the document's content and metadata unless you set it explicitly, so the same chunk written once and retrieved by both legs collapses into one entry with a boosted fused score — which is exactly the behaviour you want, since agreement between two independent retrievers is signal. If your indexing pipeline assigns ids non-deterministically, or writes the same text twice with differing metadata, dedup silently stops working and duplicates survive the join. ## `weights`, `top_k` and `sort_by_score` `weights` takes one number per incoming list and biases the fusion — useful when your corpus is jargon-heavy and the lexical leg deserves more say, or the opposite for conversational queries. Tune it against a labelled query set, not by intuition. `top_k` on the joiner caps the fused output; setting it well below the sum of the legs' `top_k` is the point of the exercise, since the fused list is the candidate pool for a downstream reranker. `sort_by_score` controls whether the output list is ordered by the resulting score; leaving it on is normal, but it matters if a later component depends on incoming order. ## Running the pipeline A hybrid pipeline has two entry points. The BM25 retriever needs `query`, and the dense leg's text embedder needs `text`: ```python pipe.run({"bm25": {"query": q}, "text_embedder": {"text": q}}) ``` Omit one and Haystack will simply not execute that branch, leaving you with a "hybrid" pipeline that is quietly doing single-leg retrieval and mediocre results with no error. A cheap guard is an integration test asserting the joiner received two lists — or, more practically, asserting a known lexical-only match (an exact product code, say) survives to the top of the fused list. ## Where the joiner sits The usual production shape is: two retrievers → `DocumentJoiner` → a similarity ranker → prompt builder → generator. The joiner widens recall cheaply; the ranker restores precision on the widened pool. Fusing and then feeding the raw fused list straight to the generator wastes most of the benefit, because RRF is good at surfacing candidates and mediocre at final ordering.

  • After a reciprocal_rank_fusion join, why is it wrong to apply the score threshold you tuned on cosine similarity?
    Because the joiner overwrites each document's score with a rank-derived fusion value that has no relationship to cosine magnitude — it depends on list length and position, not on semantic closeness. A cutoff tuned on cosine will either drop everything or nothing. If you want a meaningful cutoff, put a similarity ranker after the joiner and threshold on its score instead.
  • How do you decide top_k on each leg versus top_k on the joiner?
    Each leg should fetch deep enough that a document only one retriever likes still appears — typically well beyond what you intend to keep. The joiner's top_k then caps the fused candidate pool handed to the reranker. Widening the legs costs retrieval time and reranking time, so tune both against a labelled query set and a latency budget rather than guessing.
  • Two retrievers return the same passage but dedup does not collapse it. What would you check?
    Document ids. Haystack derives the id from content plus metadata unless you set it, so the same text indexed twice with different metadata produces two ids and survives the join as a duplicate. Check the indexing pipeline for non-deterministic ids or duplicated writes, and consider a DuplicatePolicy on the writer so the corpus itself holds one copy.
  • When would you pick distribution_based_rank_fusion over reciprocal_rank_fusion?
    When you want the magnitude of the scores to carry information rather than only their order. Distribution-based fusion normalises each list against its own score distribution, so a document that is far ahead of the pack in one leg keeps that advantage, whereas rank fusion flattens it to first place. It is worth trying when one leg is much better calibrated than the other.

saying these in an interview costs you the question

  • Thinks each retriever needs its own numbered joiner input
  • Treats the post-fusion score as a similarity value
  • Normalises BM25 and cosine scores by hand before joining
  • Forgets to pass inputs for both branches in run()
  • Sends the fused list straight to the generator with no reranking

context