skip to content

Retrievers & Rankers

You will learn Haystack's retrieval stage: BM25 and embedding retrievers bound to a store, hybrid pipelines joined by reciprocal rank fusion, and rankers that reorder or diversify the result. Interviewers ask how you improved recall without wrecking precision, and this is the concrete answer.

part ofAI agent & RAG frameworksoverview, primer and where to startread it →
on this pageshow

questions

6

In Haystack, how does wiring InMemoryBM25Retriever differ from InMemoryEmbeddingRetriever?

level: juniorimportance: must knowfreq 74%

answer

  1. two retrievers, two different input sockets
  2. one consumes text, one consumes a vector
  3. something must run before the dense leg
  4. same model on both sides of the corpus
  5. query versus query_embedding

basics

~20 s

InMemoryBM25Retriever takes the raw query string on its query input and scores lexically over document text. InMemoryEmbeddingRetriever takes a query_embedding, so a text embedder must run first in the pipeline and feed its vector into that socket.

solid answer

~50 s

Both retrievers are bound to the same `InMemoryDocumentStore` at construction and both emit a `documents` list, but their inputs differ, and that difference is the whole wiring story. `InMemoryBM25Retriever` has a `query` input of type `str`: you pass the user's text straight into `pipeline.run({"bm25": {"query": q}})` and it scores lexically over the stored document text. `InMemoryEmbeddingRetriever` has a `query_embedding` input of type `list[float]`, so a query-side embedder such as `SentenceTransformersTextEmbedder` must sit in front of it and you connect `text_embedder.embedding` to `retriever.query_embedding`. That embedder must use the same model as the document embedder used at indexing time, otherwise the vectors live in different spaces and results are noise. Both accept `top_k` and `filters`, and swapping to Elasticsearch or Qdrant means swapping the store plus the matching retriever class from its integration package, not rewriting the pipeline shape.

code

python · 22 lines
python
from haystack import Pipeline
from haystack.document_stores.in_memory import InMemoryDocumentStore
from haystack.components.embedders import SentenceTransformersTextEmbedder
from haystack.components.retrievers.in_memory import (
    InMemoryBM25Retriever,
    InMemoryEmbeddingRetriever,
)

store = InMemoryDocumentStore()

lexical = Pipeline()
lexical.add_component("bm25", InMemoryBM25Retriever(document_store=store, top_k=5))
lexical.run({"bm25": {"query": "how do rankers work"}})

dense = Pipeline()
dense.add_component(
    "text_embedder",
    SentenceTransformersTextEmbedder(model="sentence-transformers/all-MiniLM-L6-v2"),
)
dense.add_component("dense", InMemoryEmbeddingRetriever(document_store=store, top_k=5))
dense.connect("text_embedder.embedding", "dense.query_embedding")
dense.run({"text_embedder": {"text": "how do rankers work"}})

go deeper

for a junior

Know that the BM25 retriever takes a query string and the embedding retriever takes a vector, and that the vector comes from a text embedder wired in front of it. Say which socket each one uses.

for a middle

Explain the socket contract in both directions: text embedder for the query, document embedder at indexing, both retrievers emitting a documents list. Mention that the two embedding models must match.

for a senior

Show that you know the in-memory store scores over the whole corpus and is a development tool, and describe the exact swap to a backed store: new store, new retriever class, unchanged wiring. Mention a smoke test that catches an embedding-model mismatch.

for a principal

Frame the choice as a substitutability property: because both legs share a socket contract, retrieval backends become a configuration decision rather than a rewrite. Own the policy that embedding model identity is pinned in one place and versioned with the index.

## The two retrieval legs Haystack keeps retrieval explicit: a retriever is an ordinary component with typed input and output sockets, bound to one document store. There is no single "retriever" that switches modes with a flag — you pick the class that matches the kind of query signal you have. `InMemoryBM25Retriever(document_store=store, top_k=10, filters=None, scale_score=False)` is the lexical leg. Its run signature takes `query: str`. Internally the `InMemoryDocumentStore` holds the documents in Python memory and computes BM25 term scores across them; the algorithm and its parameters are configurable on the store (`bm25_algorithm`, `bm25_parameters`, `bm25_tokenization_regex`). There is no index structure to warm up and no model to load — but there is also no approximate index, so scoring is a pass over the corpus and the in-memory store is a development and small-corpus tool, not a production search backend. `InMemoryEmbeddingRetriever(document_store=store, top_k=10, filters=None, return_embedding=False)` is the dense leg. Its input socket is `query_embedding: list[float]`. It compares that vector against the `embedding` field already stored on each `Document`, using the similarity function configured on the store (`embedding_similarity_function`, e.g. dot product or cosine). ## Why the dense leg needs an extra component Because the embedding retriever consumes a vector rather than text, a query pipeline that uses it always has at least two components: ``` text_embedder.embedding -> retriever.query_embedding ``` The usual embedder is `SentenceTransformersTextEmbedder`, which takes `text` and outputs `embedding`. Note the deliberate split in Haystack between the *text* embedder (query side, one string in, one vector out) and the *document* embedder (indexing side, a list of `Document` objects in, the same documents with `embedding` populated out). Using the document embedder on a query, or vice versa, produces a socket type mismatch that `Pipeline.connect` rejects at wiring time — one of the reasons Haystack's explicit graph is easier to debug than a chain that fails at run time. The model identity matters more than the class. If the indexing pipeline embedded documents with one sentence-transformers model and the query pipeline embeds with another, nothing errors: the dimensions may even match, and you get plausible-looking but meaningless neighbours. Pin the model name in configuration and use the same value on both sides. ## What they share - **The store.** Both retrievers hold a reference to the same store object; documents are not duplicated. A hybrid pipeline can point a BM25 retriever and an embedding retriever at one `InMemoryDocumentStore` and get two views of the same corpus. - **`top_k`.** The number of documents returned, settable at construction and overridable per run. - **`filters`.** A metadata filter dictionary narrowing the candidate set before scoring. - **The output socket.** Both emit `documents: list[Document]`, which is why either can feed a joiner, a ranker or a prompt builder without adapters. - **`scale_score`.** Both can map raw scores into a 0–1 range, which matters when you want comparable numbers across legs. ## Moving off in-memory The in-memory pair exists so you can build and test a pipeline with no infrastructure. Production stores ship their own retriever classes in the corresponding `haystack-integrations` package — for example an Elasticsearch BM25 retriever or a Qdrant embedding retriever — with the same socket contract. Migrating means changing two constructor lines (store and retriever class); the `connect` calls and the downstream ranker/generator stay as they are. That substitutability is the practical argument for the explicit component graph. ## Common wiring mistakes 1. Connecting the embedder's output to the BM25 retriever, or passing `query` to the embedding retriever — both are socket type errors, caught at `connect` or run time rather than silently. 2. Forgetting that in `pipeline.run()` you address the *entry* component: for the dense leg you pass `{"text_embedder": {"text": q}}`, not `{"retriever": {"query": q}}`. 3. Feeding the same query text to both legs in a hybrid pipeline but forgetting one of the two input dict entries, so one leg silently never runs. 4. Assuming the BM25 retriever benefits from having embeddings present. It ignores the `embedding` field entirely.

  • What actually happens if the query embedder uses a different model from the one used at indexing?
    Usually nothing visible fails. If the dimensions happen to match, the retriever compares vectors from two unrelated spaces and returns near-arbitrary neighbours with plausible scores; if the dimensions differ you get a shape error at scoring time. Pin the model name in one config value and reuse it in both the indexing and query pipelines, and add a smoke test asserting a known query retrieves a known document.
  • Both retrievers point at the same InMemoryDocumentStore — does that double the memory?
    No. The retrievers are thin components holding a reference to one store object; the documents and their embeddings exist once. What does cost memory is the embeddings themselves, stored on each Document, plus the fact that the in-memory store keeps the whole corpus in the Python process. That is the real scaling limit, not the number of retrievers attached.
  • How much of the pipeline changes when you move from InMemoryDocumentStore to a real backend?
    The store constructor and the retriever class, both from that backend's integration package. Socket names and types are the same, so the connect calls, the joiner, the ranker and the generator are untouched. The indexing pipeline changes similarly at the writer. Keeping those two lines behind a factory function makes the swap a one-line configuration change.

saying these in an interview costs you the question

  • Says the BM25 retriever needs a text embedder in front of it
  • Passes a raw query string to InMemoryEmbeddingRetriever
  • Uses different embedding models for indexing and querying
  • Thinks the in-memory BM25 retriever builds an ANN index
  • Believes one retriever class handles both lexical and dense modes

context

open as a page

In Haystack, how does DocumentJoiner's reciprocal_rank_fusion mode build hybrid retrieval?

level: middleimportance: must knowfreq 66%

basics

~20 s

You wire a BM25 retriever and an embedding retriever into the same DocumentJoiner, whose documents input is variadic and accepts several connections. With join_mode set to reciprocal_rank_fusion the joiner deduplicates by document id and replaces each score with one derived from the document's rank in every incoming list.

open as a page

In Haystack, why must TransformersSimilarityRanker be warmed up before it ranks?

level: middleimportance: should knowfreq 52%

basics

~20 s

Its constructor only records the model name and device so components stay cheap to build and serialise. The model is downloaded and loaded onto the device in warm_up(), which a pipeline calls for you before the first run; calling run() standalone without it raises an error.

open as a page

In Haystack retrievers, how do init-time top_k and filters differ from run-time ones?

level: middleimportance: should knowfreq 44%

basics

~20 s

Constructor values are the component's defaults and are captured in its serialised form. Values passed in Pipeline.run's per-component input dict apply to that run only. top_k always overrides, while filters are combined or replaced according to the retriever's filter_policy setting.

open as a page

In Haystack's SentenceTransformersDiversityRanker, how do its two strategies differ?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

With strategy="greedy_diversity_order" it repeatedly picks the candidate least similar to what it has already selected, so the output spreads out. With strategy="maximum_margin_relevance" each pick balances query relevance against similarity to the selection so far, tuned by lambda_threshold.

open as a page

What does Haystack's LostInTheMiddleRanker do to an already-ranked document list?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

It reorders documents so the most relevant ones sit at the start and end of the list and the least relevant in the middle. It loads no model and computes no scores, trusting the incoming order to already be relevance-sorted, and word_count_threshold can cap how much text it passes on.

open as a page