In Haystack, how does wiring InMemoryBM25Retriever differ from InMemoryEmbeddingRetriever?
answer
- two retrievers, two different input sockets
- one consumes text, one consumes a vector
- something must run before the dense leg
- same model on both sides of the corpus
- query versus query_embedding
basics
~20 sInMemoryBM25Retriever takes the raw query string on its query input and scores lexically over document text. InMemoryEmbeddingRetriever takes a query_embedding, so a text embedder must run first in the pipeline and feed its vector into that socket.
solid answer
~50 sBoth retrievers are bound to the same `InMemoryDocumentStore` at construction and both emit a `documents` list, but their inputs differ, and that difference is the whole wiring story. `InMemoryBM25Retriever` has a `query` input of type `str`: you pass the user's text straight into `pipeline.run({"bm25": {"query": q}})` and it scores lexically over the stored document text. `InMemoryEmbeddingRetriever` has a `query_embedding` input of type `list[float]`, so a query-side embedder such as `SentenceTransformersTextEmbedder` must sit in front of it and you connect `text_embedder.embedding` to `retriever.query_embedding`. That embedder must use the same model as the document embedder used at indexing time, otherwise the vectors live in different spaces and results are noise. Both accept `top_k` and `filters`, and swapping to Elasticsearch or Qdrant means swapping the store plus the matching retriever class from its integration package, not rewriting the pipeline shape.
code
python · 22 linesfrom haystack import Pipeline
from haystack.document_stores.in_memory import InMemoryDocumentStore
from haystack.components.embedders import SentenceTransformersTextEmbedder
from haystack.components.retrievers.in_memory import (
InMemoryBM25Retriever,
InMemoryEmbeddingRetriever,
)
store = InMemoryDocumentStore()
lexical = Pipeline()
lexical.add_component("bm25", InMemoryBM25Retriever(document_store=store, top_k=5))
lexical.run({"bm25": {"query": "how do rankers work"}})
dense = Pipeline()
dense.add_component(
"text_embedder",
SentenceTransformersTextEmbedder(model="sentence-transformers/all-MiniLM-L6-v2"),
)
dense.add_component("dense", InMemoryEmbeddingRetriever(document_store=store, top_k=5))
dense.connect("text_embedder.embedding", "dense.query_embedding")
dense.run({"text_embedder": {"text": "how do rankers work"}})go deeper
Know that the BM25 retriever takes a query string and the embedding retriever takes a vector, and that the vector comes from a text embedder wired in front of it. Say which socket each one uses.
Explain the socket contract in both directions: text embedder for the query, document embedder at indexing, both retrievers emitting a documents list. Mention that the two embedding models must match.
Show that you know the in-memory store scores over the whole corpus and is a development tool, and describe the exact swap to a backed store: new store, new retriever class, unchanged wiring. Mention a smoke test that catches an embedding-model mismatch.
Frame the choice as a substitutability property: because both legs share a socket contract, retrieval backends become a configuration decision rather than a rewrite. Own the policy that embedding model identity is pinned in one place and versioned with the index.
## The two retrieval legs Haystack keeps retrieval explicit: a retriever is an ordinary component with typed input and output sockets, bound to one document store. There is no single "retriever" that switches modes with a flag — you pick the class that matches the kind of query signal you have. `InMemoryBM25Retriever(document_store=store, top_k=10, filters=None, scale_score=False)` is the lexical leg. Its run signature takes `query: str`. Internally the `InMemoryDocumentStore` holds the documents in Python memory and computes BM25 term scores across them; the algorithm and its parameters are configurable on the store (`bm25_algorithm`, `bm25_parameters`, `bm25_tokenization_regex`). There is no index structure to warm up and no model to load — but there is also no approximate index, so scoring is a pass over the corpus and the in-memory store is a development and small-corpus tool, not a production search backend. `InMemoryEmbeddingRetriever(document_store=store, top_k=10, filters=None, return_embedding=False)` is the dense leg. Its input socket is `query_embedding: list[float]`. It compares that vector against the `embedding` field already stored on each `Document`, using the similarity function configured on the store (`embedding_similarity_function`, e.g. dot product or cosine). ## Why the dense leg needs an extra component Because the embedding retriever consumes a vector rather than text, a query pipeline that uses it always has at least two components: ``` text_embedder.embedding -> retriever.query_embedding ``` The usual embedder is `SentenceTransformersTextEmbedder`, which takes `text` and outputs `embedding`. Note the deliberate split in Haystack between the *text* embedder (query side, one string in, one vector out) and the *document* embedder (indexing side, a list of `Document` objects in, the same documents with `embedding` populated out). Using the document embedder on a query, or vice versa, produces a socket type mismatch that `Pipeline.connect` rejects at wiring time — one of the reasons Haystack's explicit graph is easier to debug than a chain that fails at run time. The model identity matters more than the class. If the indexing pipeline embedded documents with one sentence-transformers model and the query pipeline embeds with another, nothing errors: the dimensions may even match, and you get plausible-looking but meaningless neighbours. Pin the model name in configuration and use the same value on both sides. ## What they share - **The store.** Both retrievers hold a reference to the same store object; documents are not duplicated. A hybrid pipeline can point a BM25 retriever and an embedding retriever at one `InMemoryDocumentStore` and get two views of the same corpus. - **`top_k`.** The number of documents returned, settable at construction and overridable per run. - **`filters`.** A metadata filter dictionary narrowing the candidate set before scoring. - **The output socket.** Both emit `documents: list[Document]`, which is why either can feed a joiner, a ranker or a prompt builder without adapters. - **`scale_score`.** Both can map raw scores into a 0–1 range, which matters when you want comparable numbers across legs. ## Moving off in-memory The in-memory pair exists so you can build and test a pipeline with no infrastructure. Production stores ship their own retriever classes in the corresponding `haystack-integrations` package — for example an Elasticsearch BM25 retriever or a Qdrant embedding retriever — with the same socket contract. Migrating means changing two constructor lines (store and retriever class); the `connect` calls and the downstream ranker/generator stay as they are. That substitutability is the practical argument for the explicit component graph. ## Common wiring mistakes 1. Connecting the embedder's output to the BM25 retriever, or passing `query` to the embedding retriever — both are socket type errors, caught at `connect` or run time rather than silently. 2. Forgetting that in `pipeline.run()` you address the *entry* component: for the dense leg you pass `{"text_embedder": {"text": q}}`, not `{"retriever": {"query": q}}`. 3. Feeding the same query text to both legs in a hybrid pipeline but forgetting one of the two input dict entries, so one leg silently never runs. 4. Assuming the BM25 retriever benefits from having embeddings present. It ignores the `embedding` field entirely.
- What actually happens if the query embedder uses a different model from the one used at indexing?Usually nothing visible fails. If the dimensions happen to match, the retriever compares vectors from two unrelated spaces and returns near-arbitrary neighbours with plausible scores; if the dimensions differ you get a shape error at scoring time. Pin the model name in one config value and reuse it in both the indexing and query pipelines, and add a smoke test asserting a known query retrieves a known document.
- Both retrievers point at the same InMemoryDocumentStore — does that double the memory?No. The retrievers are thin components holding a reference to one store object; the documents and their embeddings exist once. What does cost memory is the embeddings themselves, stored on each Document, plus the fact that the in-memory store keeps the whole corpus in the Python process. That is the real scaling limit, not the number of retrievers attached.
- How much of the pipeline changes when you move from InMemoryDocumentStore to a real backend?The store constructor and the retriever class, both from that backend's integration package. Socket names and types are the same, so the connect calls, the joiner, the ranker and the generator are untouched. The indexing pipeline changes similarly at the writer. Keeping those two lines behind a factory function makes the swap a one-line configuration change.
saying these in an interview costs you the question
- Says the BM25 retriever needs a text embedder in front of it
- Passes a raw query string to InMemoryEmbeddingRetriever
- Uses different embedding models for indexing and querying
- Thinks the in-memory BM25 retriever builds an ANN index
- Believes one retriever class handles both lexical and dense modes