skip to content

Retrievers & RAG

You will learn the retrieval half of the framework: document loaders, text splitters, vectorstore retrievers, MultiQueryRetriever and ContextualCompressionRetriever, retrieval chain patterns, and hybrid search with reranking. Interviewers ask because most RAG quality complaints are retrieval failures, and this is where chunking, query expansion and reranking decisions actually get made.

part ofAI agent & RAG frameworksoverview, primer and where to startread it →
on this pageshow

questions

7

In LangChain, what does as_retriever(search_type=...) control on a vectorstore?

level: middleimportance: must knowfreq 66%

answer

  1. Adapter from store search to one method
  2. Three strategies, one keyword
  3. Diversity knob needs a candidate pool
  4. A threshold lets you return nothing
  5. k defaults smaller than people assume

basics

~20 s

as_retriever() wraps a vectorstore as a retriever you call with invoke(query). search_type selects the algorithm — "similarity" (default), "mmr" for diversity, or "similarity_score_threshold" — and search_kwargs passes k, fetch_k, lambda_mult, score_threshold and store-specific metadata filters.

solid answer

~40 s

`as_retriever()` adapts a vectorstore into a `VectorStoreRetriever`, whose contract is simply `invoke(query) -> list[Document]`. Two arguments do the configuring. `search_type` picks the strategy: `"similarity"` returns the nearest `k` neighbours; `"mmr"` runs maximal marginal relevance, pulling `fetch_k` candidates (default 20) and greedily choosing `k` of them balancing relevance against novelty via `lambda_mult` (0 = maximum diversity, 1 = pure relevance, default 0.5); `"similarity_score_threshold"` returns only documents above a `score_threshold` you must supply, and can legitimately return fewer than `k` documents or none at all. `search_kwargs` is a passthrough dict — `k` (default 4), the MMR parameters, and any store-specific `filter` for metadata. Everything below the retriever, including how scores are computed, belongs to the store.

code

python · 5 lines
python
retriever = vectorstore.as_retriever(
    search_type="mmr",
    search_kwargs={"k": 5, "fetch_k": 30, "lambda_mult": 0.4},
)
docs = retriever.invoke("how do refunds work for damaged goods?")

go deeper

for a junior

Know that as_retriever() gives you an object you call with invoke(query), and that k in search_kwargs sets how many documents come back.

for a middle

Explain all three search types and what each parameter does, especially why MMR needs a larger candidate pool than k.

for a senior

Argue about calibration and cost: thresholds being store- and model-specific, metadata filters beating rank tuning, and k as a permanent per-request token bill.

for a principal

Treat retrieval configuration as something measured against a labelled query set and re-validated on every embedding-model or store change, not tuned by intuition.

## What the wrapper is for A vectorstore has a rich, store-specific search API. A retriever has exactly one method that matters: give it a query string, get back a list of Documents. `as_retriever()` is the adapter between them, and its value is that everything downstream — prompt assembly, compression, fusion — only has to know the narrow interface. Swapping the underlying store then does not ripple through your application. The object it returns is a `VectorStoreRetriever`, invoked as `retriever.invoke("...")`, with an async `ainvoke` alongside it. ## search_type: similarity The default. Embed the query, ask the store for the `k` nearest vectors, return them as Documents in rank order. `k` defaults to 4 — small enough that people are often surprised by how little context reaches the model until they look. This is pure nearest-neighbour, which means it is also happily redundant: if your corpus contains the same paragraph in five documents, similarity search will return all five and burn your whole context budget on one fact. ## search_type: mmr Maximal marginal relevance exists for exactly that failure. The store first fetches `fetch_k` candidates (default 20), then the retriever selects `k` of them one at a time: at each step it picks the candidate that maximises a blend of similarity to the query and dissimilarity to what is already selected. `lambda_mult` sets the blend — 1 means ignore diversity and behave like plain similarity, 0 means maximise diversity, 0.5 is the default midpoint. The practical tuning rule is that `fetch_k` must be meaningfully larger than `k` or there is nothing to diversify from; `fetch_k=20, k=4` gives the algorithm real choice, `fetch_k=5, k=4` does not. MMR costs an extra candidate fetch and a small selection computation, and it is the right default when your corpus has heavy near-duplication — release notes, versioned docs, boilerplate contracts. ## search_type: similarity_score_threshold Here you supply `score_threshold` in `search_kwargs` and the retriever drops anything below it. The point is to be allowed to return **nothing**. A RAG system that always returns its four nearest neighbours will happily hand the model four irrelevant chunks for an off-topic question, and the model will dutifully write a confident answer grounded in them. A threshold gives your application the signal it needs to say "I don't have anything on that". The catch is calibration. Scores are store-dependent, and the number that reaches the threshold is a normalised relevance score, not the raw distance some stores return from their scored-search methods — several stores measure distance, where lower is better, and convert. So a threshold that works on one store is not portable to another, and it is not portable across embedding models either. Calibrate empirically on a labelled query set; do not copy a number from a tutorial. ## search_kwargs and metadata filters `search_kwargs` is passed through to the store's search call, so its accepted keys are ultimately the store's business. `k` is universal. `fetch_k` and `lambda_mult` apply to MMR, `score_threshold` to thresholded search. `filter` carries a metadata predicate whose *syntax* is store-specific — this is the usual portability wall when migrating between vector databases. Metadata filtering deserves emphasis because it is the cheapest large quality win available. Restricting the search to the current tenant, language, product version or document type removes whole classes of wrong-but-similar results that no amount of reranking would have fixed. It also has a failure mode: a filter narrow enough to leave fewer than `k` matching vectors returns a short list, and an over-filtered query returns nothing, which reads at the application level exactly like a retrieval failure. ## What tuning here can and cannot do These knobs choose among the vectors you already have. They cannot repair chunks that were split badly, they cannot find a document that was never ingested, and they cannot fix a query embedded with a different model than the index. When retrieval quality is bad, check ingestion and embedding-model consistency before spending a week on `lambda_mult`. One more operational note: raising `k` is the most tempting lever and the one with the most hidden cost. Each extra document is prompt tokens on every single request, so `k=20` is a permanent bill and a permanent latency increase, and it pushes relevant material toward the middle of a long context where models attend to it least. Retrieve wider only if you then compress or rerank down.

  • Why does MMR need fetch_k to be much larger than k?
    MMR selects k documents from the fetch_k candidates the store returned, choosing each one to be relevant yet unlike those already chosen. If fetch_k barely exceeds k there is nothing to choose between and the output collapses back to plain similarity. A pool several times larger than k — the default is fetch_k=20 against k=4 — is what gives the diversity term something to work with.
  • A similarity_score_threshold retriever suddenly returns nothing after you swapped embedding models. Why?
    Score distributions are a property of the embedding model and the store's metric, not of your application. A new model shifts where relevant documents land on the 0-to-1 relevance scale, so a threshold calibrated for the old one can now sit above almost everything. Thresholds must be re-calibrated against a labelled query set whenever the model or the store changes.
  • When is a metadata filter a better fix than tuning k or lambda_mult?
    Whenever the wrong results are wrong for a structural reason rather than a semantic one — another tenant's data, an obsolete product version, the wrong language. Filtering removes them before ranking even happens, which is both cheaper and more reliable than hoping a relevance score separates them. Just watch for over-filtering: a narrow predicate can leave fewer than k candidates, or none.

saying these in an interview costs you the question

  • Thinks k is the number of candidates before reranking
  • Copies a score_threshold value from a tutorial
  • Sets fetch_k equal to k and expects diversity
  • Assumes filter syntax is the same across vector stores
  • Raises k to 20 without compressing the results

context

open as a page

How does LangChain's RecursiveCharacterTextSplitter decide where to cut a document?

level: middleimportance: must knowfreq 72%

basics

~20 s

It tries a separator list in order — blank lines, then newlines, then spaces, then bare characters — recursing into any piece still over chunk_size, then merges neighbouring pieces up to chunk_size while repeating chunk_overlap units of the previous chunk.

open as a page

In LangChain v1, how do you wire a retriever into a question-answering chain?

level: middleimportance: must knowfreq 62%

basics

~20 s

Retrieve, format, prompt: call retriever.invoke(question), join the returned Documents' page_content into a context string, and render it into a prompt template beside the question. The pre-1.0 create_retrieval_chain helper packaged that shape and belongs to the legacy compatibility surface, not to v1's slim core.

open as a page

In LangChain, what does a document loader return, and when do you use lazy_load()?

level: juniorimportance: should knowfreq 48%

basics

~20 s

LangChain document loaders return Document objects, each with page_content text and a metadata dict (source, page). load() builds the full list in memory; lazy_load() yields Documents one at a time, so large corpora stream into splitting and embedding.

open as a page

When would you wrap a LangChain retriever in ContextualCompressionRetriever?

level: seniorimportance: should knowfreq 46%

basics

~20 s

When you need to retrieve wide but prompt narrow. ContextualCompressionRetriever runs a base retriever, then passes its documents plus the query through a compressor that filters, reranks or trims them — so you can raise k for recall without paying for all of it in prompt tokens.

open as a page

How does LangChain's EnsembleRetriever combine BM25 and vector search results?

level: seniorimportance: should knowfreq 44%

basics

~20 s

By reciprocal rank fusion, not by score arithmetic. EnsembleRetriever runs each wrapped retriever, converts every result to its rank position, sums weight divided by (c + rank) across retrievers, and returns the deduplicated candidates ordered by that fused score.

open as a page

In LangChain, what does MultiQueryRetriever buy you, and what does it cost?

level: seniorimportance: should knowfreq 50%

basics

~20 s

MultiQueryRetriever asks an LLM to rewrite the question into several phrasings, runs the base retriever on each, and returns the deduplicated union. It buys recall against vocabulary mismatch; it costs one extra LLM call of latency and a larger, unranked document set.

open as a page