When would you wrap a LangChain retriever in ContextualCompressionRetriever?
answer
- A stage after search, before the prompt
- The query travels with the documents
- Filter, rerank or trim — three shapes
- Cost scales with candidate count
- Cannot resurrect a missed document
basics
~20 sWhen you need to retrieve wide but prompt narrow. ContextualCompressionRetriever runs a base retriever, then passes its documents plus the query through a compressor that filters, reranks or trims them — so you can raise k for recall without paying for all of it in prompt tokens.
solid answer
~50 s`ContextualCompressionRetriever(base_compressor=..., base_retriever=...)` is a post-retrieval stage. The base retriever runs normally; the compressor then receives the returned Documents *and the query* and may drop them, reorder them, or rewrite their content to only the query-relevant part. That query-awareness is the point — a vectorstore ranks by embedding proximity, while a compressor can apply a much stronger relevance signal to a small candidate set. Which compressor you pick sets the cost curve: `EmbeddingsFilter` cuts by cosine threshold for the price of one embedding per document; `CrossEncoderReranker` or a hosted reranker scores each query-document pair with a dedicated model and keeps `top_n`; `LLMChainFilter` keeps or drops whole documents with one LLM call each; `LLMChainExtractor` rewrites each document down to relevant sentences, also one call each. `DocumentCompressorPipeline` chains several. It can never recover a document the base retriever missed.
go deeper
Know that it sits after retrieval and cuts the retrieved documents down before they reach the prompt, using the query to decide.
Distinguish the compressor families — embedding filter, cross-encoder rerank, LLM filter, LLM extract — and say what each costs per candidate.
Argue the retrieve-wide/prompt-narrow tradeoff with numbers, order a compressor pipeline cheap-first, and name the extraction risk to citation fidelity.
Own the latency and cost budget of the whole retrieval path and decide where a stronger ranking signal is worth a dedicated model in the request path.
## Retrieve wide, prompt narrow There is a structural tension in RAG. Raising `k` improves the chance the answer is somewhere in the retrieved set, but every retrieved document is prompt tokens on every request, and long contexts dilute attention so the relevant passage buried at position nine gets used less than the irrelevant one at position one. Contextual compression resolves the tension by splitting the decision into two stages with different cost profiles. Stage one, the base retriever, is cheap and approximate: cast a wide net with a large `k`. Stage two, the compressor, is expensive per document but only sees the candidates, and applies a far stronger relevance signal to cut back to what actually goes in the prompt. ## The interface `ContextualCompressionRetriever` takes a `base_retriever` and a `base_compressor` and is itself a retriever — `invoke(query)` in, `list[Document]` out — so it drops into any pipeline the base retriever fitted. Internally it calls the base retriever, then hands the resulting Documents *and the original query* to the compressor's `compress_documents` method. The compressor may return fewer documents, reordered documents, or documents whose `page_content` has been rewritten. "Compression" therefore covers three quite different operations — filtering, reranking, and extraction — under one interface. ## Choosing a compressor **EmbeddingsFilter.** Embeds the query and each candidate and drops those below a similarity threshold (and/or keeps a top `k`). One embedding call per document, no generation, milliseconds. Its relevance signal is the same kind of signal the vectorstore already used, so it mostly trims the tail rather than genuinely re-judging relevance. Cheapest useful option, and a good first filter in a pipeline. **Cross-encoder reranking.** `CrossEncoderReranker(model=..., top_n=...)` runs a model that sees the query and the document *together* and outputs a relevance score, keeping the best `top_n`. Because the pair is encoded jointly rather than as two independent vectors, this is a materially stronger signal than cosine similarity, and it is usually the single highest-quality-per-millisecond upgrade available to a RAG pipeline. Cost is one model forward pass per candidate, which is why you rerank twenty or fifty candidates, not two thousand. Hosted equivalents exist as compressors too, trading local compute for a network call and a per-query fee. **LLMChainFilter.** Asks the LLM, per document, whether it is relevant, and keeps or drops it whole. One generation call per candidate: accurate, slow, expensive, and parallelisable at best. Reasonable for small candidate sets or offline pipelines. **LLMChainExtractor.** Asks the LLM, per document, to return only the sentences relevant to the query. This is the only compressor that genuinely shrinks token count *within* a document, which matters when your chunks are large. It is also the riskiest: the model can drop the qualifying clause that changed the meaning, or subtly reword the extract, and now your citation no longer matches the source text. Do not use it where the exact wording is contractually or clinically important. **DocumentCompressorPipeline.** Chains transformers and compressors in order — for example a redundancy filter to remove near-duplicate chunks, then an embeddings filter, then a reranker. Order matters: put the cheap filters first so the expensive stage sees fewer documents. ## What it cannot do Compression is strictly subtractive with respect to the candidate set. If the base retriever did not return the document that contains the answer, no compressor will produce it — there is no second search, no query rewriting, no traversal to neighbouring chunks. This is the most common misconception in interviews, and the correct instinct follows directly from it: when adding a compressor, **raise the base `k` at the same time**. Compression without widening the net just gives you fewer documents than you had before. ## Cost model to state out loud Let the base retriever return `k` candidates. An embeddings filter costs `k` embeddings; a cross-encoder costs `k` forward passes; an LLM-based compressor costs `k` generations. So `k=50` is entirely reasonable in front of a cross-encoder and financially absurd in front of an LLM extractor. Latency compounds too, and unlike the base retrieval it scales linearly with `k` rather than logarithmically. Budget the stage explicitly: candidate count in, document count out, milliseconds and currency per query. ## Where it fits against the alternatives If the right document is retrieved but ranked badly, rerank. If the right document is retrieved but drowned in irrelevant neighbours, filter. If chunks are large and mostly boilerplate, extract. If the right document is *not retrieved at all*, none of these help — fix ingestion, chunking, the embedding model, or add lexical retrieval alongside the dense one.
- Why does a cross-encoder reranker beat cosine similarity on the same candidates?Because it encodes the query and the document together and scores the pair directly, so it can model interactions between them — negation, qualifiers, which entity the question is actually about. Bi-encoder retrieval compresses each side into an independent vector before they ever meet, which is what makes it fast enough to search millions of items and weaker at fine-grained judgement over a handful.
- What is the risk of LLMChainExtractor specifically?It rewrites document content, so the text reaching the prompt is no longer verbatim source text. The model can drop a qualifying clause that inverted the meaning, or paraphrase in a way that breaks a citation's exact match against the original. Where wording is legally or clinically load-bearing, prefer compressors that keep or drop whole documents rather than editing them.
- You add a compressor and answer quality drops. What is the likely cause?You compressed without widening. Compression only subtracts from what the base retriever returned, so leaving k at 4 and then filtering down to 2 gives the model less evidence than before. The pattern is to raise the base k substantially — enough candidates for the stronger signal to choose from — and let the compressor cut back to the budgeted number.
saying these in an interview costs you the question
- Thinks the compressor re-queries the vectorstore
- Adds compression without raising the base k
- Uses an LLM compressor over fifty candidates
- Assumes compressed text is still verbatim source
- Believes it fixes documents missing from the index