How does a reranker like CohereRerank fit into a LlamaIndex query engine?
answer
- it lives between retrieve and synthesize
- query and document read together
- cost scales with candidates, not survivors
- it reorders, it cannot discover
- the score scale changes underneath you
basics
~20 sRerankers are node postprocessors: pass CohereRerank or SentenceTransformerRerank in node_postprocessors, and they rescore the retrieved nodes with a cross-encoder and keep only top_n. They run in list order between retrieval and synthesis, and they replace the embedding scores with their own.
solid answer
~50 sA reranker is a `BaseNodePostprocessor`, so it goes in the `node_postprocessors` list of a `RetrieverQueryEngine` — `index.as_query_engine(similarity_top_k=25, node_postprocessors=[reranker])` or `RetrieverQueryEngine.from_args(retriever, node_postprocessors=[...])`. The engine retrieves, then hands the node list and the query to each postprocessor in list order, then sends whatever survives to the response synthesizer. `SentenceTransformerRerank(model=..., top_n=4)` runs a local cross-encoder; `CohereRerank(api_key=..., top_n=4)` calls a hosted one. Unlike bi-encoder retrieval, a cross-encoder reads the query and document *together*, which is why it ranks far better — and why it costs one forward pass per candidate, so latency scales with `similarity_top_k`, not `top_n`. Two consequences bite in production: it can only reorder what retrieval already found, so it fixes ranking and never coverage; and it overwrites `node.score` with its own scale, so any similarity cutoff placed after it is now filtering against a different distribution.
code
python · 12 linesfrom llama_index.core.postprocessor import SentenceTransformerRerank
from llama_index.core.query_engine import RetrieverQueryEngine
reranker = SentenceTransformerRerank(
model="cross-encoder/ms-marco-MiniLM-L-6-v2", top_n=4
)
engine = RetrieverQueryEngine.from_args(
index.as_retriever(similarity_top_k=25),
node_postprocessors=[reranker],
)
response = engine.query("what is the refund window for annual plans?")
print(len(response.source_nodes))go deeper
Know that a reranker is passed in node_postprocessors and its job is to reorder the retrieved chunks and keep only the best few before the LLM sees them.
Explain the bi-encoder versus cross-encoder difference that makes a second pass worthwhile, and that postprocessors run in list order between retrieval and synthesis.
Bring the production facts: latency scales with the candidate pool, a reranker cannot create recall, scores are replaced so downstream cutoffs need retuning, and a hosted reranker needs an explicit degrade-or-fail policy.
Justify the reranker economically — what fraction of queries it actually changes, what it adds to p95 and to per-query cost, and whether a cheaper precision fix would buy the same answer quality.
## Where a reranker sits A LlamaIndex query engine runs three stages: retrieve, postprocess, synthesize. Rerankers are the headline occupants of the middle stage. They implement the `BaseNodePostprocessor` interface — given a `list[NodeWithScore]` and the query bundle, return a possibly shorter, possibly reordered list — and you install them by passing them in `node_postprocessors`, either on `as_query_engine` or on `RetrieverQueryEngine.from_args`. The order in the list is the order of execution, and each stage's output is the next stage's input. What survives the last postprocessor is exactly what becomes `response.source_nodes` and exactly what the LLM reads. ## Why a second scoring pass helps at all Retrieval uses a bi-encoder: the document was embedded once at ingestion, the query is embedded at query time, and the score is a geometric comparison of two vectors that never met. That is what makes it fast enough to search millions of chunks — and what makes it approximate, because the document's embedding had to be a single point summarizing everything it might be relevant to. A cross-encoder scores query and document *jointly*: both texts go through the model in one pass, so it can weigh which phrase in the document answers which part of the query. It is dramatically more accurate at ranking and completely unusable as a search primitive, because you would need one forward pass per document in the corpus. Hence the two-stage shape. Retrieval is a cheap wide net (`similarity_top_k=25–50`); the reranker is an expensive accurate filter over that small candidate set (`top_n=3–5`). ## The two rerankers you will be asked about `SentenceTransformerRerank(model="cross-encoder/ms-marco-MiniLM-L-6-v2", top_n=4)` runs a HuggingFace cross-encoder in your process. No network hop and no per-call bill, but it occupies memory, competes for CPU or GPU with the rest of the service, and is slow on CPU once the candidate count grows. `CohereRerank(api_key=..., top_n=4)` from the `llama-index-postprocessor-cohere-rerank` package calls a hosted reranking endpoint. Better quality for zero local footprint, at the price of a network round trip in the hot path, a per-call cost, a rate limit and a new external dependency your query path can fail on. Decide whether a reranker outage degrades to unreranked results or fails the request — that is a design decision, not an accident. ## The three things that bite **1. Latency tracks the candidate count.** The reranker performs work proportional to `similarity_top_k`, not `top_n`. Doubling the candidate pool doubles rerank time. For hosted rerankers it also doubles the payload. This is the single knob that most often blows a p95 budget after someone "improves recall" by raising top-k. **2. A reranker cannot invent recall.** It reorders the retrieved list. If the gold passage is not in the candidates, the reranker's only effect is to promote the best of the wrong answers — often making a wrong answer *more* confident. When reranking does not help, check whether the passage was ever retrieved before blaming the reranker. **3. Scores change scale.** The reranker replaces `node.score` with its own relevance score, which is not on the embedding-similarity scale. If your postprocessor list is `[SimilarityPostprocessor(similarity_cutoff=0.7), reranker]`, the cutoff filters embedding scores — coherent. Reverse them and the cutoff is now comparing 0.7 against a cross-encoder output, and depending on the model that either drops nearly everything or nothing at all. Ordering in `node_postprocessors` is semantics, not decoration. ## Tuning it Set `top_n` from answer quality with the candidate pool fixed, and set `similarity_top_k` from recall@k with `top_n` fixed — moving both together makes the results uninterpretable. Typical landing zones are 20–50 candidates and 3–5 survivors. If quality keeps improving past `top_n=8`, that usually means the reranker is not confident and your retrieval is noisy; if it peaks at 2, your corpus has one obviously right chunk per question and the reranker may be earning its latency only on a minority of queries. Measure that minority explicitly. Log both orderings for a sample of traffic and compare where the reranker actually changed the top few nodes. It is common to find the reranker is decisive on 10–20% of queries and inert on the rest — which is still worth it if those queries matter, and is worth knowing before you pay for it on every request. ## Composing with other postprocessors Rerankers coexist with the rest of the postprocessor family — similarity cutoffs, metadata-based filters, and LLM-based rerankers that ask a model to pick relevant nodes (more accurate, far more expensive, and another LLM call in the hot path). Keep the chain short: every postprocessor is latency on every query, and a chain nobody can explain the ordering of is a chain that will be misordered.
- Why can a reranker afford a cross-encoder when retrieval cannot?Because it only ever sees a small candidate set. A cross-encoder needs one forward pass per query-document pair, which is impossible across millions of chunks but trivial across the twenty-five the retriever returned. Retrieval must use precomputed document embeddings compared to a query embedding, which is fast but approximate; the reranker buys back the accuracy on a tiny slice.
- Reranking is enabled and quality did not improve. How do you tell whether the reranker is the problem?Check whether the gold passage was in the retrieved candidates at all by calling `retriever.retrieve(q)` directly. If it is absent, the reranker was never given a chance and the fix is recall — higher top-k, hybrid retrieval, better chunking. If it is present but ranked low after reranking, then the reranker or its model choice is genuinely at fault.
- What breaks if you order node_postprocessors as [reranker, similarity cutoff]?The cutoff is applied to the reranker's scores, not the embedding similarities it was tuned against, because the reranker overwrites `node.score` with a different scale. Depending on the model that silently drops everything or nothing. Put score-scale-dependent filters before the reranker, or retune the threshold against the reranker's own distribution.
- How would you decide between a local cross-encoder and a hosted reranking API?Weigh latency, footprint and failure mode. A local model has no network hop or per-call fee but consumes CPU or GPU inside the service and is slow on CPU at larger candidate counts. A hosted one adds a round trip, a bill and a rate limit, and puts a third party in the query path — so you must decide explicitly whether an outage degrades to unreranked results or fails the request.
saying these in an interview costs you the question
- Expecting a reranker to surface documents retrieval never returned
- Thinking rerank latency scales with top_n rather than candidate count
- Placing a similarity cutoff after a reranker without retuning it
- Treating a hosted reranker as free of latency and failure risk
- Tuning similarity_top_k and top_n at the same time and reading the result