What does a cross-encoder reranker buy over the vector index's own ranking?
answer
- separate encodings versus a joint one
- nothing precomputable, cost per candidate
- two stages, two different metrics
- cannot rescue what was never retrieved
- candidate count is the latency dial
basics
~20 sA cross-encoder reads the query and the candidate together, so it can judge fine-grained relevance that independently-encoded vectors cannot. It raises precision at the top of the list but cannot recover anything the first stage failed to retrieve, and its cost grows with the number of candidates scored.
solid answer
~50 sThe vector index uses a bi-encoder: query and chunk are embedded separately, so the chunk's vector was computed long before your query existed and can carry no query-specific signal. That is what makes it fast and precomputable, and also what makes it coarse. A cross-encoder scores a query-chunk *pair* in one pass, with full attention across both, which is far more discriminating — but nothing can be precomputed, so cost is linear in candidates scored. Hence the two-stage shape: retrieve broadly, then rerank the head. A typical setup retrieves and fuses around 100 candidates, reranks them, and passes the best dozen to the model. Two consequences matter in an interview. First, reranking fixes *precision*, never *recall* — if the answering chunk was not in the 100, no reranker can rescue it, so first-stage recall@N is the metric that bounds the whole pipeline. Second, the candidate count is a live latency and cost dial, typically adding tens to a few hundred milliseconds per query.
go deeper
Know that retrieval happens in two stages — a fast index pass that gathers candidates and a slower, more accurate model that reorders the best of them before anything reaches the user.
Explain the bi-encoder versus cross-encoder distinction: independent encodings that can be precomputed against joint encoding that cannot, and why that forces reranking to run only over a small candidate set.
Demonstrate the metrics split — recall@N bounds the pipeline and the reranker is measured on ordering quality — and derive the candidate count from a recall curve against a p95 latency budget, with a defined degradation path on timeout.
Own the quality-versus-cost frontier across the retrieval tier: where late interaction, a purpose-built cross-encoder or an LLM reranker each pay, which business signals belong in the final ordering, and how the tier degrades under load.
## Bi-encoder versus cross-encoder A **bi-encoder** encodes each text independently into a vector. Chunks are embedded once at ingestion and stored; at query time only the query is embedded, and the search is a nearest-neighbour lookup in an index. Because the two encodings never see each other, the chunk's representation must be a single fixed summary that works for every possible future query. That is a strong compression, and it is where the coarseness comes from: two chunks that both sit near the query in the space may differ enormously in whether they actually answer it. A **cross-encoder** takes the query and one candidate concatenated as a single input and produces a relevance score, with attention running across both texts. It can see that the query's constrained noun phrase appears as the subject of the candidate's key sentence, that a negation flips the meaning, that a date qualifier does not match. Nothing is precomputable — every query-candidate pair is a fresh forward pass — so the cost is N model calls for N candidates. That asymmetry dictates the architecture. You cannot cross-encode millions of chunks per query, so you use the cheap index to cut the corpus to a hundred candidates and spend the expensive model only there. ## Where it sits in the pipeline In a hybrid setup the shape is: lexical retriever and dense retriever each return their top candidates, the two lists are fused, the fused head is reranked, and the best few are handed to the generation step. In an e-discovery search the concrete numbers might be 100 fused candidates reranked down to 12 shown to the paralegal or fed to the model. ## Precision, not recall This is the point interviewers actually probe. Reranking reorders a set; it cannot conjure a document into it. If the responsive chunk ranked 340th in the first stage, retrieving 100 candidates means the reranker never sees it and the pipeline fails no matter how good the reranker is. So the two stages have different metrics: measure **recall@N** for the retrieval stage — what fraction of known-relevant chunks appear anywhere in the N candidates — and measure ordering quality, such as nDCG@10, for the reranker. Optimizing the reranker while first-stage recall@100 sits at 70% is wasted effort; the ceiling is 70%. ## Cost and latency Reranking adds a synchronous stage to every query. A small cross-encoder over 100 short passages, batched on a GPU, typically lands in the tens of milliseconds; a larger model, longer chunks, a CPU deployment or a hosted rerank service pushes it into the hundreds. Cost scales with candidates times chunk length, so the two dials are N and how much text each candidate contributes. Practical levers: rerank fewer candidates for interactive paths and more for batch or high-stakes ones; truncate each candidate to the region that matched; batch aggressively; and cache reranked results for repeated queries. ## Choosing N N is an explicit recall-versus-latency trade. Derive it, do not guess it: on a labelled query set, plot first-stage recall@N as N grows. The curve rises steeply then flattens. Pick N just past the knee, because beyond it you are paying linear rerank cost for negligible additional recall. Re-derive N whenever the embedding model, the chunking or the corpus changes, since the shape of that curve is a property of the retrieval stage. ## Alternatives and variants **Late-interaction models** (the ColBERT family) sit between the two extremes: they store per-token vectors and compute a cheap interaction at query time, buying much of the cross-encoder's discrimination at lower query cost, at the price of a substantially larger index. **LLM rerankers** prompt a general model to score or order candidates, sometimes listwise over a whole batch. They can be very good and need no specialized model, but they are slower and dearer than a purpose-built cross-encoder, and they inherit judge-style biases such as sensitivity to candidate order. **Feature-based reranking** — blending the relevance score with recency, authority, custodian or access-tier signals — is often what production actually needs. A pure semantic reranker has no idea that a 2019 draft matters less than a 2024 signed version. ## Failure modes to watch The reranker is another model with its own domain sensitivity: one trained on web passages may underperform on dense legal or clinical prose, and it is worth evaluating on your own labelled set rather than trusting a leaderboard. It also introduces a second point of drift — when retrieval quality changes, the reranker's input distribution changes with it. And because it sits synchronously in the request path, its failure mode must be decided in advance: on timeout, degrade to the fused first-stage order rather than failing the query.
- Your reranker is excellent but end-to-end quality is flat. Where do you look first?At first-stage recall@N. The reranker can only reorder what it receives, so if the relevant chunk is not among the candidates the pipeline is already capped. Measure recall@N over a labelled set; if it is low, the fix is upstream — better chunking, a hybrid lexical leg for exact terms, a larger N, or a different embedding model — not a better reranker.
- How do you choose how many candidates to rerank?Plot first-stage recall@N against N on a labelled query set and take the knee of the curve, then check that the resulting rerank latency fits the p95 budget. Beyond the knee each extra candidate costs a full model pass for almost no recall. Re-derive it after any change to chunking, embedding model or corpus size, and consider a lower N for interactive traffic and a higher one for batch or high-stakes queries.
- What should happen when the reranker times out?Degrade, do not fail. The fused first-stage ordering is a usable answer, just less precise, so serve it and record the degradation as a metric rather than returning an error. Give the reranker a hard timeout well inside the request budget, and if degraded responses become common, treat that rate as a capacity signal for the reranking tier.
- When would you prefer a late-interaction retriever over a cross-encoder reranker?When query-time budget is tight but you still need finer matching than single-vector similarity gives. Late-interaction models such as the ColBERT family store per-token vectors and compute a lightweight interaction at query time, recovering much of the cross-encoder's discrimination without a full forward pass per candidate. The trade is index size and operational complexity, since storing per-token vectors can multiply storage by an order of magnitude.
saying these in an interview costs you the question
- A reranker improves recall as well as ordering
- Reranking is free because it only touches the top results
- Cross-encoder scores can be precomputed and stored like embeddings
- If reranking is in place, first-stage retrieval quality stops mattering
- One reranker is equally good on any domain's text