Why does a semantic search pipeline retrieve with a bi-encoder and rerank with a cross-encoder?
answer
- two stages, two different jobs
- one side can be precomputed, one cannot
- recall first, precision second
- the shortlist is a hard ceiling
- joint encoding is why it cannot be cached
basics
~20 sA bi-encoder embeds documents independently, so their vectors are built offline and searched in milliseconds. A cross-encoder reads query and document together: much more accurate, but impossible to precompute, so it only rescores the shortlist the first stage returns.
solid answer
~50 sThe two stages exist because the accurate model is too expensive to run over the whole corpus. A **bi-encoder** encodes the query and each document separately into vectors, so document vectors are computed once at index time and a nearest-neighbour search over millions of them costs milliseconds — but the two sides never interact, so the match is approximate. A **cross-encoder** takes the query and one document as a single joint input and emits a relevance score; attention runs across the pair, so it can align a query phrase with a differently-worded passage. That score only exists once the pair exists, so it cannot be precomputed — scoring a whole corpus means one forward pass per document. Hence the funnel: the cheap stage maximises recall into a shortlist of tens to a couple of hundred, the expensive stage supplies precision by reordering it. Anything the first stage misses is unrecoverable.
go deeper
Know that search happens in two passes: a fast one that pulls a few dozen candidates from the whole corpus, and a slower, smarter one that reorders just those candidates. Be able to say why the smart one cannot be run over everything.
Explain the mechanical reason: the bi-encoder encodes each side independently so document vectors are precomputed and indexed, while the cross-encoder scores a query-document pair jointly and so must run at query time. Name the shortlist as the boundary between recall and precision.
Show you can diagnose with it. Split a quality complaint into retrieval failure versus ranking failure by checking candidate-set membership, size the shortlist from measurements rather than habit, and account for the whole latency budget including the candidate-text fetch people forget.
Own the tradeoff between pipeline complexity and payoff. Be ready to argue when a second stage is not worth its serving cost, infrastructure and failure surface, and to set the quality-per-millisecond bar a reranking tier must clear before it ships.
## The two encoder shapes A **bi-encoder** (also called a dual encoder) runs the query through an encoder and each document through an encoder — usually the same weights — and produces one vector per side. Relevance is then a cheap arithmetic comparison of two vectors. The crucial property is that the document side never sees the query, so every document vector can be computed once, ahead of time, when you build the index. At query time you pay for exactly one forward pass (the query) plus a lookup in an approximate-nearest-neighbour index. A **cross-encoder** does something structurally different: it concatenates the query and one candidate document into a single input, runs them through the model together, and emits a scalar relevance score rather than a vector. Because attention runs across the pair, the model can directly relate a phrase in the query to a differently-worded phrase in the document — a job posting that says "run our Kubernetes cluster" against a résumé that only ever says "container orchestration" — instead of hoping the two independently-produced vectors happen to land near each other. That interaction is exactly why the score cannot be precomputed. It does not exist until the pair exists. Scoring a million-document corpus for one query means a million forward passes. For any corpus above toy size that is off the table as a first stage. ## Why the funnel So you build a funnel with two different objectives: - **Stage one — recall.** Cheap, approximate, run over everything. Its job is not to get the order right; its job is to make sure the genuinely relevant items are somewhere in the candidate set. Typical shortlists run from a few tens to a couple of hundred candidates. - **Stage two — precision.** Expensive, accurate, run only over the shortlist. Its job is to reorder those candidates so the best ones surface into the top few positions the user actually reads. The single most important consequence in production: **the reranker can only reorder what stage one returned.** If the right résumé never entered the shortlist, no amount of reranking quality recovers it. When search is failing, the first diagnostic question is always "was the answer in the candidate set at all?" — because that splits the problem into a retrieval bug and a ranking bug, which have completely different fixes. ## Sizing the shortlist Shortlist depth is the main tuning dial connecting the stages. Deeper shortlists monotonically improve the chance the right item is present, and monotonically increase reranking cost and latency, which grows roughly linearly with candidate count. The practical method is empirical: sample real queries with known good answers, and measure how often the good answer appears anywhere in the candidate set at depth 20, 50, 100, 200. That curve flattens somewhere; take the depth just past the knee, then check the resulting latency against your budget. Choosing a depth by intuition is the most common reason a pipeline is simultaneously slow and inaccurate. ## Where the latency goes Accounting for the end-to-end budget usually shows four components, and engineers routinely misattribute them: 1. **Embedding the query** — one forward pass, but if the embedding model is a remote service, this is a network round trip and can dominate a fast pipeline. 2. **The nearest-neighbour search** — normally the smallest term, single-digit milliseconds on a well-built index. 3. **Fetching candidate text** — the reranker needs the actual document text, not the vector, so there is a store lookup for every candidate. This is frequently forgotten at design time and then discovered as the bottleneck. 4. **Reranking** — usually the dominant term, and the one that scales with shortlist depth. If the pipeline is too slow, the lever with the best cost-to-quality ratio is almost always the shortlist depth, not the index. ## When you can skip the second stage A reranker is not mandatory. Skip it when the corpus is small enough that first-stage ordering is already good, when the latency budget is brutal (type-ahead suggestions), or — most honestly — when an offline comparison on your own queries shows it buys no measurable lift. Conversely, when your first stage is retrieving topically-similar-but-useless results, a reranker is usually a bigger win per unit of effort than swapping the embedding model. ## Failure modes to name in an interview - Reranking a shortlist that never contained the answer, and concluding the reranker is bad. - A reranker trained on general web relevance applied to a specialist domain, where it may score worse than the first stage. - **Text mismatch between stages**: you embedded a condensed summary of each document but hand the reranker the raw 3,000-word original, so the two stages are judging different objects. - Treating first-stage scores and reranker scores as comparable numbers and mixing them in one ordering; they come from different models and different scales.
- If reranking is so much more accurate, why not just make the shortlist very deep?Because reranking cost and latency grow roughly linearly with shortlist depth, while the chance of adding a new relevant item flattens out quickly. Past the knee of that curve you pay proportionally more for almost no quality. Measure how often the correct item appears at increasing depths on real queries, then pick the depth just beyond the flattening point that still fits your latency budget.
- Search is returning topically related but useless results. How do you tell whether that is a retrieval problem or a ranking problem?Check whether the correct item is present anywhere in the first-stage candidate set. If it is present but ranked low, it is a ranking problem — add or improve the reranker. If it is absent, no reranker can help; the fix belongs in the first stage: what text you embedded, which embedding model, shortlist depth, or a missing lexical retrieval leg.
- Can the same model serve as both stages?Not usefully. The bi-encoder shape exists so document vectors can be computed once and searched; the cross-encoder shape exists so the query and document interact. A cross-encoder has no per-document vector to index, and a bi-encoder gains nothing from being run over a shortlist it already ranked. They are different architectures chosen for different cost profiles, even when fine-tuned from the same base model.
saying these in an interview costs you the question
- Claims cross-encoder scores can be cached per document in the index
- Thinks reranking can recover items the first stage never returned
- Sizes the shortlist by intuition instead of measuring candidate coverage
- Compares first-stage similarity scores directly against reranker scores
- Assumes a reranker always improves quality regardless of domain