skip to content

Cross-Encoder Reranking

A cross-encoder reads the query and a candidate passage together rather than embedding them separately. That joint attention makes it much more accurate — and far too slow to run over a whole corpus.

on this pageshow

questions

4

Why can a cross-encoder reranker score query-passage pairs more accurately than a bi-encoder?

level: middleimportance: must knowfreq 72%

answer

  1. two towers versus one input
  2. when do the sides meet?
  3. attention runs across the pair
  4. score, not an embedding
  5. nothing to precompute or index

basics

~20 s

A cross-encoder feeds the query and the passage through one model together, so every query token can attend to every passage token. A bi-encoder encodes each side separately and compares two fixed vectors, so the sides never interact before the similarity is computed.

solid answer

~50 s

A bi-encoder is a two-tower design: the passage is encoded once, offline, into a single vector, the query is encoded at search time, and relevance is a cheap similarity between the two vectors. All the interaction between query and document is squeezed into that one dot product, so anything the passage vector did not happen to preserve is lost. A cross-encoder concatenates the query and the passage into one input and runs a single forward pass over the pair, so attention runs **across** the two — the model can see that this passage's "dosage" clause is the one the query's "how much" is asking about. It emits a single relevance score rather than an embedding. The price is structural: because the representation depends on the query, there is no passage vector to precompute or index, so a cross-encoder can only ever rescore a candidate list someone else produced.

code

python · 13 lines
python
# Bi-encoder: passage vectors are computed once and reused for any query.
doc_vecs = [[0.1, 0.9], [0.8, 0.2]]
query_vec = [0.2, 0.8]
bi_scores = [sum(q * d for q, d in zip(query_vec, doc)) for doc in doc_vecs]

# Cross-encoder: the score is a function of the pair, so nothing is reusable.
def cross_encoder_score(query, passage):
    # stand-in for one joint forward pass over [query; passage]
    overlap = set(query.split()) & set(passage.split())
    return len(overlap) / (len(query.split()) or 1)

print(bi_scores)
print(cross_encoder_score("dosage for cats", "recommended dosage for adult cats"))

go deeper

for a junior

Know the shape: a bi-encoder turns each passage into a vector ahead of time and compares vectors; a cross-encoder reads the query and the passage together and returns one relevance score. Be able to say why only the first one can be indexed.

for a middle

Explain the mechanism — joint attention across query and passage tokens in every layer versus a single dot product between two independently-computed vectors — and derive the consequence that no passage vector exists to precompute.

for a senior

Show you reason about the staging as a system: first stage owns recall, reranker owns top-of-list precision, and a recall miss is unrecoverable downstream. Be ready to say which metric you would look at to tell the two failures apart.

for a principal

Frame it as an axis rather than two boxes: how early query and document interact trades precomputability for accuracy, with late-interaction designs in between. Own the call about how much of the index budget and online compute you are willing to spend for that accuracy.

## Two ways to compare a query and a passage Every retrieval system has to answer one question: how well does this passage answer this query? There are two architectural answers, and the difference between them explains nearly everything about where each one is used. A **bi-encoder** (also called a dual encoder, or two-tower model) runs the query through an encoder to get one vector, runs the passage through an encoder to get another vector, and scores the pair with a cheap similarity function — usually a dot product or cosine. The crucial property is *separability*: the passage vector does not depend on the query at all. So you compute every passage vector once, at index time, store them, and at search time you embed only the query and compare it against the stored vectors. That is what makes billion-scale search possible: the expensive neural work happened offline, and the online work is arithmetic over vectors that an approximate-nearest-neighbour index can prune aggressively. A **cross-encoder** refuses that separation. It concatenates the query and the passage into one input sequence — conceptually `[query; separator; passage]` — and runs the whole thing through the transformer in a single forward pass. On top of the joint representation sits a small scoring head that emits one number: a relevance logit for that specific pair. The output is a *score*, not an embedding. ## Why joint attention buys accuracy Inside the transformer, self-attention lets every token attend to every other token. In a cross-encoder, that means query tokens attend to passage tokens and vice versa, in every layer. The model can perform soft term matching ("how much" against "5 mg per kg"), resolve which of several entities in the passage the query is about, notice a negation that flips relevance, and weigh a phrase's importance conditionally on what was asked. None of that can happen in a bi-encoder, because by the time the two sides meet they have already been compressed to fixed-length vectors that were computed in ignorance of each other. The bi-encoder's passage vector must be a *query-independent* summary — good enough for every query it might ever face. The cross-encoder's representation is built for exactly one query. This is why, on standard passage-ranking evaluations, a cross-encoder reranking a bi-encoder's top candidates usually lifts ranking quality substantially over the bi-encoder's own ordering, often with a model that is no larger. The gain comes from the interaction, not from parameter count. ## The cost is structural, not incidental The same property that buys accuracy forbids indexing. There is no standalone document representation to store, because the representation is a function of the pair. You cannot precompute scores either, unless you enumerate every (query, passage) combination, which is unbounded. So the work is strictly online and strictly per pair: scoring `k` candidates costs `k` forward passes, and scoring a whole corpus of `N` passages would cost `N` forward passes per query. For any realistic corpus that is impossible within an interactive latency budget. Hence the standard **retrieve-then-rerank** staging. A cheap, indexable first stage (a bi-encoder over an ANN index, a lexical retriever, or both) produces a candidate list with high recall but imperfect ordering. The cross-encoder then rescores only that list and reorders it. The first stage optimizes for *recall* — anything it misses is gone forever, since the reranker never sees the rest of the corpus. The second stage optimizes for *precision at the top* — getting the genuinely best passages into the few slots the generator will actually read. ## What this means in practice - The reranker cannot rescue bad recall. If the right passage is not in the candidate list, no amount of reranking helps; that is a first-stage problem. - Caching works only at the pair level. A repeated identical query over an unchanged candidate set can reuse scores; a slightly different query cannot reuse anything. - Scores are not comparable to cosine similarities and often are not calibrated across queries. Treat them as an ordering signal for one query unless you have checked otherwise. - Index freshness is cheap for the reranker: adding documents changes nothing about the reranker, because it stores nothing. It changes the first-stage index only. ## The middle ground exists Between the two extremes sit late-interaction designs that keep per-token document representations and defer a cheaper interaction to query time. They trade index size for some of the cross-encoder's accuracy. For interview purposes, the point to hold onto is the axis itself: the earlier the query and the document interact, the more accurate and the less precomputable the system becomes.

  • If a cross-encoder is more accurate, why not use one for first-stage retrieval?
    Because scoring is per pair, first-stage use would mean one forward pass per document in the corpus, per query — millions of transformer passes for a single search. There is also nothing to index: the passage representation depends on the query, so no offline structure can prune the candidate set. Cross-encoders are only affordable over a short candidate list someone cheaper already produced.
  • What does a cross-encoder actually emit for a pair, and is it comparable across queries?
    A single scalar relevance score from a small head on top of the joint representation — not an embedding and not a cosine similarity. Scores are trained to order passages within one query, so they are reliable for ranking but often poorly calibrated across queries. If you want an absolute cutoff, verify calibration on your own labelled data rather than assuming a fixed threshold transfers.
  • Can reranking compensate for a first stage with poor recall?
    No. The reranker only ever sees the candidates handed to it, so a relevant passage missing from the candidate list is unrecoverable. Diagnose the stages separately: measure recall of the first stage at the candidate depth you feed forward, and measure top-of-list quality after reranking. Recall failures are fixed by the retriever — better embeddings, hybrid lexical matching, chunking — not by a stronger reranker.

saying these in an interview costs you the question

  • Says the cross-encoder produces document embeddings you can store in a vector index
  • Claims the difference is just that cross-encoders are bigger models
  • Describes cross-encoder scores as cosine similarity between two vectors
  • Assumes passage scores can be precomputed offline like embeddings
  • Believes a strong reranker removes the need for a good first-stage retriever

context

open as a page

How do you size cross-encoder rerank depth k against a 400ms p95 latency budget?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Cross-encoder cost is O(k) forward passes over query-passage pairs, so measure milliseconds per pair at your real sequence length and batch size, then solve for the k that fits after subtracting retrieval and generation time from the budget. Do the arithmetic at p95, not at the mean.

open as a page

Why can an MS MARCO-trained cross-encoder underperform your bi-encoder on a specialist corpus?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A reranker learns a notion of relevance from its training distribution. MS MARCO is short, general web-search queries over web prose, so on a specialist corpus with unfamiliar terminology and differently-shaped queries the reranker can confidently reorder good candidates into a worse order.

open as a page

Hosted rerank API or self-hosted cross-encoder — how do you decide between them?

level: principalimportance: should knowfreq 33%

basics

~20 s

Decide on three axes: unit economics, since hosted reranking is billed per document scored and rerank depth multiplies every query; tail latency and control, since a network hop adds variance you cannot tune; and data exposure, because reranking sends your corpus passages out, not just the query.

open as a page