skip to content

Why does RAG retrieval use a bi-encoder rather than scoring every document with a cross-encoder?

level: juniorimportance: must knowfreq 75%

answer

  1. the cost turns on where the model runs
  2. documents are encoded before the query exists
  3. independent encoding is what allows an index
  4. pair scoring costs one pass per document
  5. precompute once versus rescore per query

basics

~20 s

A bi-encoder embeds every document once, offline, so a query is answered by a nearest-neighbour lookup over precomputed vectors. A cross-encoder must run the model on each query-document pair, so scoring a whole corpus per query is infeasible.

solid answer

~50 s

The two arrangements differ in *when* the model sees the pair. A bi-encoder (dual encoder) encodes the query and the document **independently** into fixed-length vectors, so documents can be embedded once at index time and the online cost is one query encoding plus an approximate-nearest-neighbour lookup — sublinear in corpus size. A cross-encoder concatenates query and document and runs a forward pass over the pair, which is more accurate because every query token can attend to every document token, but produces no reusable document vector: scoring a four-million-chunk corpus would mean four million forward passes per query. That is why production RAG is two-stage — a bi-encoder retrieves a broad candidate set cheaply, and a slower query-aware scorer re-ranks only the few dozen candidates that survive. The bi-encoder's price for that speed is an information bottleneck: the document vector must be committed before anyone knows what will be asked.

go deeper

for a junior

Be able to say plainly that documents are embedded once ahead of time and a query is matched against those stored vectors, and that comparing the query against every document with a full model pass would be far too slow.

for a middle

Explain the mechanism: independent encoding is what makes precomputation and approximate-nearest-neighbour search possible, and give the cost arithmetic for a realistic corpus rather than just saying one approach is faster.

for a senior

Show the operational consequences — the index is pinned to a model version, query encoding sits on the hot path, and first-stage recall caps everything downstream — and justify the candidate-set size you would actually run.

for a principal

Own the cascade as a cost-quality frontier: decide how much of the accuracy budget belongs in the retriever versus the re-scorer, given corpus size, query volume, latency SLO and the re-indexing cost of ever changing the retriever.

## Two ways to compare a query and a document Any neural relevance model has to bring a query and a document together somewhere. There are two arrangements, and the choice determines the entire shape of the retrieval system. In a **bi-encoder** (also called a dual encoder) the query and the document are pushed through the encoder *separately*. Each becomes a fixed-length vector — a few hundred to a couple of thousand floats. The relevance score is then a cheap arithmetic operation between the two vectors, almost always a dot product or cosine similarity. In a **cross-encoder** the query and the document are joined into a single input and pushed through the model *together*. The model reads them jointly and emits one relevance score. There is no separable document representation to store: the score exists only for that specific pair. ## Why independence is the whole point Because a bi-encoder encodes a document without knowing the query, the document's vector can be computed at **index time**, long before any user shows up. Ingestion becomes a batch job: chunk the corpus, embed once, write the vectors into a vector index. At query time you pay for exactly one forward pass (the query) plus a similarity search. That search is not a linear scan either. Because every item is a point in the same vector space, an approximate-nearest-neighbour index can prune most of the corpus and return the top *k* in roughly logarithmic or sublinear time. Precomputation plus ANN is what makes semantic search over millions of chunks answerable in tens of milliseconds. ## The cost arithmetic Make it concrete with a four-million-chunk corpus of 800-token chunks. - **Bi-encoder:** four million forward passes *total*, paid once at ingestion (and again only when you change models). Per query: one short forward pass and an index lookup. - **Cross-encoder over the full corpus:** four million forward passes over ~800-token pairs *per query*. Even at a thousand pairs per second, a single query takes over an hour of GPU time. The asymmetry is not a small constant factor; it is the difference between a feasible system and an impossible one. ## What the bi-encoder gives up The independence that buys the speed is also the weakness. A document's vector must be a single summary of everything the chunk might ever be asked about, committed before the question exists. Fine-grained interactions — an exact identifier, a negation, a rare term that decides relevance — get averaged into one point. A cross-encoder, by contrast, lets every query token attend to every document token, and its accuracy advantage on hard pairs is consistently measurable. The standard resolution is a **cascade**: use the bi-encoder for high-recall, low-precision retrieval of maybe the top 50-200 candidates, then spend a query-aware scorer only on that small set. You get corpus-scale reach at bi-encoder cost and pair-level precision where it matters. ## Consequences you inherit by choosing a bi-encoder 1. **The index is tied to a model version.** Vectors from model A are meaningless in a space built by model B. Switching embedding models means re-embedding the whole corpus, which is why the choice is expensive to revisit. 2. **Query encoding sits on the hot path.** Its latency and per-call cost are paid on every request, so a big encoder is much more affordable at index time than at query time. 3. **Both sides must use the same model and the same conventions** — same version, same normalization, same query/document prefix scheme if the model expects one. Mixing them silently degrades ranking rather than erroring. 4. **ANN adds its own approximation.** Even a perfect bi-encoder loses a little recall to the index's search parameters, so retrieval quality is a property of model *and* index together. ## What interviewers listen for The weak answer is "bi-encoders are faster" with no account of *why*. The strong answer names independence-and-precomputation as the mechanism, states the cost arithmetic, and then concedes the accuracy gap honestly instead of pretending the bi-encoder is simply better — because the concession is what motivates the two-stage pipeline every real RAG system ends up with.

  • If the bi-encoder is less accurate, why not just retrieve a larger candidate set and let the second stage sort it out?
    Because the second stage's cost is linear in candidate count, so widening from 50 to 5,000 candidates multiplies re-ranking latency and spend by a hundred while recall improves only marginally past a point. There is also a ceiling: a document the bi-encoder never surfaces can never be recovered downstream. The right move is usually to fix first-stage recall — better chunking, hybrid retrieval, a stronger encoder — not to brute-force a wider funnel.
  • What breaks if you embed documents with one model and queries with a different one?
    Nothing errors, which is the danger. The two models produce vectors in unrelated geometric spaces, so dot products between them are essentially noise and ranking collapses to near-random while the pipeline keeps returning results. It usually surfaces as a mysterious quality drop after a deploy. Guard against it by storing the model name and version in index metadata and asserting at query time that the query encoder matches.

saying these in an interview costs you the question

  • Says the bi-encoder is simply more accurate than pair scoring
  • Thinks the query and document are encoded together in a bi-encoder
  • Believes an ANN index can store pair-level relevance scores
  • Suggests running pair scoring across the entire corpus per query
  • Assumes vectors from different models are interchangeable

context