Why can an MS MARCO-trained cross-encoder underperform your bi-encoder on a specialist corpus?
answer
- relevance is learned from a distribution
- web queries versus clinical shorthand
- reordering can only hurt precision
- always keep an unreranked baseline
- mine negatives from your own retriever
basics
~20 sA reranker learns a notion of relevance from its training distribution. MS MARCO is short, general web-search queries over web prose, so on a specialist corpus with unfamiliar terminology and differently-shaped queries the reranker can confidently reorder good candidates into a worse order.
solid answer
~50 sPublic rerankers are trained on large web-search relevance data, most commonly MS MARCO, whose queries are short general-purpose questions and whose passages are ordinary web text. When your corpus is a veterinary formulary — dosage tables, drug abbreviations, contraindication clauses — and your users ask in clinical shorthand, both sides of the pair are out of distribution. The reranker still emits confident scores; they are just less trustworthy, and because reranking *reorders*, a wrong opinion actively demotes passages the retriever had ranked correctly. Detect it by always evaluating reranked versus unreranked order on a labelled domain set, and look at per-query wins and losses rather than only the average, since a small mean gain can hide a tail of badly damaged queries. If it loses, the options are fine-tuning the reranker on domain pairs with hard negatives mined from your own retriever, using a smaller reranker as a lighter touch, or simply not reranking — an honest and underused answer.
go deeper
Know that a reranker is a trained model, so it works best on text that resembles what it was trained on, and that reranking only reorders the candidates it is given.
Explain why general web-search training transfers poorly to a specialist corpus — unfamiliar terminology, differently-shaped queries and passages, a different relevance criterion — and why that can leave results worse than not reranking.
Demonstrate the evaluation discipline: unreranked baseline retained permanently, per-query rank deltas and sliced results rather than a single average, and a concrete fine-tuning plan built on hard negatives mined from your own retriever.
Own the buy-versus-own decision. A fine-tuned reranker is a model you must version, evaluate and retrain as the corpus drifts; weigh that permanent cost against the measured gain, and be willing to ship no reranker at all.
## Relevance is learned, not universal A cross-encoder does not compute relevance from first principles. It was trained on a dataset of queries, passages and relevance judgements, and what it learned is *that dataset's* notion of which passage answers which query. The dominant public training set for passage rerankers is MS MARCO, built from real Bing queries with human-judged passages. Everything about it shapes the resulting model: queries are short natural-language questions, passages are general web prose of a few sentences, and relevance means roughly "this paragraph contains the answer a general reader wanted". That is an excellent prior for a general question-answering assistant over web-like content. It is a much weaker prior for a corpus that looks nothing like the web. ## What goes wrong out of domain Take a veterinary formulary. Passages are dense, semi-structured entries: drug name, species, dosage per kilogram, route, contraindications, often as tables flattened into text. Queries are clinical shorthand — abbreviations, species plus drug plus a numeric question — not fluent web questions. Several failure paths open at once: - **Vocabulary drift.** Terms that carry the entire meaning (a drug abbreviation, a species code) are rare or absent in the training distribution, so the model has weak representations for exactly the tokens that decide relevance. - **Passage shape.** The model learned to reward flowing prose that states an answer. A dosage table that *is* the answer may look, to it, like a low-quality passage. - **Query shape.** Terse keyword-ish queries are unlike the natural-language questions the model saw, so its query encoding is less reliable. - **A different relevance criterion.** In many specialist settings, relevance means "applies to this exact species/jurisdiction/version", a constraint the general model never learned to enforce and will happily ignore. Meanwhile the bi-encoder in your first stage may have been adapted to the domain, or may simply be doing lexical-adjacent matching that happens to work. So you get the counterintuitive result: the strictly more powerful architecture makes the results worse. ## The damage is asymmetric A reranker only ever *reorders*, so a bad reranker cannot improve recall but can absolutely destroy precision at the top. If it demotes the single correct passage from rank 1 to rank 8 and you pass three passages to the generator, the answer is now unsupported — a worse outcome than not reranking at all. This is why "we added a reranker and quality went down" is a real and common report, not a misconfiguration story. ## How to find out before your users do Make unreranked order a permanent baseline in the eval, not a one-off check. Score a labelled query set both ways at the depth you actually pass to the generator. Two things to insist on: 1. **Per-query deltas, not just the mean.** A +2 point average can consist of many small wins and a handful of catastrophic losses, and the losses are what users notice. Count queries where the correct passage's rank got worse. 2. **Slices.** Break the eval by query type, by document family, by whether the query contains domain abbreviations. Out-of-domain damage is rarely uniform; it concentrates in the slices furthest from web prose. Also watch the score distribution. Reranker scores that all cluster tightly, or that rank near-duplicates above the one distinguishing passage, are a signal the model has no real opinion on your text. ## The repair options, roughly in order of cost - **Skip reranking.** If the first stage is already domain-adapted and the reranker does not beat it, the cheapest, fastest and most reliable configuration is the one without it. Say this out loud in an interview; candidates rarely do. - **Try a different checkpoint.** Multilingual and newer general rerankers differ substantially in how they transfer. This is a cheap experiment with a real chance of working. - **Fine-tune on domain pairs.** The high-value move. Collect queries (from logs, from support tickets, or synthesized from documents by an LLM and then human-checked) with a known relevant passage, and mine **hard negatives** from your own retriever's top results — passages that scored well but are wrong. Training against those teaches the distinctions your corpus actually requires. Random negatives teach almost nothing, because they are trivially separable. - **Rebuild the label set as you learn.** Domain relevance criteria surface during labelling: whoever adjudicates will discover that species, dosage form or document version is decisive. Encode that in the labelling guidance and the fine-tuning data. A fine-tuned reranker also carries an ongoing cost: it is now a model you own, version, evaluate and retrain when the corpus shifts. That is a legitimate reason to prefer the no-reranker configuration when the gain is small.
- You decide to fine-tune the reranker. Where do the negative examples come from?From your own first-stage retriever's top results — passages it ranked highly for a query that are nonetheless wrong. These hard negatives teach exactly the distinctions the deployed system gets wrong. Random passages from the corpus are trivially separable and teach the model almost nothing. Refresh the negative pool after retraining, since the errors the retriever makes will have shifted.
- Your reranked results show a +2 point average improvement. Why is that not enough to ship?Because the mean hides distribution. A small average gain often consists of many marginal wins plus a tail of queries where the correct passage was demoted out of the window passed to the generator, producing unsupported answers users notice. Report per-query rank deltas and the count of regressions, and slice by query type — out-of-domain damage concentrates in the slices least like the reranker's training data.
- How would you generate a labelled evaluation set for a specialist corpus with no query logs?Synthesize candidate queries from documents with an LLM, then have a domain expert adjudicate which passage genuinely answers each one and discard the ambiguous items. Deliberately include the query shapes your users will actually type, including abbreviations and terse phrasing. Keep the set small but adjudicated rather than large and noisy, and treat the expert's disagreements as documentation of what relevance means in this domain.
saying these in an interview costs you the question
- Assumes a public reranker transfers to any corpus unchanged
- Reports only average ranking gains and ignores per-query regressions
- Thinks reranking can never make results worse because it adds information
- Fine-tunes on randomly sampled negatives instead of retriever-mined hard negatives
- Treats reranker scores as calibrated relevance probabilities across domains