How would you choose between two embedding models for your own corpus?
answer
- public averages are not your distribution
- the shortlist comes from elsewhere than the decision
- a few hundred labelled real queries
- measure at the k you actually inject
- rank order flips on in-domain data
basics
~20 sBuild a few hundred query-and-relevant-chunk pairs from real usage, then measure recall at your actual retrieval depth and nDCG on your own corpus. Public leaderboards only narrow the shortlist; in-domain ranking frequently inverts the leaderboard order.
solid answer
~50 sTreat the leaderboard as a **candidate generator, not a decision**. Public suites such as MTEB average dozens of tasks over public corpora, so a model's rank there reflects a text distribution that is almost certainly not yours, and top entries are close enough that the gaps are within noise. The decision comes from an in-domain set: mine 300-500 real queries from logs or support tickets, label the chunk that actually answers each, and score every candidate with recall@k at the *k* you truly feed the model plus nDCG@10 for ordering quality. In practice the ranking inverts often — re-scoring MTEB shortlist candidates on a 500-query in-domain set routinely promotes a model that sat several places lower. Then apply the non-quality constraints that decide deployability: maximum sequence length against your chunk sizes, vector dimension against storage and index memory, query latency and cost at your request volume, multilingual coverage, licence and self-host feasibility, and how much re-embedding a future switch would cost.
go deeper
Know that embedding models differ in quality and that the right one has to be tested on your own data, not just picked from a public ranking.
Describe the actual comparison: a few hundred labelled query-chunk pairs from real usage, recall at your production k plus nDCG, everything else held constant between candidates.
Show how the eval set is built and sliced, explain why public benchmark ranks invert in-domain, and factor sequence length, dimension cost, latency and re-index cost into the decision.
Own the choice as a long-lived commitment: the index is pinned to this model, so set the bar for what future gain would justify a re-embed, and decide how much evaluation infrastructure the organization maintains permanently.
## Why the leaderboard cannot decide this Public embedding benchmarks — MTEB is the reference point, with BEIR as its retrieval heritage — are genuinely useful and genuinely insufficient. Four reasons: 1. **Distribution mismatch.** The tasks are built on public web, Wikipedia, scientific and forum text. If your corpus is insurance claim narratives, patent claims, incident postmortems or internal policy documents, its vocabulary, sentence shape and query style are unlike anything in the average. 2. **Averaging hides the axis you care about.** A headline score is a mean over many task types — classification, clustering, reranking, retrieval, summarization-adjacent tasks. A model can lead the average while trailing on asymmetric retrieval specifically. 3. **Contamination and optimization pressure.** Public test sets leak into training corpora, and a public leaderboard is a target; entries can be tuned toward it in ways that do not transfer. 4. **Compressed margins.** Near the top, models sit within a point or two of each other — inside the noise band of any evaluation you would run yourself. So use it for what it is good at: eliminating clearly weaker models and surfacing three to five plausible candidates. ## Building the in-domain set The artifact that actually decides the question is a labelled set from your own data. A workable recipe: - **Harvest real queries.** Search logs, support tickets, chat transcripts, the questions your existing assistant already gets. Real queries are short, misspelled, jargon-laden and ambiguous in ways synthetic ones are not. - **Aim for 300-500 queries.** Enough for the differences you care about to clear noise; small enough that a couple of people can label it in days. - **Label the relevant chunk(s).** For each query, mark which chunk from your indexed corpus genuinely answers it. Where a query has several acceptable answers, record them all — graded labels sharpen nDCG considerably. - **Stratify deliberately.** Include head queries and long-tail ones, queries hinging on an exact identifier, queries phrased in user vocabulary that never appears in the documents, and queries whose answer is genuinely absent so you can watch what a model retrieves when nothing is right. - **Keep a held-out slice.** If you later fine-tune, you will need a set that was never used to select or train. Synthetic queries generated from your chunks are an acceptable bootstrap when logs do not exist, but they carry a bias: they were written *from* the chunk, so they share its vocabulary and overstate every model's recall. Replace them with real queries as soon as you have them. ## The metrics that matter - **recall@k at your real k.** If you inject ten chunks into the prompt, recall@10 is the number that predicts whether the generator can possibly be right. Recall@100 flatters everything and tells you little. - **nDCG@10.** Rewards putting the right chunk near the top rather than merely somewhere in the window — this matters because context position affects how well the model uses a chunk. - **MRR** when there is exactly one right answer per query. - **Per-slice breakdowns.** An average can hide that one model is far better on identifier-bearing queries and worse on paraphrased ones. The slices tell you whether a hybrid or a re-scoring stage would close the gap more cheaply than switching encoders. Evaluate the *end state* fairly: same chunking, same prefix conventions per model, same ANN settings, same *k*. Changing two things at once makes the comparison meaningless. ## The constraints that are not quality A model that wins on nDCG can still be the wrong choice: - **Maximum sequence length** versus your chunk-size distribution — silent truncation is a quality bug disguised as a config detail. - **Vector dimension** drives index memory and storage across millions of chunks, and therefore infrastructure cost. - **Query-side latency and cost** are paid per request forever; index-side cost is paid once. - **Multilingual coverage** if any part of the corpus or query stream is not English. - **Licence, self-hosting and data residency** — a hosted API may be disqualified before quality is discussed. - **Stability of the provider or checkpoint.** A model that may be deprecated forces a re-index on someone else's schedule. - **Switching cost.** Because the index is pinned to the model, the *next* change means re-embedding everything. That makes the first choice worth doing carefully and makes marginal gains a poor reason to churn. ## Running the comparison Embed the corpus once per candidate (this is the expensive step — sample the corpus if full embedding is too costly, but keep the sample large enough that ANN behaviour is realistic), run the query set, and put the numbers in one table alongside dimension, latency and cost. Then make the call explicitly, and write down the in-domain score of the winner: it is the baseline every future change — a reranker, hybrid retrieval, fine-tuning — has to beat. This reflects practice as of mid-2026; the specific leaderboards move, the method does not.
- You have no query logs at all. How do you get an evaluation set off the ground?Bootstrap synthetically: sample chunks across the corpus and have a model write the question each chunk answers, then have a human filter out the unanswerable and the trivially lexical ones. Accept that this set overstates recall because the query was written from the chunk's own wording. Treat it as a relative comparator between candidates, ship behind it, and replace it with mined real queries within the first weeks of traffic.
- Two candidates land within a point of each other on your in-domain set. What decides it?Nothing about quality — the gap is inside your noise band with a few hundred queries. Decide on the operational axes: dimension and therefore index cost, maximum sequence length against your chunks, query latency and price at projected volume, licence and hosting constraints, and how likely the checkpoint is to be deprecated. If you want the quality gap resolved, enlarge or re-slice the eval set rather than trusting the point estimate.
- Why is recall@100 a misleading way to compare retrievers for a RAG system?Because you never inject a hundred chunks. At large k almost every reasonable model finds the answer somewhere, so the metric saturates and stops discriminating. What determines whether the generator can answer is whether the right chunk is inside the window you actually pass — and where in it. Measure recall at your production k and pair it with nDCG@10 so ordering counts.
saying these in an interview costs you the question
- Picks the top model on a public leaderboard and stops there
- Evaluates on synthetic queries written from the chunks themselves
- Reports recall@100 when the pipeline injects ten chunks
- Changes chunking and embedding model in the same comparison
- Ignores sequence length, dimension cost and licence constraints