When should a RAG reranker return nothing, and how do you pick that score cutoff?
answer
- top-m can never say "nothing here"
- relative scores rank, absolute scores decide
- two query distributions, one overlap region
- the operating point is a product choice
- re-calibrate whenever the model changes
basics
~20 sReturn nothing when the best candidate's absolute relevance score falls below a calibrated floor — the corpus simply has no answer. Set the floor from labelled in-scope and out-of-scope queries, choosing the point that meets your tolerated false-answer rate.
solid answer
~50 sTop-m selection always returns m passages, even for a question the corpus cannot answer, which is how a grounded system ends up citing the least-bad chunk to justify an invention. The fix is an **absolute score threshold** applied after reranking: if the top candidate scores below a floor, return an empty set and let the application say it does not know. Calibrate the floor empirically — score a labelled mix of in-scope and deliberately out-of-scope queries, plot the two score distributions, and pick the cutoff that meets your tolerated rate of confidently wrong answers. That is a product decision, not a modelling one: a clinical or legal assistant tolerates far more abstentions than a brainstorming tool. Thresholds are model-specific and drift, so re-calibrate on every reranker or embedding change and monitor the abstention rate in production.
go deeper
Know that a RAG system can return no passages at all, and that answering from the least-bad chunk is worse than saying the corpus has nothing. Be able to say a score floor is what makes that possible.
Explain why raw similarity scores are relative rather than absolute, distinguish a relative cutoff from a calibrated floor, and describe fitting the floor from in-scope and out-of-scope query distributions.
Show you have operated one: pick the operating point from the tolerated false-answer rate, re-calibrate on model swaps, monitor abstention rate as a regression signal, and mine abstained queries as content gaps.
Own abstention as a policy: decide the acceptable rate of confidently wrong answers per product surface, make the threshold an auditable artifact tied to the model version, and design the user-facing behaviour when the system declines to answer.
## The problem with top-m A selection stage that always emits its best m candidates has no way to express "there is nothing here". Ask a corporate policy assistant about a policy that does not exist and it still hands the model five chunks — the five least-irrelevant paragraphs in the corpus. The model, primed to answer from provided context, then produces a fluent, cited answer to a question its knowledge base cannot support. Citations make it worse, not better: they make the fabrication look verified. Abstention is the fix, and it is a legitimate result. "I don't have anything on that" is a correct output for an out-of-scope question. ## Relative versus absolute scores The reason this is subtle is that most ranking machinery is *relative*. Similarity scores and reranker outputs are meaningful for ordering candidates within one query; they are not automatically comparable across queries. A cosine of 0.62 might be a strong match for one query phrasing and mediocre for another. So a naive cutoff ("drop anything under 0.7") behaves differently for short queries, long queries, and different languages. That gives you three practical options: 1. **Calibrated absolute threshold.** Convert reranker output into something comparable across queries — many cross-encoders emit a logit you can pass through a sigmoid, and you can fit a calibration on labelled data — then set one floor. This is the standard approach and the one interviewers expect you to describe. 2. **Relative cutoff.** Keep candidates within some fraction of the top score, or cut where the largest score gap appears. This adapts to per-query scale but cannot detect the case where *everything* is bad, since the top score is always 100% of itself. Relative rules control set size; they do not implement abstention. 3. **A separate decision.** Ask a cheap classifier or a small model "does this passage actually answer this question?" as a final gate. More robust, more expensive, and its own thing to evaluate. Use a relative rule for how many to keep and an absolute rule for whether to keep any. They solve different problems and a strong answer says so. ## Calibrating the floor The procedure is mechanical: - Assemble two query sets: **in-scope** questions the corpus genuinely answers, and **out-of-scope** questions it genuinely does not — plausible-sounding neighbours work best (a question about a policy the company almost has), not absurdities. - Run the full cascade and record the top reranked score for each query. - Plot the two distributions. They overlap; the overlap is your unavoidable error region. - Sweep the threshold across the overlap and produce a curve: at each cutoff, what fraction of out-of-scope queries are wrongly answered, and what fraction of in-scope queries are wrongly refused? - Choose the point matching the tolerated error, then verify end to end — a threshold that looks right on scores can still be wrong once the generator's own hedging is in the loop. The choice of point is a product decision. A clinical decision-support tool or a legal-research console should sit far toward abstention: refusing to answer is cheap, answering wrongly is not. A brainstorming or exploratory-search tool can sit far the other way, because a mediocre passage still gives the user something to react to. Ask which one you are building before quoting a number. ## Operating a threshold A calibrated cutoff is a piece of tuned state, and it rots. - **It is model-specific.** Swap the reranker, the embedding model, or even the chunk size and the score distribution moves. Re-calibrate as part of the change; never carry a threshold across models. - **It drifts with the corpus.** As documents are added, previously out-of-scope questions become answerable and the in-scope distribution shifts. - **Monitor the abstention rate** as a first-class metric. A sudden climb usually means an ingestion or embedding regression, not a change in user behaviour; a sudden fall can mean the threshold is now effectively disabled. Sample abstained queries for review — they are the cheapest source of "content we should add". - **Log the top score** with every request so you can re-tune from production traffic rather than a stale offline set. ## What the product does with nothing Abstention only pays if the surrounding application handles it well. Returning an empty context and letting the model improvise defeats the purpose — the prompt must instruct the model to say it has no source, or the application should short-circuit the model entirely and render a fixed message. Good behaviours: state plainly that the corpus has no coverage, show the closest near-misses explicitly labelled as "related, not an answer", offer to widen the search or route to a human, and capture the query as a content gap. Bad behaviour: silence, or an unlabelled answer built from the near-misses. ## The interview-ready shape Name the failure (top-m cannot express "nothing"), distinguish relative cutoffs from absolute thresholds, describe the two-distribution calibration with a stated error tolerance, tie the operating point to the product's cost of being wrong, and close on re-calibration and abstention-rate monitoring.
- Why is "keep everything within 90% of the top score" not enough to implement abstention?Because it is relative to the best candidate, which is always 100% of itself. If every candidate is irrelevant, the rule still keeps the top one and any close companions. Relative cutoffs are good at controlling set size and trimming a long tail after a score cliff, but only an absolute, calibrated floor can express that the whole pool is bad.
- You swap the cross-encoder for a newer one and end-to-end quality drops sharply. What would you check first?The threshold. Reranker scores are not on a shared scale across models, so a floor calibrated for the old one may now reject almost everything — the symptom is a spike in abstentions — or accept almost everything, producing confident wrong answers. Re-run the two-distribution calibration as part of the model swap and treat the threshold as part of the model artifact.
- How would you choose the operating point differently for a clinical assistant versus an internal brainstorming tool?By the asymmetry of the costs. For clinical guidance a wrong grounded answer can cause harm and a refusal costs a little time, so set the floor high and accept many false refusals. For brainstorming, a weak passage is still a useful prompt for a human and a refusal is pure friction, so set it low. Same machinery, opposite tuning, driven by the product not the model.
saying these in an interview costs you the question
- Assuming a similarity score is comparable across different queries
- Using a single hardcoded cutoff copied from a blog post
- Carrying a threshold across a reranker or embedding-model change
- Treating a refusal to answer as a system failure rather than a valid result
- Returning an empty context but letting the model answer anyway