How do you normalize BM25 and cosine scores onto one scale before blending them?
answer
- one score is unbounded, the other is not
- the usual fix is also the usual bug
- every query's best hit becomes one
- one outlier stretches the denominator
- decide what a missing leg contributes
basics
~20 sMap each leg onto a comparable range before weighting: min-max or z-score over a fixed candidate window, or a corpus-calibrated bound. Per-query min-max is the usual choice and the usual bug, because it forces a 1.0 onto every query's best hit however bad it is.
solid answer
~60 sYou cannot add the two directly: a lexical score is unbounded and depends on term rarity, document length and corpus statistics, while cosine sits in a bounded range with its own compressed distribution. The standard move is **min-max over each leg's returned window** — `(s - min) / (max - min)` — then a weighted sum. It is easy and it is dangerous, because the normalization is per query: the best hit always becomes 1.0, so a query where nothing matched looks exactly as confident as a perfect one, and one outlier score squashes everything below it. **Z-score** over the window is more robust to a single outlier but is still window-relative. If you need scores comparable *across* queries, you need a fixed transform — a corpus-calibrated bound or a distribution fitted offline — not per-query rescaling. You also have to decide, explicitly, what a document found by only one leg gets for the missing leg. Given all this, rank fusion is the safer default unless you specifically need magnitude.
code
python · 12 linesdef minmax(scores):
lo, hi = min(scores.values()), max(scores.values())
span = hi - lo
if span == 0: # single candidate, or all tied
return {d: 1.0 for d in scores}
return {d: (s - lo) / span for d, s in scores.items()}
def blend(lexical, dense, alpha=0.5, fill=0.0):
lex, den = minmax(lexical), minmax(dense)
ids = set(lex) | set(den) # union, not intersection
return {d: alpha * den.get(d, fill) + (1 - alpha) * lex.get(d, fill)
for d in ids}go deeper
Know that a keyword relevance score and a cosine similarity live on different scales and cannot simply be added. Recognise min-max normalization as the usual way to bring two score lists into the same range.
Explain why a lexical score is unbounded and corpus-dependent while cosine is bounded and compressed, and walk through the min-max formula and its per-query nature. Be ready to name the missing-leg fill value as a decision you have to make.
Demonstrate that you have been burned by per-query min-max: manufactured confidence, outlier sensitivity, window dependence. Show how you would pin the window, pick a fill value deliberately, and tune the blend weight on a judgment set rather than by intuition.
Own the choice between rank fusion and score fusion as an architectural commitment: what the product needs a score to mean, what calibration and re-tuning that obliges the team to do on every model and corpus change, and who maintains it.
## Why the raw scores cannot be added The two legs produce numbers that share no axis. A BM25 score is a **sum of per-term contributions**, each shaped by inverse document frequency, term frequency saturation and document-length normalization. It has no upper bound in practice, it grows with the number of query terms, and its magnitude depends on corpus statistics — the same document, same query, in a different index, scores differently. A score of 12 means nothing on its own. A cosine similarity is **bounded** (in [-1, 1], and effectively in a narrow band near the top for most trained encoders). It does not grow with query length, and it carries no notion of term rarity. Add them and the lexical leg dominates by sheer magnitude on multi-term queries and is drowned on single-term ones — the weighting you *intended* is replaced by whatever the scales happen to be that day. ## Per-query min-max, and its failure The usual normalization is min-max over each leg's returned candidate window (say the top 100): ``` norm(s) = (s - min_window) / (max_window - min_window) ``` Both legs now land in [0, 1] and a weighted sum `α·norm(dense) + (1-α)·norm(lexical)` behaves as intended *within* one query. The failures are systematic: **It manufactures confidence.** The best result of every query becomes exactly 1.0. A query where the top hit is barely relevant is scored identically to one with a perfect match. Any downstream threshold, any "did we find anything?" logic, any confidence surfaced to a user or a generator, is now meaningless. **It is outlier-sensitive.** One dominant top score stretches the denominator and compresses every other document into the bottom of the range, so the second through hundredth results become near-indistinguishable — the ordering survives but the *weighting* against the other leg collapses. **It depends on the window.** Normalize over the top 10 and over the top 1000 and you get different numbers for the same document, so the fused order changes with a parameter that has nothing to do with relevance. The window must be fixed and identical between offline evaluation and production, or your tuned α does not transfer. **It breaks on short lists.** With two or three candidates, min and max are almost arbitrary; with one candidate the formula divides by zero. ## Z-score and other options **Z-score**: `(s - mean_window) / stdev_window`. More robust to a single outlier than min-max, and it preserves relative spacing better, but it is still computed per query over a window, so it also carries no cross-query meaning, and it produces negative values that a naive weighted sum handles badly. **Theoretical or calibrated bounds**: lexical scoring does admit per-term score upper bounds — the same bounds that dynamic pruning algorithms use to skip postings — so a query-dependent maximum can be computed rather than observed. Dividing by that gives a value with a stable interpretation ("how close to the best achievable score for this query"). It is more work and less common, but it is the honest way to get a number that means something. **Distribution fitting**: the classic information-retrieval approach models the score distribution of relevant and non-relevant documents offline and maps raw scores to an estimated probability of relevance. Done properly this is the only normalization that is genuinely comparable across queries, and it needs labelled data and periodic refitting. ## The missing-leg problem A document that the lexical leg found and the dense leg did not has no dense score. This is not a corner case — it is most of the union. You must choose a fill value, and every choice biases the result: - **Fill with the normalized floor (0)**: treats "not in the top 100" as "maximally dissimilar", which over-penalizes a document that was at rank 101. - **Fill with the window's minimum score**: gentler, but still window-dependent. - **Re-score the document with the missing leg**: exact, and usually affordable only for a small union because it means computing a similarity per document. - **Keep only the intersection**: destroys recall, which is the entire reason you built a hybrid system. Rank fusion sidesteps this cleanly — absence simply contributes nothing. ## When to normalize anyway Despite all of the above, score fusion beats rank fusion in specific situations: when you need a **relevance threshold** (rank fusion cannot give you one), when one leg's *magnitude* carries real information you want to preserve (an exact identifier match should not be one rank step ahead of a fuzzy one — it should be far ahead), or when you want a continuous confidence value for downstream logic. The discipline that makes it work: fix the candidate window and use the same one everywhere; prefer a calibrated or theoretical bound over per-query rescaling if the score must mean something across queries; choose the fill value deliberately and document it; tune the blend weight on a judgment set rather than by eye; and re-tune after any change to the embedding model, the analyzer or the corpus, because every one of those moves the distributions you calibrated against.
- When is score normalization worth the trouble over rank fusion?When you need magnitude. Rank fusion cannot give you a relevance threshold, cannot express that an exact identifier match is far better than the runner-up rather than one step better, and cannot produce a confidence value for downstream logic. If any of those matter, normalize — and accept the calibration work, the fixed window, the fill-value decision and the re-tuning after every model or corpus change.
- Why does a tuned blend weight stop working after an embedding model upgrade?Because the weight was tuned against a specific score distribution. A new encoder shifts the mean and spread of similarities, so the same α now gives the dense leg a different effective influence, and any absolute thresholds calibrated alongside it are invalid too. Treat the blend weight and any thresholds as artifacts of the model version, re-derive them on a judgment set, and gate the rollout on offline metrics before shipping.
- What goes wrong if you normalize over the top 10 offline and the top 200 in production?The min and max differ, so identical raw scores normalize to different values and the fused ordering diverges from what you evaluated. The blend weight you tuned offline no longer corresponds to the influence each leg actually has. The window size is part of the fusion configuration and must be pinned identically in both places.
saying these in an interview costs you the question
- Adds a lexical score and a cosine similarity together directly
- Assumes min-max normalized scores are comparable across queries
- Ignores what a document missing from one leg scores
- Tunes the blend weight by eye instead of on a judgment set
- Thinks bounded cosine values need no normalization at all