skip to content

A marketplace blends keyword matches and embedding neighbours into one candidate list — why not normalise the two scores and add them?

level: seniorimportance: should knowfreq 46%

answer

  1. two sources, no shared scale
  2. unbounded against bounded
  3. per-query normalisation is scale-free
  4. merge on position, not magnitude
  5. margin discarded; scoring stage restores it

basics

~20 s

The two scores share no scale. A keyword score has no fixed ceiling and moves with query length and term rarity, while a similarity score is bounded but query-dependent, so per-query normalisation makes a weak list look exactly like a strong one. Rank fusion merges on position instead.

solid answer

~50 s

Keyword relevance scores such as BM25 are unbounded and their magnitude depends on the query itself — more terms and rarer terms produce larger numbers — so the same value means different things for two different queries. Embedding similarity is bounded, but the band that counts as "close" also shifts by query and by corpus region. Min-max normalising each source per query does not fix this: it rescales *every* list to fill the same range, so a source that returned nothing good gets the same top-of-range score as a source that returned an exact match. **Reciprocal rank fusion** sidesteps the problem by merging on the one thing the sources agree about — position — summing `1/(k + rank)` across sources. The trade is that you discard the margin inside each list: a runaway top hit fuses exactly like a narrow one, so the second-stage scoring model has to re-establish the real ordering with shared features.

code

pseudocode · 16 lines
pseudocode
K = 60
fused = empty_map(default = 0)

FOR each source IN [keyword_results, embedding_results]:
    FOR rank FROM 1 TO size(source):
        item = source[rank]
        fused[item] = fused[item] + 1 / (K + rank)

# worked example with K = 60
#   item A: rank 1 keyword, rank 30 embedding
#           1/61 + 1/90 = 0.01639 + 0.01111 = 0.02750
#   item B: rank 3 keyword, rank  3 embedding
#           1/63 + 1/63 = 0.01587 + 0.01587 = 0.03175
#   B outranks A: agreement beats one source's top hit

shortlist = take(sort_descending_by_value(fused), 400)

go deeper

for a junior

Remember that the two retrieval sources produce numbers on different scales, so their scores cannot simply be added together.

for a middle

Explain why per-query normalisation is scale-free and therefore hides how good each list actually was, and how rank fusion avoids the issue.

for a senior

Work the fusion arithmetic on a concrete example, name what the margin loss costs on head queries, and say where the scoring stage puts it back.

for a principal

Decide whether a trained blender is worth the extra model to operate, given how often either retrieval source is replaced and who owns it.

## Why the two candidate sources exist at all On a marketplace search surface, keyword retrieval and embedding retrieval fail in opposite directions. Keyword matching nails exact model numbers, part codes and brand-plus-attribute queries, and returns nothing at all when the shopper's wording misses the catalogue's wording. Embedding retrieval catches paraphrase and intent ("something warm for camping") and will happily return its nearest neighbours for a query that has no good answer anywhere. Running both and merging is the standard design because the union has much better recall than either alone. The merge is where the design gets interesting, because the two sources hand you numbers that are not comparable. ## What each score actually is | | keyword relevance score | embedding similarity score | |---|---|---| | range | unbounded above; no natural ceiling | bounded (a similarity measure) | | grows with | more query terms, rarer terms, term frequency in the item | closeness in the embedding space | | comparable across queries? | no — a two-term query and a six-term query live on different scales | not really — the useful band shifts by query and corpus region | | behaviour on a hopeless query | returns few or zero matches | still returns k nearest neighbours | A single fixed weight (`0.7 x keyword + 0.3 x similarity`) is therefore meaningless: the same weight behaves differently on every query. ## Why per-query normalisation does not rescue it The obvious patch is to min-max normalise each source's scores within the request, mapping both to `[0, 1]`, and then add. This quietly destroys exactly the information you needed: - Normalisation is **scale-free by construction**. The best item in each list becomes 1.0 whether it was an exact match or the least-bad of ten irrelevant neighbours. - It is **unstable in the tails**. One outlier at the top compresses everything below it, so the gap between ranks 2 and 3 changes meaning from query to query. - It **hides the empty case**. A keyword source that found two weak matches contributes two 1.0-ish scores, which is precisely the signal you wanted to *lose confidence* on. In short, normalising per query converts "how good was this list?" into "where in this list was this item?" — while pretending it is still a score. If you are going to end up with ranks, use ranks honestly. ## Reciprocal rank fusion, and what it costs Rank fusion assigns each item `1/(k + rank)` in each source that returned it and sums across sources; `k` is a damping constant, commonly around 60, that stops rank 1 dominating everything. Worked with `k = 60`: - item A: rank 1 in keyword, rank 30 in embedding — `1/61 + 1/90 = 0.01639 + 0.01111 = 0.02750` - item B: rank 3 in both — `1/63 + 1/63 = 0.01587 + 0.01587 = 0.03175` B wins. That is the property you are buying: **agreement across independent sources outranks one source's confident top hit.** It is a reasonable prior for candidate generation, where the goal is a high-recall shortlist rather than a final order. The costs are real and worth naming: 1. **Margin is discarded.** A source that found a perfect match and a source that found a mediocre one contribute identically at the same rank. 2. **The damping constant is a tuning knob** with no principled value; it decides how much a top-1 position is worth against agreement. 3. **Coverage skew.** A source that returns 1,000 candidates gets more chances to appear than one returning 100, so list lengths should be comparable or normalised. ## The alternative, and when it is worth it The principled alternative is to make the scores comparable rather than to abandon them: train a light model over shared features — each source's score, each source's rank, whether both sources returned the item, simple query statistics — against labels that exist on both surfaces. That recovers the margin rank fusion threw away, at the cost of a model to train, monitor and retrain whenever either source changes. Rank fusion needs none of that and degrades gracefully when a source is swapped out, which is why it is the usual default at the **candidate** stage. Either way, the merge is not the final ordering. Its job is to produce a shortlist with good recall; the heavy scoring stage then re-scores every survivor on one shared feature set, which is the only place in the funnel where a single consistent score genuinely exists.

  • What does the damping constant in reciprocal rank fusion control?
    How much a top position is worth relative to appearing in several sources. A small constant makes rank 1 dominate; a large one flattens the curve so agreement across sources matters more. It has no principled value, so it is tuned against the shortlist's recall at the size you actually pass to the scoring stage.
  • When is training a blending model worth it instead?
    When the margin you are discarding is costing you — typically on head queries where one source is confidently right and fusion buries it behind consensus mediocrity. A light model over each source's score, each source's rank and whether both returned the item recovers that, in exchange for a model you must monitor and retrain whenever either retrieval source changes.
  • One source returns 1,000 candidates and the other 100. What breaks?
    Coverage skew: the longer list has ten times as many chances to contribute a reciprocal-rank term, so it dominates the tail of the fused shortlist regardless of quality. Either cap the sources to comparable lengths before fusing, or weight each source's contribution explicitly rather than letting list length decide.

saying these in an interview costs you the question

  • Adding a keyword score and a similarity score with fixed weights
  • Believing min-max normalisation makes two sources comparable
  • Assuming a similarity score means the same thing for every query
  • Treating the fused order as the final ranking
  • Fusing sources of wildly different lengths without capping them