How do Weaviate's HybridFusion.RANKED and RELATIVE_SCORE differ in hybrid results?
answer
- positions versus magnitudes
- one strategy throws scores away
- reciprocal rank, damped near the top
- normalization is relative to the candidate set
- the default changed at 1.24
basics
~20 sHybridFusion.RANKED merges the keyword and vector lists using only each object's position, so score magnitude is discarded. HybridFusion.RELATIVE_SCORE normalizes each list's raw scores into a comparable range first, so how far ahead a hit is changes the final order.
solid answer
~50 sBoth fusion strategies merge the keyword list and the vector list, but they read different signals. `HybridFusion.RANKED` is reciprocal-rank fusion: an object contributes roughly `1 / (60 + rank)` from each list it appears in, so only positions matter — the runaway top keyword hit and a barely-better-than-second one are treated almost identically. `HybridFusion.RELATIVE_SCORE` normalizes each list's raw scores against that list's own maximum and minimum, then adds the two normalized values weighted by `alpha`, so **score gaps survive** into the final ordering. Relative-score fusion is the server default from Weaviate 1.24 onward. The practical consequence is that ranked fusion is stable and forgiving of one badly-calibrated half, while relative-score fusion is more responsive but sensitive to the candidate set — an object matched by only one half gets zero from the other, and changing `limit` can shift normalization and therefore order.
code
python · 19 linesimport weaviate
from weaviate.classes.query import HybridFusion, MetadataQuery
client = weaviate.connect_to_local()
articles = client.collections.get("Article")
for fusion in (HybridFusion.RANKED, HybridFusion.RELATIVE_SCORE):
res = articles.query.hybrid(
query="index compaction stalls",
alpha=0.5,
fusion_type=fusion,
limit=5,
return_metadata=MetadataQuery(score=True, explain_score=True),
)
print(fusion)
for obj in res.objects:
print(" ", obj.properties["title"], obj.metadata.score)
client.close()go deeper
Know that Weaviate offers two ways to merge the keyword and vector results and that one uses positions while the other uses normalized scores. Naming the two options and the direction of the difference is enough here.
Explain the mechanics: reciprocal-rank contributions versus per-list min/max normalization, and give the observable consequence — magnitudes are discarded in one and preserved in the other.
Demonstrate that you know relative-score ordering depends on the candidate set, so results can shift after an ingest or a limit change, and that you would set fusion_type explicitly rather than inherit a server default across upgrades.
Own the evaluation stance: a fusion change is a ranking change that must be measured on a labelled set, and any downstream score threshold is a liability because the fused score has no stable meaning across strategies.
## The problem fusion solves A hybrid query produces two ranked lists whose scores are not on the same scale. BM25 scores are unbounded positive numbers that depend on term frequency, document length and corpus statistics — a rare-term match can score 14, a common-term match 0.9. Vector scores live on a completely different scale set by the similarity metric. You cannot simply add them. The fusion strategy is the rule for making them comparable, and Weaviate's `fusion_type` argument on `hybrid()` selects between two rules. ## HybridFusion.RANKED — reciprocal rank fusion Ranked fusion throws the raw scores away and keeps only **positions**. Each object's contribution from a list is a decreasing function of its rank in that list — Weaviate uses the reciprocal-rank form `1 / (60 + rank)` — and the contributions from the keyword list and the vector list are summed, weighted by alpha. Properties that follow directly: - **Scale-free.** Because no raw score is read, it does not matter that BM25 and vector similarity are numerically incomparable. There is nothing to calibrate. - **Magnitude-blind.** A document that is overwhelmingly the best keyword match and one that scraped into first place produce the same first-place contribution. Confidence is invisible. - **Flat near the top.** The gaps between the contributions of ranks 1, 2 and 3 are small (the constant 60 deliberately damps them), so a document that appears in *both* lists at middling positions frequently beats a document that is first in one list and absent from the other. That "consensus wins" behaviour is the strategy's main personality. - **Robust.** If one half is badly calibrated or noisy, ranked fusion limits the damage — the bad half can only contribute rank positions, not an inflated score. ## HybridFusion.RELATIVE_SCORE — normalized score fusion Relative-score fusion keeps the magnitudes. Within each list it normalizes scores against that list's own range — the top result becomes 1, the bottom becomes 0, everything else lands proportionally between — and then combines the two normalized values with the alpha weighting. Properties that follow: - **Confidence survives.** A keyword hit that is dramatically better than everything else keeps a large lead in the fused ordering, instead of being flattened to "rank 1". - **An object missing from one list scores zero on that side.** It is not penalized beyond that, but it also gets no credit, so single-list hits need a strong showing in their own list to compete with dual-list hits. - **Order depends on the candidate set.** Normalization is relative to the results actually retrieved. Change `limit`, add a filter that removes the previous top hit, or ingest a new document, and the min/max of a list moves — which can reorder objects whose raw scores did not change at all. This is the single most surprising behaviour in practice, and it is why the same query can reorder after an unrelated ingest. - **Sensitive to a mis-calibrated half.** If your vector scores are all squeezed into a narrow band, normalization stretches that band to the full 0–1 range and amplifies differences that are effectively noise. ## Which to choose Relative-score fusion is the server default from Weaviate 1.24, and it is usually the better starting point: it uses more information, and for most corpora the strong-signal-should-win behaviour matches human expectations. Reach for ranked fusion when: - One of your two halves is unreliable or poorly calibrated and you want its influence capped. - Your evaluation shows relative-score results churning as the corpus changes and you value stability over peak quality. - You want the consensus effect — documents that both halves like should beat documents only one half loves. Set it explicitly rather than relying on the default when the choice matters, because the default is a server-version property and a cluster upgrade can change your ranking underneath you. ## Consequences for the rest of your system The fused score is a *strategy-dependent* number. Its meaning, and its range, differ between the two modes and change with the candidate set under relative-score fusion. So: - Never hard-code an absolute score threshold and expect it to hold across a fusion change. - Do not compare scores between two different queries, even under the same fusion mode. - When you A/B a fusion change, evaluate on ranking metrics over a labelled set, not on score values. ## Debugging the choice Request `return_metadata=MetadataQuery(score=True, explain_score=True)` on the hybrid call. `object.metadata.explain_score` returns a human-readable breakdown of what each half contributed to that object, which is the fastest way to answer "why did this rank here" — in particular whether a surprising result came from one half only, or from the fusion arithmetic doing exactly what you asked.
- Under relative-score fusion, why can the same query reorder after an unrelated document is ingested?Because normalization is relative to the retrieved candidate list, not to a fixed scale. The top and bottom scores in each list define the 0–1 range, so a new document that enters a list and moves its maximum rescales every other object's normalized score. Two objects whose raw scores never changed can swap places. Ranked fusion is immune to this, since it only reads positions.
- What happens to an object that the keyword half finds but the vector half does not return at all?Under relative-score fusion it gets a zero contribution from the vector side and must win on its normalized keyword score alone, so it competes at a disadvantage against objects present in both lists. Under ranked fusion it likewise contributes only its keyword rank term, and because rank contributions near the top are close together, a dual-list object at middling positions often overtakes it.
- Would you hard-code a minimum hybrid score to drop weak results?Not as an absolute constant. The fused score's scale depends on the fusion strategy and, under relative-score fusion, on the candidate set — so a threshold tuned today breaks after a fusion change, a limit change, or a corpus shift. If you need a cutoff, derive it relative to the top result within the same response, or tune retrieval so the result count itself is the control.
saying these in an interview costs you the question
- Saying both fusion modes just add the raw BM25 and vector scores together
- Believing ranked fusion preserves how much better the top hit was
- Assuming relative-score normalization uses a fixed global scale
- Treating the fused score as stable across fusion modes or limits
- Thinking fusion type and alpha are two names for the same knob