skip to content

A Weaviate hybrid query ranks an exact keyword match below a fuzzy one — how do you debug it?

level: seniorimportance: should knowfreq 40%

answer

  1. ask the engine before tuning
  2. two metadata fields, one call
  3. bisect with the alpha endpoints
  4. did the keyword half match at all
  5. coverage and tokenization before weights

basics

~20 s

Ask Weaviate to explain itself: request score and explain_score metadata on the hybrid call to see what each half contributed per object. Then run the query at alpha=0 and alpha=1 separately to find out which half is misbehaving before changing any tuning parameter.

solid answer

~40 s

Start with evidence, not tuning. Add `return_metadata=MetadataQuery(score=True, explain_score=True)` to the `hybrid()` call; `object.metadata.explain_score` returns a breakdown of what the keyword and vector halves contributed to that object, which usually settles the question immediately. Then bisect: run the same query at `alpha=0` and at `alpha=1`. If the exact match is absent from the `alpha=0` result, the keyword half never matched it — look at `query_properties` coverage, whether the property is searchable, and tokenization of the token you are searching. If it ranks first at `alpha=0` but loses in the blend, the issue is weighting or fusion: a vector-leaning alpha is drowning it, or a rank-based fusion strategy has flattened its lead into a rank-1 contribution no larger than the fuzzy match's. Only after locating the failing stage should you touch alpha, boosts or `fusion_type`.

code

python · 20 lines
python
import weaviate
from weaviate.classes.query import MetadataQuery

client = weaviate.connect_to_local()
docs = client.collections.get("Doc")
q = "ERR4471 pool timeout"

for alpha in (0.0, 1.0, 0.7):
    res = docs.query.hybrid(
        query=q,
        alpha=alpha,
        limit=5,
        return_metadata=MetadataQuery(score=True, explain_score=True),
    )
    print(f"--- alpha={alpha}")
    for obj in res.objects:
        print(obj.properties["title"], obj.metadata.score)
        print(obj.metadata.explain_score)

client.close()

go deeper

for a junior

Know that Weaviate can report a per-result score and an explanation of how it was reached, and that you can run the query keyword-only or vector-only to see each half separately.

for a middle

Be able to name the metadata you request and walk the bisect: alpha=0 shows the keyword ranking, alpha=1 the vector ranking, and the difference tells you which stage went wrong.

for a senior

Show the full diagnosis discipline — evidence before tuning, coverage and tokenization before weights, and a fix applied at the layer that actually failed rather than at whichever knob is nearest.

for a principal

Own the regression risk: hybrid tuning trades query classes against each other, so the deliverable is a labelled evaluation set and a per-class routing policy, not a new global constant.

## Why this symptom is confusing A hybrid query has several independent places where an exact keyword match can lose, and they produce nearly identical output: results come back, nothing errors, and the wrong document is on top. Guessing at alpha is the reflex and it is usually wrong. The efficient path is to identify **which stage failed** before adjusting anything. The stages, in order: 1. **Candidate generation, keyword half** — did BM25 match the object at all? 2. **Candidate generation, vector half** — did the vector search return it? 3. **Weighting** — how much did alpha scale each half's contribution? 4. **Fusion** — how did the chosen strategy combine the two into one order? ## Step one: make the engine explain itself Request the metadata explicitly: `return_metadata=MetadataQuery(score=True, explain_score=True)` `object.metadata.score` gives the fused score, and `object.metadata.explain_score` gives a human-readable string describing what each half contributed to that object. Read it for both the document that won and the document you expected to win. The two most common readings are decisive: - The expected document shows **no keyword contribution at all** → stage 1 failed; this is a coverage or tokenization problem, and no tuning will fix it. - The expected document shows a **healthy keyword contribution that was outweighed** → stage 3 or 4; this is a tuning problem. ## Step two: bisect with the alpha endpoints Run the identical query three times: at `alpha=0` (keyword only), at `alpha=1` (vector only), and at your production alpha. This isolates each retrieval. **Exact match missing at alpha=0.** The keyword half genuinely does not find it. Check, in order: - **Property coverage.** If the call passes `query_properties`, that list replaces the default set — a property left off contributes no keyword matches. Re-run with the argument omitted to test. - **Searchability.** A text property configured without a searchable index cannot participate in BM25, however you query it. - **Tokenization.** The query token and the indexed token must actually be the same token. Identifiers with hyphens, dots, underscores, casing differences or embedded punctuation are the usual failure: `ERR-4471` may index as two tokens, or as one, and your query may do the opposite. Test with a single bare token you are certain appears. - **Property type.** Numeric, boolean and date properties are matched by filters, not by keyword scoring. **Exact match ranks first at alpha=0 but loses in the blend.** The keyword half worked and the blend overrode it: - **Alpha is too vector-leaning** for this query class. The v4 client default of 0.7 is deliberate but is not right for identifier lookups. Lower it, or route this query class to a lower alpha. - **Fusion strategy is flattening the lead.** A rank-based fusion strategy discards magnitude: being overwhelmingly the best keyword match and being narrowly the best produce the same rank-1 contribution, and a document present in *both* lists at middling positions can overtake a document that dominates only one. If your query classes are of the "one document is obviously right" kind, a score-normalizing fusion strategy preserves that lead better. - **A boost is not doing what you think.** `query_properties` boosts change keyword *scores*; under rank-based fusion, a boost that does not move the keyword *rank* changes nothing in the final order. **Exact match missing at alpha=1 as well.** Expected and harmless on its own — rare literal tokens are exactly what embeddings represent poorly. It becomes a problem only combined with a high alpha, which is the argument for not running a vector-leaning alpha on identifier traffic. ## Step three: fix at the right layer Map the finding to the fix: - Coverage or searchability → change the query's property list, or the collection's property configuration and reindex. - Tokenization → change the property's tokenization and reindex, or normalize the query text in your application. - Weighting → route this query class to a lower alpha rather than moving the global default and regressing natural-language traffic. - Fusion → set `fusion_type` explicitly and re-measure; do not leave it implicit, since the server default is a version property and can change under an upgrade. ## Guard the fix One fixed query is an anecdote. Capture the failing query and a handful of its neighbours into a small labelled set, and re-run it after each change — hybrid tuning trades one query class against another almost every time, and without a set you will discover the regression from users. Do not encode the fix as an absolute score threshold: fused scores are not comparable across alpha values, fusion strategies or result-set sizes, so a threshold that works today silently breaks at the next tuning change.

  • The keyword half finds nothing even when you search a bare token that is visibly in the document. What now?
    Tokenization or searchability, not weighting. Confirm the property is a text property with a searchable index — without one it cannot participate in BM25 regardless of the query. Then compare how the property's tokenizer split the stored text against how the query is tokenized; punctuation, casing and hyphenation are the usual divergence. The fix is a property configuration change plus reindexing, or normalizing the query text before sending it.
  • Why is lowering the global alpha often the wrong response to this symptom?
    Because alpha is a single knob shared by query classes with opposite needs. Dropping it to rescue identifier lookups degrades the natural-language traffic that motivated hybrid search in the first place. Since alpha is a per-request argument, the better fix is to classify the query in your retrieval layer and pass a low alpha for token-heavy queries while leaving conversational queries vector-leaning.
  • How would you keep a hybrid tuning change from regressing other queries?
    Keep a small labelled evaluation set covering each query class you serve — identifier lookups, short phrases, natural-language questions — and re-run it on every alpha, boost or fusion change, scoring on ranking metrics rather than raw score values. Hybrid tuning almost always trades one class against another, and without the set the regression surfaces as user complaints instead of a failing check.

saying these in an interview costs you the question

  • Reaching for alpha first without checking whether the keyword half matched
  • Assuming the fused score alone explains the ordering
  • Believing a boost fixes a property that was never searched
  • Treating a missing exact match under pure vector search as a bug
  • Encoding the fix as an absolute score cutoff

context