Why does dense vector retrieval always return k results when a keyword query can return none?
answer
- one retriever can return nothing, the other cannot
- matching is a predicate; distance is not
- there is always a nearest vector
- similarity has no calibrated zero
- empty results are information, not a bug
basics
~20 sNearest-neighbour search ranks by distance and has no match predicate, so it returns the k closest vectors however far away they are. Keyword retrieval intersects postings lists first, so query terms that appear nowhere yield an empty result set.
solid answer
~50 sThe two retrievers have different contracts. A lexical engine first computes a **match set** — the documents whose postings lists contain the query terms — and only ranks what survives; if nothing matches, the result is empty. A dense retriever has no match set. It asks "which k vectors are nearest to this one", and there is always a nearest vector, so a gibberish query still returns a full page of confidently-ordered garbage. That matters in hybrid retrieval: when the lexical leg correctly returns nothing, the fused list is pure vector output, and rank fusion will happily crown the least-bad of a bad set. The fix is not to trust the ordering but to add an explicit relevance gate — a similarity floor calibrated on your own corpus and embedding model, a knee/score-gap cutoff on the returned window, or a reranking gate — and to measure precision, not only recall.
go deeper
Be able to say that keyword search first finds documents containing the words, while vector search returns the closest vectors no matter how far away they are. Know that a vector search can return results for a query with no good answer.
Explain the mechanism: a match set built from postings lists versus a pure distance ranking with no predicate. Be ready to explain why similarity scores are not calibrated and why a fixed cutoff is fragile across queries and models.
Show how you would gate weak results in production — an empirically calibrated floor tied to the model version, a score-gap cutoff, or a rerank gate — and how you would measure precision so the failure is visible before users find it.
Own the product decision about what the system says when it does not know: honest zero-result handling versus labelled fallbacks, and the evaluation regime that keeps a model upgrade from silently invalidating every threshold in the stack.
## Two retrieval contracts Lexical and dense retrieval look interchangeable from the outside — text in, ranked documents out — but they answer different questions. A lexical engine answers **"which documents contain these terms, and how well?"** Retrieval is a two-stage act: build the candidate set by walking the postings lists of the query terms, then score the survivors with a model such as BM25. Matching is a predicate; scoring is a ranking over things that already passed the predicate. If none of the query's terms occur in the collection, the intersection is empty and the engine returns zero results. That empty result is *information*: it tells you the user's vocabulary is not in your index. A dense retriever answers **"which k stored vectors are closest to this query vector?"** There is no predicate. Every document has a position in the embedding space, every query has a position, and distances are always defined. Ask for ten neighbours of any point and you get ten, whether the nearest is nearly identical text or a completely unrelated document that happens to be the closest thing you own. ## Why there is no natural zero You might expect cosine similarity to supply the missing predicate: keep everything above 0.0, discard the rest. It does not work, for three reasons. First, embedding spaces are **anisotropic**. Trained encoders tend to place all their vectors in a narrow cone, so two completely unrelated texts routinely score 0.7–0.8 cosine. The observed similarity range is a small window near the top of the theoretical one, and where that window sits differs by model. Second, similarity is **not calibrated across queries**. A short query, a long query, a query in another language and a query full of rare proper nouns all produce different score distributions. A threshold that suppresses junk for one query suppresses everything for another. Third, the threshold is **tied to a specific model and corpus**. Swap the encoder — or re-embed with a newer version, or truncate dimensions — and every number shifts. A hard-coded cutoff silently becomes wrong at the moment of a model upgrade, which is exactly when nobody is looking at precision. ## What this does to hybrid retrieval In a hybrid system, the two legs disagree most sharply on the queries that have no good answer. Typing a phrase your corpus has never discussed yields an empty lexical list and a full dense list. Whatever fusion you use — rank-based or score-based — the fused ranking is then entirely dense, and its top entry is presented with exactly the same visual confidence as a perfect hit. Users read position as relevance. "No results" is an honest, actionable answer; "here are ten unrelated documents" is worse than nothing, and in a retrieval-augmented pipeline it becomes worse still, because a generator will try to answer from whatever it is handed. The symmetric failure exists too: dense retrieval's willingness to always answer is precisely what rescues the query whose words are absent but whose *meaning* is present. You want that behaviour for paraphrases; you do not want it for nonsense. The engine cannot tell the two apart from distance alone. ## Practical remedies **Calibrate a floor empirically.** Sample real queries, label the top results, and find the similarity value below which precision collapses — for *your* model and *your* corpus. Store it beside the model version and re-derive it on every model change. Treat it as a tuned parameter, not a constant. **Prefer relative cutoffs to absolute ones.** Look at the shape of the returned window rather than its absolute values: cut at the largest score gap (the "knee"), or drop everything more than some fraction below the top hit. These adapt per query, though they still assume the top hit is good. **Use the lexical leg as evidence.** If the query produced zero lexical matches and the dense top hit is not decisively close, that is a strong combined signal for "no results". Requiring some minimal lexical overlap for a result to be shown is crude but effective for identifier-shaped queries. **Gate with a stronger model.** A reranking stage that scores query and passage together produces far better-separated scores than bi-encoder distance, so a cutoff on rerank score is more trustworthy than one on retrieval distance. It costs a model call over the shortlist. **Measure the right metric.** Recall-oriented evaluation cannot see this failure — the dense leg never hurts recall by returning extra documents. You need precision at small k, and an explicit "should have returned nothing" class in your judgment set, or the regression is invisible until users complain.
- If cosine similarity is bounded, why can't you just reject everything below 0.5?Bounded is not calibrated. Trained encoders occupy a narrow cone of the space, so unrelated texts commonly sit at 0.7–0.8, and the useful range differs per model, per corpus and even per query length. Any workable floor is derived empirically from labelled samples for one model and re-derived when the model changes; a constant borrowed from another system is a coin flip.
- How would you detect this failure mode in production without a labelled set?Watch behavioural signals segmented by query. Queries where the lexical leg returned nothing but the system still showed results should show sharply worse click-through and higher abandonment or reformulation. Sampling that segment for manual review is cheap and finds the junk quickly. Zero-result rate falling to zero after a dense rollout is a warning sign, not a win.
- Is returning nothing ever worse than returning weak results?Sometimes. For exploratory or recovery-oriented interfaces, a few loosely related results plus a clear "no exact match — showing similar items" label beats a dead end. The decision is a product one, and it rests on labelling: presenting fallbacks honestly is fine, presenting them as matches is not.
saying these in an interview costs you the question
- Claims a cosine threshold cleanly separates relevant from irrelevant
- Thinks approximate search already drops low-similarity candidates
- Assumes zero results means the search engine is broken
- Reuses a similarity floor across different embedding models
- Evaluates dense retrieval on recall only, never precision