When one vector store returns a distance and another a similarity score, what bug follows?
answer
- which direction counts as better?
- one backend reports lower-is-better
- the sort silently keeps near-misses
- query a document with its own text
basics
~20 sDistance is better when lower and similarity is better when higher. Code that sorts both the same way returns the worst matches from one of them, and any minimum-score threshold inverts into a keep-only-the-worst filter. Nothing crashes; results merely become quietly wrong.
solid answer
~50 sThe two conventions run in opposite directions: a similarity such as cosine is best at its maximum, while a distance such as cosine distance (1 minus cosine) or Euclidean distance is best at zero. A retrieval layer that takes `results[:k]` after a single sort direction, or applies `score > 0.7` uniformly, silently returns near-misses from whichever backend uses the other convention. The failure is dangerous precisely because it is not an exception — you still get k plausible-looking documents, and the model downstream produces a fluent answer from them. The cheapest detector is a self-similarity test: embed a document, query with its exact text, and assert it comes back at rank one with a score near the perfect value for that metric. If it comes back last, the sign convention is flipped. The durable fix is an adapter at the storage boundary that converts every backend to one canonical direction, with thresholds stored next to the metric they were calibrated against.
go deeper
Know that some backends return a distance where lower is better and others a similarity where higher is better, and that you must check which one before sorting.
Name the concrete conversions — cosine similarity is 1 minus cosine distance — and explain that both the sort direction and any threshold constant have to flip together.
Show the structural fix: a conversion adapter at the storage boundary plus a self-similarity assertion in CI, and explain why manual review cannot catch a failure that returns plausible documents.
Treat retrieval correctness as needing an oracle: mandate a labelled eval set and a canonical score contract across every store adapter, so a backend swap cannot silently degrade quality without a metric moving.
## Two conventions, opposite directions Vector comparison functions come in two flavours and the field uses both names loosely. *Similarities* are better when larger. Cosine similarity is best at 1. A raw inner product is better when larger, with no upper bound. *Distances* are better when smaller. Euclidean distance is best at 0. Cosine distance is conventionally defined as 1 minus cosine similarity, so it runs from 0 (identical direction) to 2 (opposite direction) and is likewise best at 0. Some libraries also expose an inner product through a distance-shaped interface by negating it, so that smaller remains better and one sort direction works for everything they expose. A system that talks to more than one backend — a vector store here, a library there, a cached score somewhere else — will therefore see fields called `score`, `distance`, `similarity` and `_distance` that do not agree on which end of the range is good. ## What the bug actually looks like Suppose a retrieval helper does the natural thing: ``` hits = sorted(raw, key=lambda h: h["score"], reverse=True)[:k] ``` Against a similarity-returning backend this is right. Against a distance-returning backend it returns the k *least* similar documents in the candidate set. Nothing raises. The types line up, the count is right, and every returned document is a real document from the corpus. Downstream, a language model reads them and writes a confident, fluent answer grounded in irrelevant material. Thresholds fail in the same silent way. A relevance gate written as "keep hits above 0.75" is a sensible floor on cosine similarity. Applied to a cosine distance it keeps only items whose direction is *worse* than a 0.25 cosine — the exact complement of what was intended. Worse, it will often let plenty of items through, so the pipeline looks healthy. A subtler variant is the partial migration: someone swaps the backend and correctly flips the sort, but the threshold constant elsewhere in the code is left alone. Ranking is then right and filtering is inverted. ## Detecting it The single most valuable test is self-similarity. Take a document already in the index, embed its exact text as the query, and assert two things: it is returned at rank one, and its score is the perfect value for the declared metric — 1.0 for cosine similarity, 0.0 for cosine or Euclidean distance. This one assertion catches a flipped sort, a flipped threshold, and a mismatched metric between build time and query time. It is fast enough to run in CI against a tiny fixture index. Supporting checks: assert the observed score range matches the declared metric (values above 1 rule out cosine similarity; negative values rule out a distance); and assert monotonicity on a hand-built triple where you know that A is more similar to Q than B is. ## The durable fix Convert at the boundary. Every store adapter should return one canonical shape — for example, always a similarity where higher is better — and perform the conversion itself using the identity for its metric. Cosine similarity is 1 minus cosine distance. A negated inner product is negated back. Squared Euclidean distance is converted, when vectors are unit-normalized, through the relation that squared distance equals 2 minus twice the cosine. Two supporting habits make this stick. First, name the field for what it is: `cosine_similarity` or `l2_distance`, never a bare `score`, so that a reader of the call site can see the direction. Second, store any calibrated threshold together with the metric it was calibrated against, and convert it explicitly when the metric changes rather than copying the number across. ## Why this class of bug survives review Semantic search has no oracle in the loop. A wrong result set is still a set of real documents about roughly the right corpus, and human reviewers reading a handful of results often cannot tell a mediocre retriever from an inverted one without ground truth. That is why the fix has to be structural — a conversion boundary plus an automated self-similarity assertion — rather than a code-review convention that someone will forget on the next backend swap.
- What is the cheapest test that catches this in CI?A self-similarity assertion. Index a known document, query with its exact text, and assert it returns at rank one with the perfect score for the declared metric — 1.0 for cosine similarity, 0.0 for a distance. It runs against a tiny fixture index in milliseconds and simultaneously catches a flipped sort, an inverted threshold, and a metric mismatch between index build and query time.
- How do calibrated thresholds interact with this?They invert independently of the sort. A "keep above 0.75" relevance floor is a sensible cut on cosine similarity and becomes a keep-only-the-worst filter on cosine distance — and it still lets plenty of items through, so the pipeline looks healthy. Store every threshold alongside the metric it was calibrated for, and convert it through the metric's identity when you switch rather than reusing the number.
- Why does this bug usually escape code review and manual QA?Because the failure is plausible rather than loud. No exception is raised, the result count is correct, and every returned document is real. A reviewer skimming a few results from the right corpus cannot distinguish an inverted retriever from a mediocre one without ground truth. Only a labelled eval set or the self-similarity assertion makes the difference visible.
saying these in an interview costs you the question
- Assumes higher is always better in vector search results
- Treats cosine distance and cosine similarity as interchangeable numbers
- Expects a convention mismatch to surface as an error or empty result
- Flips the sort direction during a migration but leaves threshold constants untouched
- Names the returned field a bare score with no indication of direction