In Weaviate, how do the distance and certainty thresholds on near_text differ?
answer
- two scales, one similarity
- lower is closer, higher is closer
- one of them is cosine-only
- a cutoff may return zero rows
- measure with metadata before guessing
basics
~20 sdistance sets a maximum: results further than the value are dropped, and lower is closer. certainty sets a minimum on a normalised 0-to-1 similarity where higher is closer, and it is only defined when the collection uses the cosine metric. Set one, not both.
solid answer
~50 sBoth are similarity cutoffs on the same search, expressed on opposite scales. `distance=0.25` means "return nothing further away than 0.25". Lower is more similar, and the numeric range depends on the collection's distance metric — for cosine it runs 0 to 2. `certainty=0.8` means "return nothing below 0.8 similar", on a normalised 0-to-1 scale where higher is better. Certainty is a cosine-specific reinterpretation of distance, so a collection configured with dot product or L2 cannot use it. A threshold is not a replacement for `limit`; it composes with it. With `limit=10, distance=0.25` you get *at most* ten objects and possibly zero — the cutoff can legitimately empty the result set, which is the point when you would rather show nothing than show bad matches. Add `return_metadata=MetadataQuery(distance=True)` to see the actual distances and pick a threshold from data instead of guessing.
code
python · 21 linesfrom weaviate.classes.query import MetadataQuery
products = client.collections.get("Product")
# maximum distance: lower is closer, may return fewer than limit
tight = products.query.near_text(
query="noise cancelling headphones",
limit=10,
distance=0.25,
return_metadata=MetadataQuery(distance=True),
)
for o in tight.objects:
print(o.properties["name"], o.metadata.distance)
# minimum certainty: cosine collections only, higher is closer
confident = products.query.near_text(
query="noise cancelling headphones",
limit=10,
certainty=0.8,
)go deeper
Remember the direction of each scale: distance is lower-is-closer and acts as a maximum, certainty is higher-is-closer and acts as a minimum. Say that a threshold can return fewer results than the limit.
Explain that the two express the same similarity on different scales, that certainty exists only for cosine collections, and that thresholds compose with limit rather than replacing it.
Demonstrate how you actually pick a value: run labelled queries with distance metadata, find the gap between relevant and irrelevant hits, and defend empty results as correct behaviour rather than a bug to be tuned away.
Frame thresholds as versioned configuration tied to an embedding model. Own the policy question of when a system should answer "nothing found" instead of returning its least-bad match, and the re-tuning cost every model migration carries.
## Two dials on the same search Every Weaviate vector query returns objects ordered from closest to furthest. `limit` cuts that ordered list at a fixed position; a threshold cuts it wherever quality falls off. They answer different questions — "how many do I want?" versus "how bad am I willing to accept?" — and production queries usually need both. ## distance: a maximum, lower is better ```python response = collection.query.near_text( query="noise cancelling headphones", limit=10, distance=0.25, return_metadata=MetadataQuery(distance=True), ) ``` `distance` is an upper bound. Anything whose distance from the query vector exceeds the value is excluded, no matter how few results survive. The scale is defined by the distance metric the collection was created with, and this is where people get burned: 0.25 is a tight cutoff under cosine distance and a meaningless one under squared-L2, where distances are unbounded and scale with vector magnitude and dimensionality. There is no universal good value. The only honest way to pick one is to run representative queries with `return_metadata=MetadataQuery(distance=True)`, look at the distances of results you judge good and bad, and set the cutoff between them. ## certainty: a minimum, higher is better, cosine only `certainty` presents the same underlying similarity on a friendlier scale: 0 to 1, higher is closer, and it reads like a confidence. It is derived from cosine distance, so it exists only for collections using the cosine metric. Ask for certainty on a dot-product or L2 collection and the server rejects the request rather than inventing a normalisation. Because the two express the same quantity, specifying both in one query is contradictory and is not allowed. Pick the scale your team reasons about more naturally and stick to it. Teams that display a "match score" in a UI tend to prefer certainty; teams that tune retrieval offline tend to prefer distance because it is the raw number the index works with. ## Thresholds and limit compose The two are applied together, and the tighter one wins: - `limit=10, distance=0.25` → up to ten results, all within 0.25. Could be ten, could be three, could be none. - `limit=100, distance=0.9` (cosine) → probably a hundred results, most of them poor, because the cutoff is loose. - `limit=10` with no threshold → always ten results if the collection holds ten objects, however irrelevant they are. That last line is the failure mode a threshold exists to prevent. A bare `limit` on a nearest-neighbour search *always* returns something. Ask a question your corpus has no answer to and you still get the ten least-bad vectors, ranked confidently. Downstream that becomes a language model summarising documents that were never relevant. A distance cutoff is how you let a search legitimately return nothing. ## Empty results are a feature, not a bug The most common reaction to a threshold is alarm when a query returns zero rows. Before loosening the cutoff, check what the distances actually were. If the closest object sits at 0.4 and your cutoff is 0.25, the search worked correctly and told you something true: the corpus has no good answer. Loosening the cutoff does not create a good answer, it hides the absence of one. ## Practical tuning loop 1. Collect twenty or thirty realistic queries, including some that should have no answer. 2. Run them with a generous `limit`, no threshold, and `return_metadata=MetadataQuery(distance=True)`. 3. Hand-label which returned objects are actually relevant. 4. Plot or eyeball the distances: relevant hits cluster low, irrelevant hits cluster higher, and there is usually a visible gap. 5. Put the cutoff in the gap, then re-run the no-answer queries to confirm they now come back empty. Re-run that loop whenever the embedding model changes. Distances are a property of the vector space, so a new model invalidates every threshold you tuned against the old one — a migration cost that is easy to forget precisely because nothing errors. ## When a fixed threshold is the wrong tool A fixed cutoff assumes distances are comparable across queries. For heterogeneous corpora that assumption is shaky: short queries, rare terms and unusual phrasings all shift the whole distance distribution. When per-query distributions vary that much, a relative cut based on where the results drop off is a better fit than any single absolute number, and Weaviate offers that as a separate option on the same query.
- Your near_text query with limit=5 and distance=0.2 returns nothing. What do you check first?Re-run it without the threshold but with return_metadata=MetadataQuery(distance=True) and look at the closest distances. If the best hit is far above 0.2, the search behaved correctly and the corpus genuinely lacks a good match. Only if good, relevant hits are being cut should you loosen the cutoff.
- Why can a threshold that worked well suddenly filter out everything after a deploy?Almost always an embedding-model change. Distances are a property of a specific vector space, so a new model shifts the whole distribution and every tuned cutoff becomes meaningless. Re-tune thresholds as part of any re-embedding, and treat them as model-version-scoped configuration rather than constants.
- When would you deliberately use limit with no similarity threshold at all?When the downstream stage does its own quality judgment — for example a reranker or a human reviewing a candidate list — and you want a fixed-size candidate set regardless of absolute similarity. In anything that feeds text straight to a user or a model, an unbounded limit is how irrelevant results reach production.
saying these in an interview costs you the question
- Thinking higher distance means a better match
- Expecting certainty to work with any distance metric
- Passing both distance and certainty in one query
- Treating a fixed cutoff like 0.7 as universally correct
- Assuming limit alone protects against irrelevant results