Why does a highly selective filter change Weaviate's vector search cost and recall?
answer
- filter first, then search
- post-filtering would starve the result list
- selectivity and latency are not monotonic
- small permitted set means exact scan
- the inverted index supplies the allow-list
basics
~20 sWeaviate resolves the filter against its inverted index first and searches only the matching set, so limit is honoured against filtered objects rather than being trimmed afterwards. When that set is small, the engine switches to an exact scan of it, which changes latency and eliminates approximation error.
solid answer
~60 sWeaviate applies filters before and during the vector search, not after it. The filter is resolved against the inverted index into a set of permitted objects, and the vector search is constrained to that set. The practical consequences are the ones interviewers probe. First, `limit=10` with a filter returns ten *matching* objects if ten exist — you do not silently get three because seven of the top ten were filtered out, which is what post-filtering would give you. Second, when the permitted set is small relative to the collection, traversing an approximate graph is wasteful and can even miss matches, so Weaviate falls back to an exact scan over that set; recall becomes perfect and latency scales with the size of the filtered set rather than the whole collection. Third, a filter on a property that was not indexed for filtering has no fast set to build from, which is where surprising slowness comes from. So a tight filter is usually cheaper and more accurate, and a loose filter over a huge collection is where cost lives.
code
python · 18 linesfrom weaviate.classes.query import Filter, MetadataQuery
docs = client.collections.get("Document")
# validate the filter alone before trusting a filtered vector search
matching = docs.query.fetch_objects(
filters=Filter.by_property("tenant_id").equal("acme"),
limit=1,
)
print("filter matches something:", len(matching.objects) > 0)
# limit counts objects that satisfy the filter, not survivors of a post-filter
hits = docs.query.near_text(
query="renewal terms",
limit=10,
filters=Filter.by_property("tenant_id").equal("acme"),
return_metadata=MetadataQuery(distance=True),
)go deeper
Know that a filter constrains which objects a Weaviate vector search may return, so asking for ten results with a filter gives ten matching objects rather than whatever survives filtering afterwards.
Explain the difference between pre-filtering and post-filtering and why Weaviate's choice keeps result counts predictable, and mention that filters resolve through the inverted index.
Diagnose real behaviour: explain the non-monotonic latency curve across loose, middling and tight filters, the switch to exact scanning on small permitted sets, and why a filter on an unindexed property is unexpectedly slow.
Set the measurement policy — benchmarks and recall targets must be run with production filters, since unfiltered numbers do not transfer — and own the schema-level decision of which properties get filter indexes, given the storage and write cost each one adds.
## The question behind the question "Filter plus vector search" has three possible implementations and each has a different failure mode. Knowing which one Weaviate uses is what this question tests. **Post-filtering** searches the vector index first, then discards results that fail the filter. It is trivial to implement and badly behaved: ask for ten with a filter that matches one percent of the corpus and you routinely get zero, because none of the top ten passed. Result counts become unpredictable in exactly the situation where you needed the filter most. **Naive pre-filtering** finds every matching object and compares the query vector to each. Perfect results, but the cost is linear in how many objects match — fine for a thousand, ruinous for ten million. **Filtered graph traversal** walks the approximate index while consulting a set of permitted objects, skipping the rest. This keeps the sublinear behaviour while respecting the constraint. Weaviate does not post-filter. It resolves the filter into a permitted set from its inverted index and then either traverses the vector index restricted to that set, or — when the set is small enough to make traversal pointless — scans the set exactly. ## Why that choice shows up in your metrics ### Result counts stay predictable With pre-filtering, `limit` counts objects that satisfy the filter. Ten requested, ten returned, assuming ten exist. This is the property that makes tenant-scoped or permission-scoped search usable at all: a per-user filter that matches a thousandth of the corpus still fills a page of results. ### Latency is not monotonic in selectivity Engineers expect "more filtering equals more work". The curve is closer to U-shaped: - **Loose filter** (matches most of the collection): behaves almost like an unfiltered search. Cost is dominated by graph traversal, plus a small overhead for consulting the permitted set at each step. - **Middling filter** (matches a modest fraction): the worst region. The graph traversal keeps landing on objects that fail the filter, so it explores more nodes to fill the result list. - **Tight filter** (matches a small set): cheapest and most accurate, because the engine simply compares the query vector against every permitted object. Exact results, latency proportional to the size of that small set. The crossover between graph traversal and exact scan is governed by a cutoff configured on the collection's vector index, which is why two collections with identical data can behave differently under the same filter. ### Recall changes shape Approximate search means the returned neighbours are usually, not provably, the true nearest ones. Once a filter is tight enough to trigger an exact scan, that caveat disappears for those queries — recall is exact within the filtered set. Conversely, in the middling region a filtered graph traversal can miss matches that a brute-force comparison would have found, because the permitted objects may be poorly connected in a graph built over all the data. If you are benchmarking recall, benchmark it *with* your production filters; unfiltered recall numbers do not transfer. ## The operational traps **Unindexed filter properties.** The permitted set comes from the inverted index. A property that was not indexed for filtering when the collection was created cannot contribute a fast set, so the filter becomes far more expensive than its selectivity suggests. When a filter is unexpectedly slow, check the schema before blaming the vector index. **High-cardinality equality filters.** Filtering by a per-object identifier is technically supported and almost always a modelling mistake: you are using a nearest-neighbour engine to do a primary-key lookup. Fetch by id instead. **Wildcard-leading `like` patterns.** A pattern beginning with `*` cannot use the index efficiently and degrades toward a scan of the property's terms. Prefix patterns are fine; leading wildcards are not. **Filters that match nothing.** A tight filter with a typo returns an empty set, and the query dutifully returns nothing at all. Validate with `fetch_objects` using the same filter and no vector search before concluding that the vector index is broken. **Benchmarks without filters.** Load tests that omit the tenant or permission filter present in production measure a workload you do not run. Because latency is not monotonic in selectivity, such a benchmark can be optimistic *or* pessimistic — it is simply not informative. ## What to say in an interview State plainly that Weaviate pre-filters via the inverted index rather than post-filtering; explain that this is what keeps `limit` meaningful under selective filters; note the switch to exact scanning below a configured cutoff and what it does to latency and recall; and finish with the practical rule that filter properties must be indexed for filtering and that recall benchmarks must include production filters.
- A per-tenant filter matches 0.1 percent of a large collection, and latency drops rather than rising. Why?Below a configured cutoff Weaviate stops traversing the approximate graph and compares the query vector against the permitted objects directly. Work then scales with the size of that tiny set instead of with graph traversal over the whole collection, so the query gets faster and its results become exact rather than approximate.
- Why can measured recall differ between a filtered and an unfiltered version of the same query?Approximate graph traversal is built over all the data. When only a subset is permitted, the reachable path through the graph can be poor, so genuine matches are occasionally missed at moderate selectivity. Very tight filters avoid this by scanning exactly. Benchmark recall with your production filters applied, never without them.
- A filter on one property is far slower than an equally selective filter on another. What do you suspect?That the slow property is not indexed for filtering in the collection schema, so no fast permitted set can be built from the inverted index. Leading-wildcard like patterns and cross-reference filters that require an extra traversal per candidate produce the same symptom. Check the schema and the operator before tuning anything vector-related.
saying these in an interview costs you the question
- Believing Weaviate filters results after the vector search
- Expecting fewer than limit results when a filter is selective
- Assuming tighter filters always mean slower queries
- Filtering on properties never indexed for filtering
- Benchmarking recall without the filters production actually sends