How does the filter inside an Elasticsearch knn search differ from a post_filter?
answer
- One constrains the search, one trims the output
- Ask when the constraint is evaluated
- Selective constraints can starve a greedy graph walk
- Lucene has an exact fallback for tiny matching sets
- A sibling query clause is not an AND
basics
~20 sThe knn filter is applied during the HNSW graph traversal, so all k neighbours already match it. A post_filter runs after the neighbours are chosen and simply discards some, often leaving far fewer than k hits.
solid answer
~50 s`filter` inside the `knn` clause is a **pre-filter**: Lucene consults it while walking the graph, so non-matching documents never enter the candidate queue and you get k neighbours that all satisfy the constraint. `post_filter` runs on the result list after kNN has already picked its k, so a selective filter can leave you with three hits when you asked for ten — or none. There is a third trap: a `query` alongside a top-level `knn` does **not** restrict the kNN either; the two result lists are combined and their scores added. Pre-filtering is not free, though. A very selective filter starves the greedy walk, because most of the graph's neighbourhoods are ineligible, so recall drops unless you raise `num_candidates`. When the filter matches few enough documents, Lucene abandons the graph and brute-forces the matching set exactly, which is usually the right call.
code
json · 12 linesPOST /articles/_search
{
"knn": {
"field": "body_embedding",
"query_vector": [0.11, 0.42, -0.9],
"k": 10,
"num_candidates": 200,
"filter": {
"term": { "status": "published" }
}
}
}go deeper
Know that the constraint belongs inside the knn clause's filter, and that a post_filter can leave you with fewer hits than the k you asked for.
Explain that the knn filter is evaluated during graph traversal while post_filter runs on the finished list, and that a sibling query clause combines results rather than intersecting them.
Diagnose filtered kNN in production: recognise recall collapse under selective filters, raise num_candidates for filtered traffic specifically, and know that Lucene falls back to exact search over small matching sets.
Own the data layout when filtering is inherent to the workload — tenant routing, index-per-large-tenant, or separate graphs — rather than accepting degraded recall as the cost of multi-tenancy.
## Three places a constraint can live Given a kNN search and a constraint like `status: published`, there are three syntactically similar places to put it, and they behave completely differently. **Inside the knn clause (`knn.filter`) — a pre-filter.** The filter is evaluated during graph traversal. Only matching documents are eligible to enter the candidate queue, so the search returns k neighbours that all match. This is almost always what you want. **As `post_filter` — a post-filter.** kNN runs unconstrained, produces its k nearest neighbours, and then the filter deletes the ones that do not match. If only 20% of your corpus is published, a k of 10 typically yields around two hits, and there is no mechanism to backfill. The result count becomes a random variable driven by how the vectors happen to be distributed. **As a sibling `query` to a top-level `knn`.** This looks like an AND but is not one. The `knn` clause and the `query` clause are executed separately, their result sets are combined, and the scores are added. A document matching only the query still appears; a kNN neighbour failing the query still appears. Candidates routinely write this expecting an intersection and get a union. ## Why pre-filtering is hard, and what Lucene does about it HNSW works because the graph's edges encode proximity: a greedy walk toward the query vector converges quickly. Filtering breaks that assumption. If the eligible documents are scattered thinly across the vector space, the walk spends its budget stepping through ineligible neighbourhoods, and the bounded candidate queue fills with dead ends. Recall degrades, sometimes sharply, exactly when the filter is most selective. Lucene mitigates this with a fallback: when the filter matches few enough documents relative to the work the graph search would do, it abandons approximate search and performs an **exact** comparison over just the matching documents. Brute-forcing a few hundred vectors is cheap and gives perfect recall, so this is a strictly better outcome. The pathological zone is in between — a filter selective enough to hurt the graph walk but broad enough that brute force is still expensive. Practical consequences: - Measure recall for **filtered** traffic separately from unfiltered traffic. They are different workloads with different tuning. - Filtered searches usually need a larger `num_candidates` than unfiltered ones to hit the same recall, because non-matching candidates consume traversal budget without filling the top-k. - Very high-cardinality per-user filters (one tenant per document, thousands of tenants) are a known bad fit for a single shared graph. Routing tenants to their own shards or indices, so each graph contains mostly eligible documents, often beats fighting the filter. ## Writing the filter The `knn` filter takes a normal query, and it runs in filter context — no scoring, cacheable. Keep it to `term`, `terms`, `range` and boolean combinations over `keyword`, numeric and date fields. Putting an expensive scripted condition there means paying it during traversal, once per candidate considered, which is a very different cost profile than paying it once per matching document in a normal search. Security filters deserve a note: document-level security is applied by Elasticsearch itself and constrains what a user can see, so it behaves as a constraint on the candidate set rather than as something you can forget to add. It has the same recall implications as any other selective pre-filter. ## Diagnosing the symptom in production "The kNN search returns fewer results than k" has a short list of causes, and you can walk it quickly: 1. A `post_filter` is trimming the list — move the constraint into `knn.filter`. 2. The filter is so selective that fewer than k documents match at all — no configuration fixes that. 3. The index genuinely holds fewer than k vectors on some shards. "The kNN search returns the wrong neighbours when filtered" is the recall problem instead: compare against an exact baseline over the filtered set, raise `num_candidates`, and if that is not enough, revisit `m` and the quantization choice. ## The one-line summary Pre-filter to constrain the search, post-filter only when you deliberately want the unfiltered neighbourhood computed first (for example to keep aggregation facets over the full neighbourhood), and never assume a sibling `query` clause narrows a kNN search.
- Why can a highly selective knn filter reduce recall rather than just narrowing results?HNSW converges by following edges to nearby vectors. When most neighbourhoods are ineligible, the greedy walk burns its bounded candidate queue on documents that cannot be returned, and it may terminate before reaching the eligible region of the space. The fix is a larger `num_candidates` for filtered traffic, a better-connected graph via `m`, or partitioning the data so each graph is mostly eligible.
- When is a post_filter on a kNN search actually the right choice?When you deliberately want the unfiltered neighbourhood computed first — for example when aggregations should reflect the full nearest-neighbour set while the displayed hits are narrowed, since `post_filter` runs after aggregations are collected. Outside that case it is a bug: it makes your result count depend on how the vectors happen to fall rather than on what you asked for.
- You have thousands of tenants and every kNN search filters to one of them. What do you do?Stop relying on the filter alone. Route each tenant's documents to a specific shard, or give large tenants their own index, so the graph a search traverses contains mostly eligible vectors. That restores recall and cuts traversal waste. Keep the filter as a correctness guarantee, but do not make it the only thing standing between the walk and 99% ineligible data.
A pre-filter is telling the recruiter up front which departments to search; a post-filter is interviewing the ten best people company-wide and then throwing out everyone from the wrong department.
saying these in an interview costs you the question
- Thinks a sibling query clause restricts the kNN results
- Says post_filter and knn filter are equivalent
- Assumes pre-filtering has no effect on recall
- Puts expensive scripted conditions in the knn filter
- Blames missing hits on k rather than on the post_filter