How do you restrict a $vectorSearch to documents matching a metadata condition?
answer
- The predicate belongs inside the stage, not after it
- The index must know which paths are filterable
- Not every query operator is accepted there
- A trailing match filters an already-truncated list
basics
~20 sPass a filter document to the $vectorSearch stage, and declare every path it touches as a filter field in the vector index definition. The condition is then applied during the search, so all limit results already satisfy it.
solid answer
~50 s`$vectorSearch` takes a `filter` option holding a MongoDB query predicate over metadata — tenant, language, status, a date range. Two things make it work. First, each path used must be declared in the vector index definition as `{ "type": "filter", "path": ... }` alongside the `vector` field; an undeclared path is not usable in the filter. Second, only a documented subset of query operators is supported there — equality and comparison operators such as `$eq`, `$ne`, `$gt`, `$gte`, `$lt`, `$lte`, `$in`, `$nin`, combined with `$and`/`$or` — not the full query language, so no `$regex`, `$expr` or `$text`. The payoff is that the condition is applied while candidates are selected, so the `limit` documents you get back all satisfy it. Writing the predicate as a `$match` after the stage instead discards from an already-truncated list and routinely returns far fewer results than requested.
code
json · 7 lines{
"fields": [
{ "type": "vector", "path": "embedding", "numDimensions": 1536, "similarity": "cosine" },
{ "type": "filter", "path": "tenantId" },
{ "type": "filter", "path": "publishedAt" }
]
}go deeper
Recall that the metadata condition goes inside the $vectorSearch stage's filter option, and that the field must be declared in the vector index before you can filter on it.
Explain why a trailing $match returns fewer than limit documents, and name the scalar types and operator subset the filter option accepts.
Show you handle the selectivity interaction — restrictive filters need more candidates — and that a tenant filter is an authorization boundary built server-side, not a ranking hint.
Own the modelling decision: which metadata becomes filterable index state, when partitioning into separate collections beats filtering one index, and how tenant isolation is enforced so no code path can omit it.
## The stage option `$vectorSearch` accepts a `filter` field next to `path`, `queryVector`, `numCandidates` and `limit`: ```javascript { $vectorSearch: { index: "vector_index", path: "embedding", queryVector: qv, filter: { $and: [ { tenantId: "acme" }, { publishedAt: { $gte: cutoff } } ] }, numCandidates: 300, limit: 10 } } ``` The predicate is expressed in familiar MongoDB query syntax, but it is evaluated by mongot against its own index, not by `mongod` against the collection — which is why it is neither the full query language nor automatically available on every field. ## The index side: filter fields must be declared An Atlas Vector Search index definition is a `fields` array. One entry describes the vector itself — `type: "vector"`, the `path`, `numDimensions` matching the embedding model exactly, and a `similarity` of `euclidean`, `cosine` or `dotProduct`. Every metadata path you intend to filter on needs its own entry with `type: "filter"`. Filterable paths are scalar-shaped: strings, numbers, booleans, dates and objectIds (and arrays of those), which covers tenant ids, languages, statuses, tags and timestamps. Forgetting the declaration is the classic first bug: the vector index looks complete, the query looks correct, and the filter cannot be honoured. Declaring the field costs index space, so declare what you filter on, not every field in the document. ## Why it matters that filtering happens during the search This is the difference between two very different pipelines. **Filtering during the search** — the `filter` option — means the engine only ever accepts qualifying documents as candidates. Ask for `limit: 10` and you get 10 qualifying documents, assuming at least 10 exist and the search reaches them. **Filtering after the search** — a `$match` stage following `$vectorSearch` — means the engine picks its 10 nearest neighbours knowing nothing about your predicate, and then the `$match` throws away those that fail it. If 5% of the corpus belongs to the current tenant, you can expect roughly zero survivors. The failure is quiet and shape-dependent: it looks fine in a demo dataset where one tenant dominates and collapses in production. There is one legitimate use of a trailing `$match`: a predicate on a field you deliberately did not index as a filter field, applied to results you are willing to lose. Treat that as an exception with a comment, not a default. ## Multi-tenancy is an authorization boundary When the filter carries a tenant id, it stops being a relevance hint and becomes a security control. Build it server-side from the authenticated principal, never from client input, and make sure no code path can construct a `$vectorSearch` on that collection without it — a missing filter here leaks another customer's documents into a RAG answer. ## Selectivity and candidate breadth A very restrictive filter interacts with the approximate search: the traversal explores a neighbourhood in which most vectors fail the predicate, so a candidate count that was ample unfiltered can return short of `limit` or with degraded recall. The mitigation is to raise `numCandidates` for heavily filtered query classes, and to measure recall for those classes separately rather than trusting numbers from unfiltered benchmarks. Where a filter partitions the corpus permanently and coarsely, separate collections with their own indexes are sometimes a cleaner answer than one index filtered on every query. ## The parallel in Atlas Search `$search` has the same shape with different spelling: predicates belong in the `filter` clause of the `compound` operator, where mongot applies them during matching and, unlike `must` and `should`, they contribute nothing to the relevance score. Recognising these as the same idea in both stages is what an interviewer is listening for. ## What to say in an interview Name the `filter` option, state the index requirement (`type: "filter"` entries), note that only a subset of operators is supported, and explain the concrete consequence: a trailing `$match` returns fewer than `limit` because it filters a list that was already truncated.
- What happens if you filter on a path that is not declared as a filter field in the vector index?The filter cannot be honoured — the path simply is not part of the structures mongot searches. You have to add a `{ "type": "filter", "path": ... }` entry to the index definition, which triggers a rebuild. This is why filterable metadata is a design decision made with the index, not per query.
- Which query operators can you use in the $vectorSearch filter?A documented subset: equality and comparison operators such as `$eq`, `$ne`, `$gt`, `$gte`, `$lt`, `$lte`, `$in` and `$nin`, combined with logical operators like `$and` and `$or`. Full query-language features — `$regex`, `$expr`, `$text`, geospatial operators — are not available, so filterable metadata should be modelled as plain scalars.
- How does the equivalent look for a text query with $search?Predicates go in the `filter` clause of the `compound` operator, evaluated by mongot while matching. Unlike `must` and `should` clauses, `filter` clauses do not affect the relevance score, so they express structural constraints such as tenant or status without distorting ranking.
saying these in an interview costs you the question
- Applies the predicate as a $match after $vectorSearch
- Assumes any document field can be used in the filter
- Expects $regex or $expr to work inside the vector filter
- Builds a tenant filter from client-supplied input
- Keeps unfiltered benchmark numbers when the query is heavily filtered