skip to content

Metadata Filtering and Scoped Retrieval

Real corpora are multi-tenant and time-sensitive, so retrieval is rarely pure similarity. You constrain it with metadata predicates, and where that filter is applied changes both recall and latency.

on this pageshow

questions

5

In RAG retrieval, why filter by metadata instead of relying on similarity alone?

level: juniorimportance: must knowfreq 60%

answer

  1. similarity has no notion of ownership
  2. embeddings cannot express permissions
  3. near-duplicates across tenants and versions
  4. predicate decides membership, score decides order
  5. tenant, ACL, status, effective date, version

basics

~20 s

Similarity measures only how close two texts are in meaning. It says nothing about who may read a document or whether it is still current. Metadata predicates on fields such as tenant, ACL group, status and version enforce scope that embeddings cannot express.

solid answer

~50 s

A vector search returns the k nearest neighbours under a metric like cosine similarity, and that score expresses one thing: semantic closeness. It has no opinion on ownership, permissions, freshness or document status. So retrieval requests pair the query vector with a **predicate over stored metadata** — `tenant_id`, ACL groups, `status`, `effective_date`, product version — and only vectors satisfying that predicate are eligible to be returned. Three kinds of scope matter in practice. Authorization scope (tenant, ACL, region) decides what the user is allowed to see at all. Correctness scope (version, effective date, draft vs published) decides what is still true — on a fast-moving release-notes index, a passage about a two-year-old release is a perfect semantic match and a wrong answer. Precision scope (language, document type, source) simply keeps noise out of a small top-k. The predicate decides membership, not order. It is a hard constraint, not a boost.

go deeper

for a junior

Be able to say plainly that similarity only measures meaning, and that things like tenant, permissions, status and version live in metadata and are applied as a filter on the retrieval request.

for a middle

Explain the split — the predicate decides which vectors are eligible, the score decides their order — and give one authorization example and one freshness example where similarity alone produces a confidently wrong answer.

for a senior

Show you treat metadata as an index you must operate: payload fields need their own indexes, ACL copies go stale on revocation, and unset fields on backfilled documents silently exclude them from every scoped query.

for a principal

Own where scope is defined at all — which fields are contractually required at ingestion, who owns their freshness, and whether authorization is resolved from the identity system at query time rather than frozen into payloads across the whole corpus.

## What a similarity score actually measures An embedding model maps text into a vector space so that passages with similar meaning land close together. A vector search embeds the query, then returns the k nearest stored vectors under a metric such as cosine similarity or dot product. That is the entire content of the score: semantic closeness. It carries no notion of who owns a document, whether it has been superseded, which product version it describes, or whether it was ever approved for publication. Worse, similarity actively works against scope. Two passages describing parental-leave entitlement under a French policy and under a Brazilian one are near-neighbours in embedding space *precisely because* they say nearly the same thing about nearly the same subject. Similarity is what makes them confusable; it is not what separates them. The same holds for release notes across versions, contract templates across regions, and runbooks across environments. The more uniform your corpus, the less similarity can distinguish the copies. ## What metadata predicates are Alongside each vector, a vector store keeps a structured payload: fields like `tenant_id`, a list of ACL group ids, `doc_type`, `language`, `source_system`, `effective_date`, `product_version`, `status`. A retrieval request then carries two things — the query vector, and a predicate over those fields, for example `tenant_id = "acme" AND status = "published" AND version IN ("4.2", "4.3")`. Only vectors that satisfy the predicate are eligible to be returned. Among those, similarity decides the ranking. The division of labour is the point: **the predicate decides membership, similarity decides order.** A filter is not a ranking signal with a heavy weight; it is a yes/no gate. ## Three kinds of scope **Authorization scope** — tenant, ACL groups, region, classification. This is access control expressed as a query predicate. It is non-negotiable and must be derived from the authenticated session, not from anything the user or the model supplies. **Correctness scope** — version, effective date, draft vs published, deprecated vs current. Consider an assistant over product release notes where only the last two minor versions should be answerable. Without a version predicate, a question about a config option happily retrieves the passage from an old release that described the option before it was renamed. The retrieval is semantically excellent and the answer is wrong. Recency predicates are the cheapest available fix for that class of error, far cheaper than trying to teach the model to notice dates buried in chunk text. **Precision scope** — language, document type, source system. These do not protect anything; they just stop a small top-k from being consumed by material that could never answer the question, such as changelog boilerplate crowding out the actual policy text. ## Why encoding scope in the chunk text does not work A common shortcut is to prepend `Tenant: acme. Version 4.3.` to each chunk so the embedding "knows" its scope. This fails for two reasons. First, it is a soft signal: the model may weight it near zero next to several hundred words of body text, and the nearest neighbours can still be from the wrong tenant or the wrong version. Second, it spends embedding capacity and chunk tokens on boilerplate that is identical across thousands of chunks, blurring them together. Structured scope belongs in structured fields where it can be checked exactly. ## What it costs to run Metadata is a second index to maintain. Payload fields usually need their own indexes for a predicate to be evaluated efficiently, and every field must be kept in sync with its source of truth. The dangerous version of drift is an ACL list copied into the payload at ingestion time: when a permission is revoked, the vector store keeps honouring the stale copy until the document is reindexed. Teams handle that either by pushing permission-change events into the index or by storing stable group ids and resolving the user's current groups at query time. A second cost is emptiness. Once scope is enforced, some queries legitimately match nothing, and the pipeline needs a defined behaviour for that case rather than accidentally returning whatever was nearest. ## Failure modes to recognise No predicate at all, on a multi-tenant corpus, is a cross-tenant disclosure waiting to happen. A predicate applied as a score boost rather than a gate leaks in exactly the cases that matter — a strongly matching out-of-scope document outranks the boost. A predicate on a field that is unset for older documents silently excludes the entire backfill. And a version predicate that nobody refreshes quietly narrows the answerable corpus to nothing as releases move on.

  • Why not just prepend the tenant and version into the chunk text before embedding it?
    Because that makes scope a soft signal. The embedding may weight a short header near zero against hundreds of words of body text, so nearest neighbours can still come from the wrong tenant or version. It also wastes chunk tokens on boilerplate repeated across thousands of chunks, which blurs them together. Structured scope belongs in payload fields where a predicate can check it exactly.
  • An ACL group list is stored in each document's payload. What breaks when a user's access is revoked?
    Nothing breaks visibly, which is the problem: the store keeps honouring the stale payload until that document is reindexed, so the revoked user keeps retrieving it. Two common fixes are pushing permission-change events into the index so payloads update promptly, or storing stable group ids in the payload and resolving the caller's current groups from the identity system at query time.
  • Give a case where a filter improves answer correctness rather than security.
    An assistant over product release notes that admits only the last two minor versions. Without the version predicate, a question about a configuration option retrieves the passage from an old release that described it before it was renamed — a semantically excellent match and a wrong answer. The predicate removes the superseded material from contention entirely, which no amount of ranking can do.

saying these in an interview costs you the question

  • Claims a good embedding model already separates tenants
  • Puts tenant and version in chunk text instead of payload fields
  • Treats a filter as a score boost rather than a gate
  • Assumes the newest document always ranks highest by similarity
  • Never refreshes version or date predicates as the corpus moves

context

open as a page

In vector search, how do pre-filtering, post-filtering and filtered ANN traversal differ?

level: middleimportance: must knowfreq 72%

basics

~20 s

Pre-filtering restricts the candidate set first and searches only inside it. Post-filtering runs the ANN search unaware of the predicate and drops non-matching hits afterwards, so results shrink. Filtered traversal evaluates the predicate during graph search, returning a full in-scope top-k without over-fetching.

open as a page

In multi-tenant RAG, why must a tenant filter be an authorization boundary, not a ranking hint?

level: seniorimportance: must knowfreq 58%

basics

~20 s

A ranking hint can be outvoted. Boost in-tenant documents and a strongly similar out-of-tenant passage still enters top-k, reaches the prompt, and gets quoted back. Scope must be a hard predicate the engine enforces on every query, derived from the authenticated session.

open as a page

When scoped retrieval returns zero documents after metadata filtering, what should the system do?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Return an explicit empty-scope signal rather than silently proceeding. Widening is permitted only on preference predicates such as recency or document type, and the relaxation must be labelled; security predicates like tenant and ACL are never relaxed, and the generation step must know the context is empty.

open as a page

Why does a metadata filter matching 0.01% of vectors slow ANN search down?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Because the search must keep exploring until it collects k matching vectors. When almost every node it visits fails the predicate, traversal wanders across most of the index to find enough hits, doing roughly full-scan work through a structure built to avoid full scans.

open as a page