In multi-tenant RAG, why must a tenant filter be an authorization boundary, not a ranking hint?
answer
- preferences can be outvoted, boundaries cannot
- leakage proportional to similarity
- near-duplicate policies across regions
- scope from the session, not the request
- stale ACL in payload survives revocation
basics
~20 sA ranking hint can be outvoted. Boost in-tenant documents and a strongly similar out-of-tenant passage still enters top-k, reaches the prompt, and gets quoted back. Scope must be a hard predicate the engine enforces on every query, derived from the authenticated session.
solid answer
~60 sTreating tenant as a score boost makes disclosure a function of similarity, which is exactly the wrong dependency: the closer the wrong tenant's document is to the query, the likelier it is to leak. In a shared HR assistant, a France-only parental-leave policy and its Brazilian counterpart are near-identical text, so a boost of any realistic size will sometimes rank the wrong one first — and once it is in the context window, the model will quote it fluently. So the tenant condition has to be a hard predicate applied inside retrieval, evaluated before or during the search, on every query path. Three properties matter. It must be derived from the authenticated session server-side, never from a request field the client can set and never from a tenant the model infers from the question. It must be enforced on every path, including reindexing jobs, eval harnesses and debug endpoints. And the payload it checks must not be stale — a revoked group in a document payload keeps granting access until that document is reindexed. An empty in-scope result is a correct outcome here, never a reason to widen.
go deeper
Know that tenant and permission conditions are hard filters applied during retrieval, and that they come from who the user is signed in as, not from anything the request or the model supplies.
Explain why a boost cannot substitute for a gate: a sufficiently similar out-of-scope document beats the boost margin, and near-duplicate documents across regions or tenants are exactly where that happens.
Demonstrate operational thinking — enforce the predicate on every path including jobs and debug tools, handle ACL staleness on revocation, and prove scoping with a poisoned-corpus test rather than asserting it.
Own the isolation strategy itself: which customer segments get physical separation versus a shared filtered collection, what the contractual commitments are, and how scope enforcement is audited and evidenced across the whole retrieval surface.
## Why the distinction is not pedantic Retrieval systems have two kinds of conditions and it is easy to blur them. A *preference* says one document should outrank another; it participates in scoring and can be outvoted by a strong enough signal on the other side. A *boundary* says a document must not be returned at all; nothing outvotes it. Tenant, ACL group, region and classification are boundaries. Recency and document type are usually preferences dressed as filters. The failure mode when you implement a boundary as a preference is precise: leakage becomes a function of similarity. A boost of, say, +0.15 on in-tenant documents means any out-of-tenant document whose similarity exceeds the best in-tenant one by more than that margin is returned. On a corpus of near-duplicate documents, that margin is crossed regularly. Consider a shared HR assistant serving employees in multiple countries, where a France-only parental-leave policy and the Brazilian equivalent are structurally identical documents differing in a few clauses. A Brazil-based employee's question about leave entitlement is close to *both*. The boost decides which one wins, and sometimes it decides wrong — and when it does, the model does not hedge. It reads the retrieved passage as ground truth and states French entitlements as the employee's own, with a citation. ## Where the predicate must live **Server-side, from the authenticated session.** The tenant and group set must be resolved from the caller's verified identity inside the retrieval service. Any design where the client passes `tenant_id` in the request body, or where the orchestration layer lets the model choose a scope, is forgeable — the model in particular is an untrusted component here, because prompt content it has read can influence what it asks for. If an agent has a retrieval tool, the tool's scope arguments must be narrowing options within the session's scope, never the source of the scope itself. **Before or during the search, not after — and never optional.** Post-filtering does technically remove out-of-tenant documents, so it is not itself a disclosure; but it makes correctness depend on a step that lives outside the index, which is easy to skip. Every path that reads the collection must carry the predicate: the chat endpoint, the async summarisation job, the eval harness that replays production queries, the internal debug console. A single unfiltered path is the whole breach. **Against payload that is current.** The most common real leak is not a missing filter but a stale one. Copying a document's ACL group list into the vector payload at ingestion time means that revoking a group's access changes nothing in the index until that document is reindexed — an interval that can be hours or days. Two robust patterns: push permission-change events into the index and update payloads promptly, or store stable group identifiers in the payload and resolve the *caller's* current groups from the identity system at query time, so revocation takes effect on the next query. ## Isolation strategies above the predicate A shared collection with a tenant predicate is the dense, cheap option: one index, good memory utilisation, one thing to operate. The alternative is physical isolation — a collection, namespace or database per tenant — where cross-tenant retrieval is impossible by construction rather than by correct code. Physical isolation costs per-tenant overhead and scales badly into the thousands of small tenants, and it complicates any legitimately cross-tenant feature. Most teams run a shared collection with a hard predicate and reserve physical isolation for tenants with contractual or regulatory separation requirements. The honest framing in an interview is that this is a spectrum chosen per customer segment, not a single right answer. ## Proving it works Scope enforcement is testable in a way most RAG quality is not. Seed the index with a poisoned corpus: documents in tenant B that are deliberately near-duplicates of tenant A's, then assert that a large sweep of tenant-A queries returns zero tenant-B ids at any k. Run that in CI on every retrieval-path change. Add a runtime invariant that inspects returned ids against the session scope and fails loudly — cheap, and it catches the path someone added without the predicate. Log the predicate alongside every retrieval so an audit can reconstruct what a given answer was allowed to see. ## What not to do when the scope is empty When the tenant predicate leaves nothing, the correct outcome is an empty result. Re-running without the predicate, softening it to a penalty, or returning the nearest out-of-scope passages "for context" each convert an unhelpful answer into a disclosure. Widening is only ever legitimate on preference predicates such as recency, never on boundaries.
- Would you isolate tenants into separate collections instead of filtering one shared index?It is a spectrum, not a rule. A shared collection with a hard predicate gives dense memory use and one system to operate, and is what most teams run. Physical isolation makes cross-tenant retrieval impossible by construction, which is worth the per-tenant overhead for customers with contractual or regulatory separation — but it scales badly across thousands of small tenants and complicates any legitimately cross-tenant feature.
- An agent has a retrieval tool. May it pass a tenant argument?It may pass narrowing arguments within the session's scope — a project, a document type — but never the scope itself. The model is an untrusted component: content it has just read can influence what it asks for, so a model-chosen tenant is an injection target. Resolve the tenant from the authenticated session inside the tool implementation and intersect any model-supplied filter with it.
- How would you prove in CI that tenant scoping actually holds?Seed a test index where tenant B holds deliberate near-duplicates of tenant A's documents, then sweep a large set of tenant-A queries and assert zero tenant-B ids come back at any k. Pair it with a runtime invariant that checks returned ids against the session scope and fails loudly, which catches the new code path someone added without the predicate.
saying these in an interview costs you the question
- Implements tenant scope as a score boost or weighting
- Reads the tenant from a client-supplied request field
- Lets the model choose which tenant to retrieve from
- Relies on the prompt telling the model to ignore other tenants
- Drops the ACL predicate when the scoped result is empty