In a multi-facility RAG corpus, where must the tenant filter be enforced, and why?
answer
- retrieved is already disclosed
- filter before, not after
- identity from the session, never the model
- rebuilds drop the tenant tag silently
- canary queries that must return zero
basics
~20 sAt query time in the retrieval layer, as a pre-filter derived from the caller's verified identity — or as physically separate indexes. Anything above that, including instructions to the model, is advice rather than a boundary, because a retrieved chunk is already disclosed.
solid answer
~50 sOnce a chunk enters the context window it is disclosed: the model may quote it, paraphrase it, or let it shape an answer, and nothing downstream reliably undoes that. So the access decision has to happen **before** retrieval returns, using the caller's authenticated identity rather than anything the model or the user supplies. In practice that is a metadata pre-filter applied inside the query, or a separate index/namespace per tenant when the isolation must be visible in the infrastructure. Two things then need continuous care. Permission metadata must be attached at ingest **and** repopulated on every rebuild, because an index rebuilt without the facility tag silently stops excluding anything. And permissions drift: a document restricted in the source system after indexing leaves a stale grant behind unless you re-sync or re-check at query time. Test it as a security control — cross-tenant probes that must return zero rows, run in CI.
go deeper
Know that a retrieved chunk counts as disclosed, so the access check has to happen in the retrieval query itself, using the logged-in user's identity.
Contrast per-tenant indexes with metadata pre-filtering, and explain why filtering after retrieval both leaks less safely and quietly degrades results for legitimate users.
Show the operational failure modes you would design against: the reindex that drops permission tags, stale ACLs after upstream changes, semantic caches keyed without tenant, and the canary tests that catch them in CI.
Decide where isolation must be structural rather than logical, and defend the cost: separate indexes per tenant for records, shared indexes for shared material, and an explicit, signed-off bound on permission staleness.
## Why the boundary sits under the model A nurse at one facility of a nine-site hospital network asks the clinical-note assistant a question. Retrieval returns chunks; the model composes an answer. If a chunk from another facility's records reaches the context window, the disclosure has already happened. The model may quote it verbatim, may paraphrase it, may reveal it indirectly by answering a question it could otherwise not have answered, and may surface it in a citation list. There is no reliable way to un-see a chunk. That single observation determines the whole design: authorization must be enforced at or before retrieval, in code, using an identity the model cannot influence. This is a privacy boundary, not a relevance preference. Relevance is about ranking; a boundary is about a set of documents that must never be candidates at all. ## Enforcement points, ranked **Separate index or namespace per tenant.** The strongest option, because isolation is structural: a query issued against one facility's index cannot physically reach another's. It costs operational overhead — more indexes to build, migrate and monitor — and it fragments corpora that legitimately should be shared, such as network-wide clinical guidelines. Common compromise: a shared index for genuinely shared material, per-tenant indexes for records. **Metadata pre-filter inside the query.** The usual choice. Every chunk carries a facility (and often unit or role) tag written at ingest, and the query includes a filter derived from the caller's session. The critical property is that the filter is applied *during* the search, restricting the candidate set, rather than after results come back. It depends on two things being right — the tags and the filter — and both are easy to break silently. **Post-retrieval filtering in application code.** Weaker but not worthless. It defends against a filter bug in the search layer, and it is a good place for a second check against the system of record. Its flaw is that if it is the *only* control, a top-k search has already spent its slots on documents that will be discarded, so recall for the legitimate user quietly collapses. **Instructing the model.** Not a control. "Only answer using documents from the user's own facility" is a probabilistic nudge in a channel that also carries untrusted document text. It belongs, at most, as defence in depth after the real filter. **Filtering the generated answer.** Too late by construction, and easily defeated by paraphrase. Checking output for another facility's name catches the clumsiest leak and nothing else. ## Where these systems actually break **The rebuild regression.** This is the classic. Someone re-embeds the corpus after a model or chunking change and the pipeline that populated the facility tag is not part of the new path — or is, but writes a null. Every filtered query now matches everything, or matches nothing, depending on filter semantics. Nothing errors. Retrieval quality even looks fine, because the extra documents are topically similar. The only detection is a test that asserts a cross-tenant query returns zero results, plus an ingest-time invariant that rejects any chunk lacking a tenant tag. **ACL drift.** Permissions live in the source system and are copied into the index at ingest. When a document is restricted, moved or deleted upstream, the copy is stale until the next sync. For high-sensitivity corpora, re-check the returned identifiers against the system of record before they enter the prompt; for lower-sensitivity ones, bound the drift with an aggressive sync and accept the window explicitly rather than by accident. **Derived artefacts forgetting their origin.** Summaries, extracted entities, embeddings, caches, evaluation sets and fine-tuning corpora derived from tenant data inherit its sensitivity — and frequently lose its tags. A cross-tenant summary index built "for efficiency" is a leak with a friendly name. Semantic caches are a specific trap: a cache keyed on question similarity alone will serve one facility's answer to another's user. Tenant identity must be part of every cache key. **Identity from the wrong place.** The filter must come from the authenticated session on the server. If the tenant identifier arrives as a request parameter, a client-side value, or worse a field the model fills in when calling a retrieval tool, the boundary is decorative — an agent that can choose its own filter has no filter. **Role inside the tenant.** Facility separation is rarely the whole requirement. A nurse may reach her own unit's records but not another department's; a research role may see de-identified extracts only. That means the filter is generally a predicate over several attributes, and it must be derivable from the session without a round trip the retrieval path cannot afford. ## Proving it holds Treat isolation as a tested control. Seed each tenant's fixtures with a distinctive canary string, then assert that a query engineered to be maximally similar to another tenant's canary returns nothing across every retrieval path — including any summary index, cache and evaluation harness. Run it in CI and after every reindex. Add an ingest assertion that no chunk is written without a tenant tag, and alert on any query issued with an empty or wildcard filter. Log the filter that was applied alongside the request so an audit can answer what the caller was permitted to see, not merely what they saw.
- Is post-retrieval filtering in application code worth keeping if the query already filters?Yes, as defence in depth against a bug in the search layer, and as the natural place to re-check identifiers against the system of record. What it must not be is the only control: filtering after a top-k search means the legitimate user's slots were spent on documents that get thrown away, so recall degrades quietly as the corpus grows.
- How does a semantic cache break tenant isolation?By keying on question similarity alone. Two clinicians at different facilities ask a near-identical question, the cache scores a hit, and one facility's grounded answer is served to the other — bypassing retrieval and therefore bypassing the filter entirely. The tenant, and any role attribute in the filter, must be part of the cache key, and the same applies to any prompt or answer cache.
- An agent has a retrieval tool. Where does the tenant identifier come from?From the server-side session that executes the tool, never from a tool argument the model fills in. If the model can supply or alter the filter, untrusted text inside a retrieved document can influence it, and the boundary evaporates. Keep the parameter out of the schema entirely so there is nothing to set.
- How would you handle documents whose permissions change upstream after indexing?Either re-check the returned document identifiers against the system of record before they enter the prompt, or bound the staleness with a fast sync and state the accepted window explicitly. Deletions need particular care: removing the source document must also remove the chunks, the derived summaries and any cached answers built from it.
saying these in an interview costs you the question
- Relies on a system-prompt instruction to keep tenants apart
- Filters results only after the model has seen the chunks
- Takes the tenant identifier from a client-supplied parameter
- Assumes a reindex preserves permission metadata automatically
- Forgets summaries, caches and eval sets inherit tenant sensitivity