skip to content

Why bind the tenant filter on RAG retrieval server-side, not in the prompt?

level: middleimportance: must knowfreq 58%

answer

  1. authorization belongs below the model
  2. untrusted text can steer any model-written argument
  3. tool takes a query string, nothing else
  4. per-tenant namespace beats a filter flag
  5. cache keys inherit the scope too

basics

~20 s

Anything the model can influence can be widened by text that reaches it. Derive the tenant scope from the authenticated session and bind it at the query layer, so the model contributes only a search string and never the filter that decides which corpus is readable.

solid answer

~50 s

A retrieval scope expressed as an instruction, or as a tool parameter the model fills in, is only as reliable as the model's compliance, and the model's context contains untrusted text from documents, tool results and user input. In a multi-tenant real-estate CRM assistant, the brokerage id must come from the authenticated session and be applied by the retrieval service itself, so a crafted query cannot widen the corpus beyond the caller's tenant. Practically that means the tool signature exposes only a query string, the filter is attached below the tool, and the strongest form is a separate namespace, collection or index per tenant so cross-tenant results are not merely filtered out but unreachable. Add a check at the result boundary that every returned chunk carries the expected tenant id, which catches mislabelled documents from ingestion rather than model misbehaviour. Cache keys and embedding caches must be scoped the same way.

code

sql · 5 lines
sql
SELECT chunk_id, content
FROM listing_chunks
WHERE brokerage_id = $1
ORDER BY embedding <=> $2
LIMIT 10;

go deeper

for a junior

Be able to say that which documents a session may read is an authorization decision, so it must come from the logged-in user's identity and not from anything the model writes.

for a middle

Explain why a model-populated filter argument is unsafe, how the scope is bound at the query layer instead, and why the tool schema should have no scope parameter at all.

for a senior

Bring the operational layers: per-tenant namespaces, a result-boundary check that catches ingestion mislabelling, tenant-keyed caches, and a test that a crafted cross-tenant query returns nothing.

for a principal

Own the tradeoff between per-tenant partitioning and a shared index with bound filters at your tenant count, and be clear that scoping bounds what is at risk while egress controls decide whether it can leave.

## The rule Authorization decisions belong below the model. A retrieval layer that decides which documents a session may read is making an authorization decision, so it must derive its scope from the authenticated principal, not from anything the model emits. This is the same principle that says a database query must not take its user id from a request parameter the client controls; the LLM is simply a new, unusually persuadable client. ## Three places the scope can live, and only one is safe **In the system prompt.** Telling the model to search only the current brokerage's listings is guidance. It survives ordinary use and fails under adversarial input, and it also fails under plain confusion on a long context. Nothing enforces it. **As a tool parameter.** Exposing a retrieval tool whose arguments include a tenant id or a filter expression looks more rigorous, because the value is now structured. It is not. The model populates it, the model's context contains text from retrieved documents and user messages, and a value the model writes is a value that untrusted text can steer. This shape is common because it is convenient during development, and it is the specific defect an interviewer is usually probing for. **Bound at the query layer.** The retrieval service receives the tenant identity from the session or a signed token, applies it to the query itself, and the tool signature simply has no parameter for scope. The model contributes a search string and nothing else. Now the worst an injected instruction can achieve is a differently worded search inside the same corpus. ## Filtering versus partitioning Binding a filter is the minimum. Stronger is physical or logical partitioning: a separate collection, namespace or index per tenant, selected by the same session-derived identity. The difference matters for two reasons. First, a partition failure is loud, whereas a filter failure is silent and returns plausible results. Second, a partition removes a class of bug where the filter is applied to metadata that ingestion set incorrectly; if the document is not in the tenant's index at all, a mislabelled metadata field cannot expose it. The cost is operational, since many small tenants means many small indexes, and per-tenant partitioning interacts with how the vector store handles memory and index rebuilds. That is a real tradeoff to state rather than wave away. ## Verify at the result boundary After retrieval returns, check that every chunk carries the expected tenant identifier before any of it is placed in the context, and treat a mismatch as an incident rather than a warning. This is not redundant with the bound filter, because it catches a different failure: ingestion writing the wrong tenant on a document, a backfill that dropped the field, or a reindex that lost metadata. Those are ordinary data bugs, not attacks, and they are the most common real cause of cross-tenant leakage in retrieval systems. ## The paths teams forget Scoping the primary query is the easy part. The leaks tend to come from adjacent state that inherited no scope. Caches are the classic case: a semantic or embedding cache keyed on the query text alone will happily serve one tenant's cached results to another, so the tenant identity must be part of every cache key. Reranker and summarization stages that fetch additional context need the same scope. Long-lived agent memory written during one tenant's session and read in another's is the same bug in a different store. Evaluation and trace tooling that replays production retrievals sits outside the request path and often has no tenant concept at all. ## What this control does not do Scoping retrieval per tenant prevents the model from reading data the caller was never entitled to. It says nothing about what happens to data the caller *is* entitled to once it is in the context: that data can still leave through a rendered URL, an outbound tool call or a trace log. Scoping is the containment control that decides how much is at risk; the egress controls decide whether what is at risk can leave. A good answer names both and does not claim either alone is sufficient. ## How to demonstrate it The convincing version of this answer is operational. Show that the tool schema physically has no scope parameter, that the retrieval service reads the tenant from a verified session token, that there is a test asserting a query crafted to request another brokerage's listings returns nothing, and that the result-boundary check exists and is alerted on. An interviewer hears the difference between someone who has configured a filter and someone who has designed the scope to be unreachable from the model.

  • The filter is bound server-side, yet a customer still saw another tenant's listing text. What do you check first?
    Ingestion and indexing. A bound filter is only as good as the tenant metadata on the documents, so a backfill that wrote the wrong brokerage id or a reindex that dropped the field produces exactly this symptom with no attacker involved. That is why a result-boundary check comparing each returned chunk's tenant id against the session's belongs in the path, alerting rather than silently dropping.
  • Where do caches break tenant scoping?
    Any cache keyed on query text or embedding alone will serve one tenant's results to another, including semantic caches, embedding caches and reranker caches. The tenant identity has to be part of every cache key, and the same applies to agent memory written in one session and read in another. These are the leaks that survive a correctly scoped primary query.
  • Is per-tenant partitioning always better than a filter?
    It is stronger, because a mislabelled document cannot be reached at all rather than merely being filtered out, and failures are loud instead of silent. But it costs operationally: many tenants means many indexes, with memory, rebuild and cold-start consequences that vary by store. The honest answer is that partitioning is the default for high-sensitivity corpora and small tenant counts, with a bound filter plus a result-boundary check as the pragmatic alternative at scale.

The model is a researcher with a library card. You decide which shelves the card opens; you do not ask the researcher to promise to stay in one aisle.

saying these in an interview costs you the question

  • Putting the tenant restriction in the system prompt
  • Exposing a filter or tenant-id argument the model populates
  • Assuming a bound filter is safe without checking ingestion metadata
  • Caching retrieval results on query text without the tenant in the key
  • Claiming retrieval scoping alone prevents data exfiltration

context