How do scope and metadata filters on agent memory reads prevent cross-user leakage?
answer
- similarity does not know who owns a row
- bind it from the session, not the model
- filter inside the query, not after
- the shared tier is where writes go wrong
- canary memories, asserted in CI
basics
~20 sEvery memory read must be constrained by a scope key — user, tenant, session or project — derived from the authenticated session, applied inside the query before ranking. Semantic search has no notion of ownership, so nothing but an explicit filter keeps one user's memories out of another's context.
solid answer
~60 sSemantic similarity is ownership-blind: a query about allergies matches every patient's allergy note equally well, so the only thing separating them is a filter. Three rules make that filter trustworthy. First, **the scope key comes from the authenticated session, never from a model-authored tool argument**. If the agent passes the user id, prompt injection can change it; the harness must bind it from the request context and reject or ignore any model-supplied override. Second, **filter inside the query, not after ranking**. Post-filtering leaks in two ways: the results still transit logs and traces, and a scoped search that returns k after filtering silently degrades to fewer results, which people then fix by widening the search rather than the filter. Third, **make the shared tier explicit**. Most systems have per-user memories plus organization-wide knowledge, and the bug is usually a memory written to the wrong tier rather than a broken filter. Then test it: seed canary memories under one identity and assert in CI that no query under another identity can retrieve them.
code
python · 15 linesdef similarity(query, text):
q, t = set(query.lower().split()), set(text.lower().split())
return len(q & t) / (len(q | t) or 1)
STORE = [
{"subject": "patient-a", "scope": "clinical", "text": "penicillin allergy on file"},
{"subject": "patient-b", "scope": "clinical", "text": "no known allergies"},
]
def recall(store, query, *, subject, scope, limit=5):
allowed = [m for m in store if m["subject"] == subject and m["scope"] == scope]
return sorted(allowed, key=lambda m: -similarity(query, m["text"]))[:limit]
# subject comes from the session, never from a model-authored argument
print(recall(STORE, "any allergy on file?", subject="patient-b", scope="clinical"))go deeper
Know that every memory read must be limited to the current user or tenant, and that semantic search will happily return someone else's data if nothing constrains it.
Explain the difference between filtering inside the query and discarding results afterwards, name the scope dimensions a real store uses, and say why the shared tier needs to be explicit.
Demonstrate the trust boundary: scope bound from the authenticated session rather than a model-authored argument, per-tenant namespaces so a missing filter fails closed, and canary memories asserted in CI. Be ready to walk an incident from symptom to root cause.
Own isolation strategy across tenants, background jobs and subagents — shared index with filters versus physical partitioning, what auditability you owe, and how scope survives delegation in a multi-agent system.
## Why this is a retrieval problem and not just an access-control problem A memory store is queried by meaning. "Does this person have any allergies on file?" is close in embedding space to every allergy note ever written, for every patient, in every session. The store has no concept of who owns a row unless someone tells it. That makes memory retrieval a place where a conventional authorization mistake produces an unconventional consequence: the leaked data does not appear in an API response where a test might catch it, it appears inside a model's context, is paraphrased into an answer, and reaches the wrong person as fluent prose with no obvious provenance. The canonical incident shape is small: a clinic assistant retrieves an allergy note recorded under one patient while handling a different patient's booking, because the recall call carried the query but not the patient scope. Nothing crashed, nothing 500'd, and the only artefact is a sentence in a transcript. ## What scope actually means Scope is rarely one dimension. A production memory store usually partitions along several at once: - **Tenant or organization** — the hard boundary; crossing it is a reportable incident. - **User or subject** — whose memories these are, which for an assistant acting on behalf of staff is not the same as who is operating the agent. - **Session or thread** — scratch memories that should not outlive the conversation. - **Project, case or entity** — the working context, such as one patient file or one incident. - **Agent or role** — which agent in a multi-agent system may read this. And metadata beyond identity matters for the same reason: sensitivity class, source, consent state and retention flag all constrain what may be read into a given context even when the identity checks pass. A memory the user asked to be forgotten should be unreachable by every read path, not merely absent from the UI. ## The three rules that make the filter trustworthy **Bind scope from the session, not from the model.** This is the rule people get wrong when they move to agentic recall. Exposing `search_memory(query, user_id)` as a tool schema puts the security-critical parameter in the model's hands, where any injected instruction sitting in retrieved content or tool output can set it. The harness should accept only the query from the model and inject the scope itself from the authenticated request. If the agent legitimately needs to work across subjects, model that as an explicit, audited capability with its own allowlist, not as a free parameter. **Filter inside the search, not after it.** Pre-filtering — restricting the candidate set before ranking, or querying a per-tenant namespace or index — keeps foreign rows out of the result path entirely. Post-filtering fetches k globally and discards non-matching rows afterwards, which leaks into logs, traces and eval dumps, and quietly returns fewer than k useful results, tempting an engineer to raise k rather than fix the scope. Strong isolation goes further and gives each tenant its own namespace so a missing filter fails closed with no results rather than open with someone else's. **Make the shared tier explicit.** Almost every system has global material — policies, product facts, procedures — alongside per-user memories. Model that as a named scope the read explicitly opts into, so "which tier is this memory in?" is answerable, and so a write cannot land in the global tier by default. In practice the most common leak is not a broken filter at all: it is a memory extracted from one user's conversation and written to a shared or wrongly-keyed scope, after which every filter behaves correctly and the data is still in the wrong place. ## Multi-agent and tool-chain amplification Scope has to survive delegation. When an orchestrator spawns subagents, each must carry the same scope constraint, and a subagent's summary returned to the parent inherits whatever it read. A background or scheduled agent running without an interactive session is the classic hole: with no request context to bind from, someone hardcodes a broad scope and the job reads across users by construction. ## Proving it, not asserting it This is the part interviews reward. Seed **canary memories** — distinctive, synthetic facts under a known identity — and run assertions in CI that queries under every other identity, including adversarial ones that name the canary directly, return nothing. Add a runtime invariant: every read carries a scope key, and reads without one are rejected rather than defaulted. Log the scope on every retrieval so a transcript can be audited after the fact. And test the injection path explicitly, by planting content that instructs the agent to look up another subject and confirming the harness ignores it, because that is the failure the tool schema invites and the one a functional test will never find.
- Your memory tool schema takes a user_id parameter. Why is that a problem and what replaces it?Any parameter the model fills is attacker-controllable, because injected instructions can reach the model through retrieved content, tool output or a document. A tool that accepts user_id lets an injection redirect the search to another subject and the filter will faithfully honour it. Replace it by having the harness inject the scope from the authenticated session and accept only the query from the model. If cross-subject reads are a real requirement, make that a separate, allowlisted, audited capability.
- Why is post-filtering results after the search worse than filtering inside it?Two reasons. Foreign rows are fetched, so they exist in query results, logs, traces and eval dumps even if the agent never sees them, which is still a disclosure in most regimes. And a global top-k that is filtered afterwards returns fewer than k in-scope results, so the ranking silently degrades and the natural fix looks like raising k rather than fixing the query. Pre-filtering or per-tenant namespaces avoid both.
- How would you catch a scope bug before a user does?Seed canary memories: distinctive synthetic facts written under one identity, then assert in CI that no query under any other identity retrieves them, including queries that name the canary text directly. Add a runtime invariant rejecting any read that arrives without a scope key, so a missing filter fails closed. Log the scope on every retrieval for post-hoc audit, and include a prompt-injection case that tells the agent to look up another subject.
saying these in an interview costs you the question
- Vector similarity will not match another user's data anyway
- Filtering after retrieval is equivalent to filtering before it
- The model can be trusted to pass the right user id
- One shared index is fine as long as every query includes a filter
- Background jobs do not need a scope because no user is present