A red-team run against a mailbox-and-documents assistant uses a freshly created empty test mailbox and a brand-new account with no history or shared drives. Why can this understate real risk, and what do you put in that account before running?
answer
- empty tenant = false negative
- seed decoys, roles, history, volume
- one obvious crown jewel
- canary token doubles as detector
- never copy production data in
basics
~20 sAn empty account gives an exfiltration attempt nothing to steal and no permissions to abuse, so the run scores a false negative. Seed it with realistic decoy documents, contacts and message history, grant the same roles a real user holds, and mark every seeded item with a canary string you can search for.
solid answer
~50 sContainment and fidelity pull against each other, and an over-sterilised environment fails quietly. **Why the empty account lies.** Half of an agent's risk surface is the data and reach it inherits from the user. No history, nothing to leak; no shared drive, no confused-deputy read. The attack can work perfectly and still produce a null result, written up as 'no finding'. **What to seed.** Decoy documents including one that looks worth stealing, contacts, threaded history so retrieval has substrate, and roles matching a representative user. Volume matters: retrieval behaves differently over three documents than three thousand. **The safety rule.** Decoys are synthetic. Never copy real records in for realism — that creates a second copy of the data you were hired to protect, in an account you plan to hand to a hijacked agent. **The bonus.** Canary strings in the decoys give you exfiltration detection for free: if the token turns up in an outbound body or a log, the read happened.
go deeper
Recognises that an empty mailbox gives an attack nothing to steal and that fake test data should be added.
Explains the false-negative mechanism, and seeds content, roles, history and volume with synthetic decoys plus canary tokens.
Matches the account to a representative persona, notes retrieval behaviour changes with corpus size, and refuses production data even when asked for realism.
Makes the seeded tenant a maintained asset with defined personas, so results are comparable between runs and a null result carries a documented environment.
Provisioning has two failure directions, and interviews mostly probe the quieter one. **Over-exposure** is the failure everyone anticipates: the test account holds real data or real power, an attack lands, and the test causes the harm it was meant to measure. **Under-exposure** is the failure that reads as a success. The account is so barren that no attack can demonstrate impact. The run comes back clean, the report says the agent resisted injection, and the conclusion is wrong — you measured a system with nothing to lose. It is the environment-side twin of an always-success stub: there the instrumentation fakes a result, here the environment fakes a *non*-result. ### Why an empty account cannot produce a positive Roughly half an agent's risk surface is inherited from the user it acts for: the documents it can read, the mailboxes and drives it can reach, the groups whose permissions it borrows. Strip those and the mechanics of the interesting attacks have nowhere to run. No history means an exfiltration attempt has no target; no shared drive means a confused-deputy read has nothing to cross into; an empty memory or retrieval store means memory- and retrieval-mediated behaviour cannot appear at all. The injection can work perfectly and still score zero. ### What to seed - **Content** — synthetic but structurally real: invoices, a contract, an HR-shaped document, a credential-shaped string that is not a live credential. Include at least one plausible crown jewel so a model choosing what to take has something to choose. - **Reach** — the roles, group memberships and shared-drive access of a *representative* user. Not an administrator, which inflates every finding, and not a first-day joiner, which suppresses them. Reach is usually what converts a prompt injection into an incident. - **History** — threaded mail and prior conversation turns, so retrieval and memory paths have substrate to work over. - **Volume** — enough items that retrieval ranking actually has to rank. Three documents make retrieval look far more precise than it is over three thousand. - **Canaries** — a unique token per seeded document. Cheap, and it doubles as a detector: seeing that exact token in an outbound body, an external log or a returned answer is objective evidence that the read and the egress both happened. ### What it costs A decoy tenant is a build, not a config change: expect one to three engineer-days for the first one across a mail, documents and tracker surface, plus a few dollars of model calls to generate several hundred plausible documents. The recurring cost is maintenance — permissions drift, the SaaS changes its sharing model, and a stale decoy tenant quietly stops representing the population you claim to be testing. Budget the refresh, or accept that the environment is an uncontrolled variable and say so in the report. ### Where the number misleads Three distinct ways. **A null result over a sparse tenant** is uninformative, not reassuring, yet it is written up as "no exfiltration observed" and read as "exfiltration is not possible". Two null results — one over a bare account, one over a realistic one — are entirely different claims and must not share a row in a trend table. **An inflated success rate over a theatrical tenant.** The opposite error. Seed one document literally named so as to advertise itself as the secret and put it at the top of an otherwise tiny corpus, and almost any injection will retrieve it; the attack-success rate approaches 100% and is measuring your seeding, not the agent. Realistic corpora bury the crown jewel among neighbours, which is what production looks like. **Canary matching as a precise-looking detector.** Exact-string matching only catches verbatim egress. An agent that paraphrases, summarises, translates or partially quotes the decoy has still leaked it and will not trip the match, so a canary-based exfiltration rate is a *lower bound*. Treat a hit as proof and a miss as unproven, never as proof of absence. ### What you check Read the test account's effective permissions back from the API instead of trusting the provisioning form. Verify every canary is unique per document, so a hit tells you *which* document was read. Confirm no seeded credential-shaped string is live anywhere. Remember the decoys will pass through the model provider's systems like any other prompt content, so they must contain nothing you would not send there. And the hard line: never seed with copied production data. It creates a second copy of the very data you were hired to protect, inside an account you are about to hand to a deliberately hijacked agent, and it usually breaks the engagement's own data-handling terms. If realism truly demands production-shaped records, generate or mask them and record which fidelity you gave up. Benchmarks that ship their environment with the tests, such as agentdojo, make this variable explicit by construction; when you build your own tenant, it is uncontrolled unless you write down what was in it.
- How do canary strings in seeded decoys change how you score an exfiltration attempt?They move scoring from judging a transcript to matching a token. If the string appears in an outbound body, an external log or a returned answer, the read and the egress both happened — an objective check instead of a grader's opinion.
- What roles should the test account hold?Those of a representative user in the population you are assessing, not an admin and not an empty new joiner. Admin roles inflate every finding; a bare account suppresses them. If you test several personas, provision one account per persona and report per persona.
Testing an agent over a brand-new empty account is like testing a burglar alarm in an empty house: nothing was taken, but that is a fact about the house, not about the alarm.
saying these in an interview costs you the question
- Reporting 'no exfiltration possible' from a run over an empty mailbox.
- Copying real customer or employee data into the test tenant for realism.
- Giving the test account admin roles because it was easier than matching a real user.
- Seeding three documents and drawing conclusions about retrieval behaviour.
- No canary tokens, so exfiltration must be judged from transcript prose.