Before red-teaming a deployed assistant that can send email, file tickets and write CRM rows, what do you provision so that a successful attack does not touch real people or real records?
answer
- throwaway principal, not the operator's
- scratch tenant you can delete
- dry-run or test egress endpoints
- per-run marker on every artefact
- stub = safety bought with proof
basics
~20 sProvision throwaway identities the agent acts as: a scratch mailbox, a test tenant or sandbox project, and disposable CRM records. Point outbound tools at dry-run or test endpoints. Mark everything with a unique per-run tag so you can find it later, and agree who owns cleanup.
solid answer
~50 sContainment is set up **before** the first attack attempt, not improvised when something lands. The checklist an interviewer wants to hear: - **Identity** — the agent authenticates as a throwaway principal, never a real employee or the operator's own account. That principal's blast radius is the ceiling on any hijack. - **Data plane** — a scratch tenant, sandbox project or test mailbox, so writes land somewhere you own and can drop wholesale. - **Egress** — outbound tools (mail, SMS, payments, webhooks) point at test endpoints, a dry-run mode, or a local sink. Recipients you control, or none. - **Markers** — every record, subject line and file the test creates carries a unique per-run string, so cleanup is a search rather than an archaeology dig. The tradeoff to name out loud: the more you stub, the less your evidence says about the real system. A stubbed mail tool proves the agent *decided* to send, not that mail would have left the building.
go deeper
Names throwaway accounts, a test mailbox and pointing sends at a test endpoint, and knows real user data should not be in scope.
Adds the credential as the real boundary, per-run markers for cleanup, and the stub-versus-real tradeoff per tool.
Enumerates the paths that escape the harness — webhooks, downstream sync, notifications, audit logs — and pre-agrees an abort signal and cleanup owner with the system owner.
Turns it into standing practice: a required test tenant per connected system, a documented live-fire policy, and treating 'no safe environment exists' as a reportable gap rather than a reason to improvise.
Containment for an agent red-team run is not a switch in the harness. It is a set of accounts, endpoints and markers you create **before the first prompt goes out**, chosen so that every side effect the agent can produce lands somewhere you own, can find, and can delete. The harness — your own wrapper around the agent's tool loop, or a benchmark runner — only constrains calls that pass through it. The credential the agent carries is what the outside world honours, and that is a different object. ### The four layers you provision **1. Identity.** The agent authenticates as a principal created for this engagement: a service account or a test user in a directory you control, never the operator's own SSO session and never a real employee's account. The roles on that principal are the ceiling on the whole run. "The harness only exposed three tools" is not a boundary — a model can reach further by chaining a tool that itself calls another API, by following a callback, or by writing code that a code-execution tool runs, and every one of those paths carries the same token. **2. Data plane.** A scratch tenant, sandbox project or test mailbox: a container you can drop wholesale afterwards. The test is whether you personally hold the permission to delete it, not whether someone believes it is disposable. **3. Egress.** Decide per outbound tool, in advance, which of three modes it runs in: real execution against your scratch tenant; the vendor's own test or dry-run mode where one exists (test API keys, a mail provider's sandbox that accepts and discards); or a recording stub inside your harness that logs the call and returns a canned result. Write the choice down per tool — that table becomes the caveat section of the report. **4. Markers.** A unique per-run token goes into every free-text field the test can create: mail subjects, record names, file bodies, ticket titles. This is the cheapest item on the list and the one most often skipped, and without it post-run cleanup degrades to searching by timestamp, which collides with real traffic. ### What it costs The expensive resource here is engineer time, not inference. A first scratch tenant for a mail-and-documents assistant is typically half a day to two days of work plus a ticket to whoever owns the identity provider; each additional connected system (tracker, CRM, object store) adds hours and often a paid seat. The model calls in an agent engagement are trivial by comparison — a few hundred to a few thousand calls, single-digit to low-tens of dollars. On a week-long engagement the containment setup, not the tokens, is the budget line, which is precisely why teams talk themselves out of it. ### Where the number misleads Two readings of a contained run go wrong, in opposite directions. The first is a **clean run read as a safe agent**. A freshly provisioned tenant usually has no history, no shared drives and no crown jewels, so an exfiltration attempt has nothing to take and a confused-deputy attempt has no reach to abuse. The run returns zero findings, and the zero describes the *environment*, not the agent. Containment and fidelity pull against each other; any score is a property of the pair. The second is a **success count read as impact**. An attack-success rate computed over a run where most tools were stubbed counts model decisions, not realised effects. Published without the mode table it is uninterpretable and, worse, uncomparable: switch one tool from stub to live next quarter and the same system "gets worse" with no change in its security. ### What you check before the first attack - Read the throwaway principal's effective roles back from the API rather than trusting the provisioning console; defaults frequently grant more than the form suggested. - Confirm the mail path has a genuine sandbox mode. A "test" recipient address on the production relay still performs real DNS, real delivery and a real bounce to a real postmaster. - Enumerate fan-out: webhooks to partner systems, CRM-to-warehouse sync, notification rules that page an on-call human. Those escape the tenant you can delete. - Confirm you can delete the tenant, and that soft-deleted objects can be purged rather than left recoverable. - Agree a named cleanup owner, an abort signal with the system owner, and disclosure of run windows and markers to the blue team so your traffic is attributable. Simulated-environment benchmarks — agentdojo, agentharm and injecagent are of that class — dodge this entirely by shipping the environment with the tests: no provisioning, no residue, and a repeatable number. They also say nothing about your deployment's real connectors, so they complement an engagement rather than replacing one.
- Why is the agent's credential the real containment boundary rather than the harness's tool list?The harness only controls calls that go through it. Anything the model reaches by another path — a chained tool, a callback, code it writes — still carries the token. Scope the credential and the harness becomes defence in depth rather than the only defence.
- You cannot get a scratch tenant; the only environment is production. What changes?Shift heavily toward stubs and read-only probes, run in a low-traffic window, pre-agree an abort signal with the system owner, and label every write finding as unproven-at-the-write-layer. And put the missing test environment in the report as a finding of its own.
saying these in an interview costs you the question
- Running the agent with the operator's own credentials because 'it's only a test'.
- Treating harness-level tool scoping as containment when the token itself is broadly scoped.
- No per-run marker, so cleanup relies on timestamps and memory.
- Assuming a 'test' recipient address means no mail actually leaves.
- Deciding what to stub only after an attack lands.