skip to content

Containing Side Effects

Throwaway accounts and dry-run endpoints are provisioned before the run, and an attack that lands leaves mail or rows you have to clear. Interviewers ask what your test left behind.

on this pageshow

explore

questions

5

Before red-teaming a deployed assistant that can send email, file tickets and write CRM rows, what do you provision so that a successful attack does not touch real people or real records?

level: juniorimportance: must knowfreq 62%

answer

  1. throwaway principal, not the operator's
  2. scratch tenant you can delete
  3. dry-run or test egress endpoints
  4. per-run marker on every artefact
  5. stub = safety bought with proof

basics

~20 s

Provision throwaway identities the agent acts as: a scratch mailbox, a test tenant or sandbox project, and disposable CRM records. Point outbound tools at dry-run or test endpoints. Mark everything with a unique per-run tag so you can find it later, and agree who owns cleanup.

solid answer

~50 s

Containment is set up **before** the first attack attempt, not improvised when something lands. The checklist an interviewer wants to hear: - **Identity** — the agent authenticates as a throwaway principal, never a real employee or the operator's own account. That principal's blast radius is the ceiling on any hijack. - **Data plane** — a scratch tenant, sandbox project or test mailbox, so writes land somewhere you own and can drop wholesale. - **Egress** — outbound tools (mail, SMS, payments, webhooks) point at test endpoints, a dry-run mode, or a local sink. Recipients you control, or none. - **Markers** — every record, subject line and file the test creates carries a unique per-run string, so cleanup is a search rather than an archaeology dig. The tradeoff to name out loud: the more you stub, the less your evidence says about the real system. A stubbed mail tool proves the agent *decided* to send, not that mail would have left the building.

go deeper

for a junior

Names throwaway accounts, a test mailbox and pointing sends at a test endpoint, and knows real user data should not be in scope.

for a middle

Adds the credential as the real boundary, per-run markers for cleanup, and the stub-versus-real tradeoff per tool.

for a senior

Enumerates the paths that escape the harness — webhooks, downstream sync, notifications, audit logs — and pre-agrees an abort signal and cleanup owner with the system owner.

for a principal

Turns it into standing practice: a required test tenant per connected system, a documented live-fire policy, and treating 'no safe environment exists' as a reportable gap rather than a reason to improvise.

Containment for an agent red-team run is not a switch in the harness. It is a set of accounts, endpoints and markers you create **before the first prompt goes out**, chosen so that every side effect the agent can produce lands somewhere you own, can find, and can delete. The harness — your own wrapper around the agent's tool loop, or a benchmark runner — only constrains calls that pass through it. The credential the agent carries is what the outside world honours, and that is a different object. ### The four layers you provision **1. Identity.** The agent authenticates as a principal created for this engagement: a service account or a test user in a directory you control, never the operator's own SSO session and never a real employee's account. The roles on that principal are the ceiling on the whole run. "The harness only exposed three tools" is not a boundary — a model can reach further by chaining a tool that itself calls another API, by following a callback, or by writing code that a code-execution tool runs, and every one of those paths carries the same token. **2. Data plane.** A scratch tenant, sandbox project or test mailbox: a container you can drop wholesale afterwards. The test is whether you personally hold the permission to delete it, not whether someone believes it is disposable. **3. Egress.** Decide per outbound tool, in advance, which of three modes it runs in: real execution against your scratch tenant; the vendor's own test or dry-run mode where one exists (test API keys, a mail provider's sandbox that accepts and discards); or a recording stub inside your harness that logs the call and returns a canned result. Write the choice down per tool — that table becomes the caveat section of the report. **4. Markers.** A unique per-run token goes into every free-text field the test can create: mail subjects, record names, file bodies, ticket titles. This is the cheapest item on the list and the one most often skipped, and without it post-run cleanup degrades to searching by timestamp, which collides with real traffic. ### What it costs The expensive resource here is engineer time, not inference. A first scratch tenant for a mail-and-documents assistant is typically half a day to two days of work plus a ticket to whoever owns the identity provider; each additional connected system (tracker, CRM, object store) adds hours and often a paid seat. The model calls in an agent engagement are trivial by comparison — a few hundred to a few thousand calls, single-digit to low-tens of dollars. On a week-long engagement the containment setup, not the tokens, is the budget line, which is precisely why teams talk themselves out of it. ### Where the number misleads Two readings of a contained run go wrong, in opposite directions. The first is a **clean run read as a safe agent**. A freshly provisioned tenant usually has no history, no shared drives and no crown jewels, so an exfiltration attempt has nothing to take and a confused-deputy attempt has no reach to abuse. The run returns zero findings, and the zero describes the *environment*, not the agent. Containment and fidelity pull against each other; any score is a property of the pair. The second is a **success count read as impact**. An attack-success rate computed over a run where most tools were stubbed counts model decisions, not realised effects. Published without the mode table it is uninterpretable and, worse, uncomparable: switch one tool from stub to live next quarter and the same system "gets worse" with no change in its security. ### What you check before the first attack - Read the throwaway principal's effective roles back from the API rather than trusting the provisioning console; defaults frequently grant more than the form suggested. - Confirm the mail path has a genuine sandbox mode. A "test" recipient address on the production relay still performs real DNS, real delivery and a real bounce to a real postmaster. - Enumerate fan-out: webhooks to partner systems, CRM-to-warehouse sync, notification rules that page an on-call human. Those escape the tenant you can delete. - Confirm you can delete the tenant, and that soft-deleted objects can be purged rather than left recoverable. - Agree a named cleanup owner, an abort signal with the system owner, and disclosure of run windows and markers to the blue team so your traffic is attributable. Simulated-environment benchmarks — agentdojo, agentharm and injecagent are of that class — dodge this entirely by shipping the environment with the tests: no provisioning, no residue, and a repeatable number. They also say nothing about your deployment's real connectors, so they complement an engagement rather than replacing one.

  • Why is the agent's credential the real containment boundary rather than the harness's tool list?
    The harness only controls calls that go through it. Anything the model reaches by another path — a chained tool, a callback, code it writes — still carries the token. Scope the credential and the harness becomes defence in depth rather than the only defence.
  • You cannot get a scratch tenant; the only environment is production. What changes?
    Shift heavily toward stubs and read-only probes, run in a low-traffic window, pre-agree an abort signal with the system owner, and label every write finding as unproven-at-the-write-layer. And put the missing test environment in the report as a finding of its own.

saying these in an interview costs you the question

  • Running the agent with the operator's own credentials because 'it's only a test'.
  • Treating harness-level tool scoping as containment when the token itself is broadly scoped.
  • No per-run marker, so cleanup relies on timestamps and memory.
  • Assuming a 'test' recipient address means no mail actually leaves.
  • Deciding what to stub only after an attack lands.

context

open as a page

In an agent red-team harness you replace the agent's real email-sending tool with a stub that records the call and returns success. An injection attempt now shows the agent calling that stub with attacker-chosen recipients and body. What does that result prove, and what does it not?

level: middleimportance: must knowfreq 55%

basics

~20 s

It proves the model was hijacked into deciding to send, and shows the exact arguments it chose. It does not prove delivery: the real endpoint might reject the recipient, demand a confirmation, or fail an authorisation check. The stub tests the model's decision, not the system's outcome.

open as a page

A red-team run against a mailbox-and-documents assistant uses a freshly created empty test mailbox and a brand-new account with no history or shared drives. Why can this understate real risk, and what do you put in that account before running?

level: middleimportance: should knowfreq 42%

basics

~20 s

An empty account gives an exfiltration attempt nothing to steal and no permissions to abuse, so the run scores a false negative. Seed it with realistic decoy documents, contacts and message history, grant the same roles a real user holds, and mark every seeded item with a canary string you can search for.

open as a page

After a live-fire agent red-team run in which the agent really sent mail and created records, how do you establish exactly what the test left behind and remove it — and what residue can you not remove?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Do not reconstruct it afterwards. Have the harness log every outbound tool call with the identifier the system returned, so the run produces a delete list. Tag artefacts with a per-run marker, delete, then re-search until the marker returns nothing. Delivered mail, webhook fan-out and audit entries stay.

open as a page

You are planning a red-team engagement against an agent wired into production systems. How do you decide, tool by tool, which actions execute for real and which are stubbed, knowing stubs weaken your evidence and real execution leaves residue?

level: principalimportance: should knowfreq 34%

basics

~20 s

Split by reversibility and by who else sees the effect. Reads and reversible writes inside a scratch tenant run for real. Anything reaching a third party, moving money, or paging a human gets stubbed, and you argue that gap separately. Decide before the run, write the table down, publish it with the findings.

open as a page