skip to content

A client asks you to red-team an assistant whose tools can send email, file tickets and call a partner company's API on a user's behalf. How do you scope what the assistant may actually do during the test, and what goes in writing?

level: seniorimportance: should knowfreq 48%

answer

  1. scope by blast radius, not by host
  2. tool inventory with credential scopes
  3. live / sandbox / recording stub tiers
  4. test-only mailboxes and accounts
  5. who cleans up what the agent created

basics

~20 s

The assistant's tools reach systems your client does not own, so a successful injection could send real mail or write to a partner. Agree in writing which tools stay live, which are stubbed, which recipients and accounts are test-only, and who you call if an action escapes into a real third-party system.

solid answer

~60 s

Scope an agentic target by its **blast radius**, not by its URL. List every registered tool and, for each, answer three questions: what identity does it act as, whose system does it touch, and is the effect reversible. Anything that touches a system the client does not own — a partner API, an outbound mail relay, a payments or ticketing SaaS — needs its own decision, because the client cannot authorise you to write into someone else's estate. The usual settlement is tiered. Reversible tools inside the client's estate run live. Tools that reach outside are pointed at a sandbox tenant, stubbed with a recorder that proves the call *would* have been made, or restricted to allow-listed test recipients. Write down the tool inventory with its tier, the test accounts and mailboxes, the credentials' scopes, a hard stop condition, and the contact who can pull the plug. Also write down what a stub proves: 'the model was induced to invoke the transfer tool with attacker-chosen arguments' is a real finding without moving any money.

go deeper

for a junior

Recognises that an assistant with tools can cause real side effects and that some tools should not be live during a test.

for a middle

Builds a tool inventory, distinguishes reversible internal effects from external ones, and uses test-only accounts and mailboxes.

for a senior

Tiers tools into live, sandboxed and recording-stub, reads credential scopes rather than tool descriptions, and negotiates stop conditions, contacts and cleanup ownership in writing.

for a principal

Sets the standing rule that a client cannot authorise writes into a third party's estate, and defines what evidence the firm accepts in place of executing a harmful action.

## Why an agent breaks the usual scope form A conventional scope document lists hosts: attack these addresses, not those. An assistant with tools cannot be scoped that way, because the thing you are attacking is not the address — it is a **decision loop**: 1. The model receives a prompt, 2. chooses a *tool call* (a named function with arguments, emitted as structured output), 3. the application's dispatch layer executes it against some real system with some real credential, 4. and the result is appended to the context so the model can choose again. The **reachable set** is therefore the transitive closure of everything those credentials can do, iterated over as many turns as the loop allows. An indirect injection that lands in a retrieved document does not respect the boundary you drew on paper; it respects the boundary the credentials draw. ## Build the inventory before you attack anything For each registered tool record: - the tool name as the model sees it; - the identity and credential the dispatch layer uses to execute it; - the scopes on that credential; - the systems it can reach; - whether the effect is reversible; - and whether anyone outside the engagement notices when it fires. That table *is* the scope. If the client cannot produce it, stop — that gap is a finding in its own right, and a serious one, because it means nobody in the organisation knows what their assistant can do. ## Tiering, and what a stub actually proves Settle each tool into one of three tiers. - **Reversible and internal:** run live. - **Irreversible but internal** — deleting records, rotating a credential — run live only against seeded test data with a restore plan agreed first. - **Anything that touches an estate the client does not own** — a partner's API, an outbound mail relay to real addresses, a payments or ticketing SaaS — defaults to a sandbox tenant obtained from that third party, or to a *recording stub*: an interceptor at the dispatch layer that logs the tool name and the arguments the model was persuaded to pass, returns a canned result, and never contacts the real system. The stub's evidence is precise and worth stating precisely: "the model was induced to invoke the transfer tool with attacker-chosen arguments." That is the vulnerability. Executing the side effect adds no information about it and adds legal exposure inside somebody else's system. ## Where this misleads, in both directions - A stub that returns a **canned success** can manufacture a chain that could not happen. The model, told the transfer succeeded, proceeds to the next step on data that would never have existed, and your transcript shows a five-step compromise of which only step one was real. - Conversely, **"tool invoked" is not "attack succeeded"**: the real tool sits behind its own authorisation, and a call the model was tricked into emitting may be rejected by the downstream service for a reason that has nothing to do with the model. Both errors have the same root — the stub replaces the part of the system that decides consequences. So report the invocation as the finding, and report the downstream authorisation state as a separate, verified fact, never as an inference. A count of "successful tool abuses" that mixes the two is a number nobody can act on. ## What it costs Building the inventory is a day with an engineer who knows the deployment, and it is the **highest-value day** in the engagement. A sandbox tenant from a partner has procurement lead time measured in weeks, so raise it at scoping or you will not have it. Stubbing is cheap to implement and cheap to run. The cost of getting it wrong is not billed in tokens: it is a partner's ticket queue full of adversarial test items, a real customer whose data an injection touched, and a relationship the client owns and you damaged. ## Test-only identities and stop conditions Dedicated mailboxes, a dedicated tenant, an allow-list on the outbound relay, and credentials scoped down for the engagement. Never run as a real customer's account: consent from the client is not consent from that customer. Agree in advance what halts the test: - an action lands in a real third-party system, - real customer data appears in output, - or the assistant does something irreversible you did not predict — and name a person on each side who can be reached within a stated time. Agree, too, who cleans up the artefacts the assistant creates. ## What you check when scoping is done Re-read the **credential scopes**, not the tool descriptions. A description says what a tool is *for*; a scope says what it *can do*. The gap between those two sentences is where the engagement's real risk lives.

  • The client says stubbing the outbound tools makes the test unrealistic. How do you answer?
    The realism that matters is whether the model can be induced to call the tool with attacker-chosen arguments; the stub captures exactly that. Executing the side effect adds no information about the vulnerability and adds legal and cleanup risk in someone else's system.
  • Mid-test the assistant files a ticket in the partner's real ticketing tenant. What now?
    Stop that line of testing, notify the client contact immediately, record the ticket identifier and timestamps, and let the client — who owns the partner relationship — decide how it is disclosed and withdrawn. Then fix the scope gap that let it through before resuming.

saying these in an interview costs you the question

  • Scopes the engagement as a URL and never inventories what the tools can reach.
  • Lets tools that write into a partner's system stay live because the client said 'go ahead'.
  • Uses a real customer's account or mailbox as the test identity.
  • Cannot describe evidence for a tool-abuse finding short of actually performing the harmful action.
  • Has no stop condition or reachable contact for an action that escapes into a real system.

context