skip to content

You are planning a red-team engagement against an agent wired into production systems. How do you decide, tool by tool, which actions execute for real and which are stubbed, knowing stubs weaken your evidence and real execution leaves residue?

level: principalimportance: should knowfreq 34%

answer

  1. two axes: reversibility, externality
  2. vendor dry-run mode when it exists
  3. spend live-fire where severity is disputed
  4. written mode table, published with findings
  5. stubbed twice = coverage gap, report it

basics

~20 s

Split by reversibility and by who else sees the effect. Reads and reversible writes inside a scratch tenant run for real. Anything reaching a third party, moving money, or paging a human gets stubbed, and you argue that gap separately. Decide before the run, write the table down, publish it with the findings.

solid answer

~50 s

Treat it as a per-tool decision with two axes. **Reversibility.** Can you delete the artefact and verify the deletion? A row in a tenant you own: yes. A delivered SMS: no. Irreversible effects default to stubs. **Externality.** Does the effect leave the boundary you were authorised over — a third-party recipient, a partner API, an on-call page? Anything that can start someone else's incident is stubbed, however reversible it looks. **Evidence value.** Only a live execution shows whether the tool's own authorisation check would have refused. Where severity will be contested, buy that link with one controlled execution, not a live-fire campaign. **Make it a written table** — tool, mode, rationale, cleanup owner — agreed before the run and reproduced in the report, so every finding's evidence level is legible. Re-test stubbed high-value paths in a staging copy; if none exists, that absence is a finding. The failure mode is deciding mid-run, which yields findings nobody can grade and residue nobody expected.

go deeper

for a junior

Says dangerous or irreversible actions should be faked and safe reads can run for real.

for a middle

Separates reversible from irreversible and internal from externally visible, and knows a stub costs evidence about downstream controls.

for a senior

Produces a written per-tool mode table with cleanup owners, uses vendor dry-run modes where they exist, and buys a single live execution to settle contested findings.

for a principal

Treats the table as a durable artefact: budgets live-fire where evidence is scarce, keeps it frozen so runs stay comparable, and reports persistently untested tools as a coverage gap.

This is a portfolio decision, not a per-attack one, and the entire discipline is making it before anything runs. ### Classify every tool the agent can reach Including the ones it reaches indirectly — a tool that calls another API, a code-execution tool, a callback the target registers — which is the set teams routinely under-count. For each tool answer five questions: what does it change, where does the change land, can it be undone and verified undone, who else notices, and what does a stub cost me in evidence? ### Four default modes | Mode | Use for | Evidence | Residue | |---|---|---|---| | Live in a scratch tenant | reads, and writes whose artefacts you can find and delete | highest | manageable, ledgered | | Vendor dry-run / test mode | providers that ship one (test keys, sandbox send) | high — real authorisation and validation, no delivery | none | | Recording stub | irreversible or externally visible effects: payments, real-recipient messaging, deletes of data you do not own, deploys, anything paging a human | model-layer only | none | | Blocked outright | tools whose mere invocation is unacceptable | the attempt itself is the record | none | The vendor test mode is the underused one. It exercises the provider's real server-side authorisation and schema validation — precisely the links a stub cannot reach — while suppressing delivery, so where it exists it is strictly better than a home-made stub. ### What it costs Stubs are cheap and weak: an hour each, no residue, evidence that stops at the model. Live-fire is expensive and strong: the provisioned tenant, the ledger, a staffed cleanup pass, delayed verification searches, and blue-team time spent on alerts your run generated. Because live-fire is priced in people rather than tokens, budget it like a scarce resource: spend it on the handful of paths where the decisive control is server-side and invisible from the client, and on the paths whose severity the owner will dispute. Everything else can be stubbed provided the report says so. ### Where the number misleads **Mode-table drift destroys trend lines.** If last quarter's mail tool was stubbed and this quarter's ran live against a scratch tenant, a rise in findings is fully explained by the instrumentation change. Findings that always existed became visible; nothing regressed. Freeze the per-tool mode table between runs, and when it must change, report the delta separately from the trend rather than folding it in. **"Coverage" hides three denominators.** Percentage of the agent's tools exercised live, percentage of attack techniques attempted, and percentage of the agent's reachable surface touched are different numbers, and the ambiguous word lets a report claim the flattering one. Say which denominator you mean, every time. **A blanket attack-success rate mixes modes.** Aggregating stubbed and live tools into one percentage produces a figure that answers no question; report per mode. ### What you check The mode table is itself a security artefact, and reading it is part of the job. A tool that nobody was willing to test live two runs running is a tool whose control has never been exercised, and that belongs in the report as an explicit coverage gap rather than a silent omission. Before the run, confirm each live tool has a named cleanup owner and a verified delete path; confirm each stub's return shape, including error shapes, matches the real API; and confirm the abort signal with the system owner. When a stubbed finding is disputed, escalate deliberately: one scoped execution, announced, with the owner watching and the abort signal live. Never an unannounced live attempt to win an argument — that converts a professional engagement into the incident you were hired to prevent. Finally, place the benchmarks correctly. Simulated-environment suites — agentdojo, agentharm and injecagent among them — dissolve this dilemma by shipping the environment with the tests, which buys repeatability and zero residue at the price of saying nothing about your production connectors, their real authorisation checks, or their fan-out. They are a regression signal beside the engagement, not a substitute for the mode table.

  • Which tools deserve the expense of live execution?
    The ones where the decisive control is server-side and invisible from the client, and the ones whose severity the owner will dispute. A stub can never tell you whether the API would have refused; for everything else the recorded call is enough.
  • How does the stub/live split affect comparing this run with the previous one?
    It dominates the comparison. Changing a tool from stubbed to live can add findings with no change in the system's security. Freeze the mode table between runs, and if it must change, report the deltas separately from the trend.

saying these in an interview costs you the question

  • Deciding live-versus-stub per attack while the run is in flight.
  • Executing against a third party or a real recipient to 'prove' a disputed finding.
  • Full live-fire everywhere with no cleanup owner staffed.
  • Changing the mode table between quarterly runs and still drawing a trend line.
  • Omitting from the report which tools were never exercised live.

context