In an agent red-team harness you replace the agent's real email-sending tool with a stub that records the call and returns success. An injection attempt now shows the agent calling that stub with attacker-chosen recipients and body. What does that result prove, and what does it not?
answer
- stub proves decision, not delivery
- evidence chain cut at the model
- downstream authz, allowlist, confirmation untested
- always-success hides the error branch
- one live re-run settles severity
basics
~20 sIt proves the model was hijacked into deciding to send, and shows the exact arguments it chose. It does not prove delivery: the real endpoint might reject the recipient, demand a confirmation, or fail an authorisation check. The stub tests the model's decision, not the system's outcome.
solid answer
~50 sA stub cuts the evidence chain at the model boundary. **What you have proven:** the injected content reached a place the model trusts, the model followed it, and it produced concrete attacker-controlled arguments. That is a real prompt-injection finding and it is reportable on its own. **What you have not proven:** that the action would land. Everything downstream of the stub is untested — the API's own authorisation check, a recipient allowlist, a rate limit, a human confirmation dialog, a DLP scan on the outbound body. Any one of those could turn your critical finding into a defence-in-depth win. **Why it still matters:** those downstream controls are usually undocumented and often absent. 'Model fully hijacked; send path stubbed and not exercised' is honest and still actionable — a defender who cannot name the control that would have stopped it does not have one. When severity is contested, re-run that one path against a scratch tenant and a recipient you own.
go deeper
Says the stub means no real email went out, and that the agent still followed the injected instruction.
Separates the model's decision from the system's outcome and names the untested downstream controls: authorisation, allowlists, confirmations, DLP.
Adds that an always-success stub distorts multi-step behaviour, and plans a single targeted live re-run to settle contested severity.
Sets the reporting convention: instrumentation disclosed with every finding, severity phrased as which layers were exercised, so stubbed evidence is neither over-claimed nor thrown away.
A stub is a function you register in the harness under the same tool name and JSON schema the model already sees, which records the call and returns a canned result instead of forwarding it. The model cannot tell the difference: to it, the tool exists and it worked. That property is exactly what makes a stub safe and exactly what makes its evidence partial. ### The evidence chain, link by link A live-fire hit demonstrates six links in sequence: 1. untrusted content reaches the model's context (a retrieved document, a tool result, an email body); 2. the model treats that content as instruction rather than data; 3. the model emits a tool call with attacker-chosen arguments; 4. the tool layer in front of the real API authorises the call; 5. the external system accepts it; 6. an observable effect exists in the world. A stubbed run establishes links 1–3 and asserts nothing whatsoever about 4–6. That is the whole answer, and everything below is what follows from it. ### What links 4–6 might have contained The untested region is usually where the undocumented controls live: the API's own server-side authorisation check on the calling principal, a recipient allowlist or external-domain block, a rate limit, a human confirmation step in the send path, a DLP scan on the outbound body, an anti-abuse heuristic on the sending account. Any one of them could downgrade your critical finding to a defence-in-depth win. Equally, none of them may exist — and a defender who cannot name the control that would have stopped it does not have one. Both possibilities are live until someone exercises the path. ### The second distortion: the missing error branch A stub that always returns success does more than hide the outcome; it changes the run. Agents branch on tool results. A real permission error may make the agent retry with different arguments, escalate to the user, or abandon the task; a synthetic success walks it straight down the happy path. So multi-step traces gathered behind an optimistic stub are unrepresentative in both directions — they can overstate how smoothly an attack completes, and they can hide adaptive behaviour that only appears after a refusal. Returning realistic error shapes for some fraction of calls is a cheap corrective, and the fraction you chose belongs in the write-up. ### What it costs A recording stub costs an hour to write and produces zero residue, which is why it is the default. Buying back links 4–6 is the expensive part, and the options are not equally priced. Reading the tool's server-side authorisation code is free if you can get the repository. Asking the owning team what would have happened costs a meeting and yields a *claim*, which you record as a claim. Executing the path once against a scratch tenant and a recipient you control costs the provisioning already discussed plus one artefact to clean up, and is usually the right answer. Executing against production with the owner watching is the last resort and needs an agreed abort signal. Note the shape: one targeted execution, not a live-fire campaign, is what settles a disputed severity. ### Where the number misleads The reporting failures are symmetric and both are common. **Over-claiming:** "the agent exfiltrated the customer list to an attacker-controlled address" when the mail tool was a stub and the real relay only accepts internal recipients. One engineer who knows that discounts the entire report, including the findings that were sound. **Under-claiming:** dropping the finding because "no real send happened", when a fully hijacked model plus an unknown downstream control is precisely the risk the engagement existed to surface. There is a quieter numeric failure too. An attack-success rate aggregated across a run mixes tools that executed for real with tools that were stubbed, so the single percentage answers no question anyone actually has. Report the rate per instrumentation mode, or do not report it as one number. ### What you check State the instrumentation next to every finding: which tools executed for real, which were stubbed, what each stub returned, and whether error injection was used. Prefer severity phrased as coverage — "model-layer control failed; send path stubbed and not exercised" — over a single ordinal, because the ordinal invites an argument about a fact you did not establish. Before the run, confirm the stub's return shape matches the real API's, including its error shapes; a stub whose success object differs from the real one can itself change the agent's next step and produce behaviour the real system would never show.
- How can an always-returns-success stub change the agent's later behaviour, not just hide the outcome?Agents branch on tool results. A real permission error might make it retry, escalate or abandon; a synthetic success sends it straight down the happy path. Multi-step traces gathered behind an optimistic stub can therefore be unrepresentative in both directions.
- The system owner disputes the severity because 'our API would have blocked that recipient'. What do you do?Ask for the control, read it if you can, and then execute that one path for real against a scratch tenant and an address you own. One live execution converts an argument about hypotheticals into a recorded fact, and costs one artefact to clean up.
A stub is like testing an evacuation by watching someone press an alarm button whose wiring has been cut. You learn that the person decided to raise the alarm; you learn nothing about whether any bell would have rung or any door would have unlocked.
saying these in an interview costs you the question
- Reporting 'mail was exfiltrated' when the send tool was a stub.
- Discarding a clear hijack finding because no real side effect occurred.
- Not stating in the report which tools were stubbed.
- Assuming an always-success stub leaves multi-step agent behaviour unchanged.
- Treating the stub's recorded call as proof the recipient address was reachable.