You are leading a red-team engagement against a team's live agent and you must intercept its tool calls to test the boundary the agent trusts. How do you decide how invasive that interception may be, and how do you keep the findings from being dismissed as 'your harness did that'?
answer
- agree the invasiveness envelope in writing
- least invasive placement that proves it
- control run + delivery check per finding
- list what the harness changed
- few reproduced behaviours beat many raw hits
basics
~20 sAgree up front how much of the stack you may modify, and write it down. Every layer you replace is a layer the owning team can blame. Prefer the least invasive interception that still proves the claim, ship a matched control run with each finding, and state exactly what the harness changed.
solid answer
~60 sTwo decisions, taken before the run. **How invasive.** Rank the options by how much of the real path survives: planting content in data a genuine tool reads changes nothing in the agent; a proxy in front of the endpoint changes the transport; a wrapper around the framework's tool dispatch changes the agent process itself. Pick the least invasive placement that can still produce the behaviour you need to demonstrate, and get the owning team to agree to it in writing along with the environment, the blast radius and the stop conditions. **How defensible.** Interception findings are disputable by construction, so pre-empt it: every reported behaviour ships with a matched control (same shim, benign content), a delivery check showing the payload reached the model, and a list of what the harness changed -- success codes, latency, ordering, result shape. Reproduce the ones that matter with the least invasive placement. The failure to avoid is a large hit count from a heavily modified stack with no controls, which the team correctly discounts.
go deeper
Should know interception must be agreed with the owning team and that the report has to say what the harness changed.
Should pick the least invasive placement that demonstrates the behaviour and attach a control run to each finding.
Should pre-empt the attribution dispute with controls, delivery verification and a least-invasive reproduction, and know which artefacts are still worth reporting as dependencies.
Owns the invasiveness envelope and the stop conditions, defends findings against the harness-caused-it objection by design, and keeps the report to reproduced behaviours the team can act on rather than a large raw hit count.
Interception is the only way to test the channel an agent actually trusts, and it is simultaneously the strongest argument the defending team has against your report. Leading the engagement means managing both sides of that at once: how much of their stack you are permitted to replace, and how you keep the findings from being attributed to the replacement. ### Set the invasiveness envelope before any run Rank the placements by how much of the real path survives, and get written agreement on where you may sit: | placement | what changes | attributability | |---|---|---| | plant content in data a real tool reads | nothing in the agent | strongest -- the whole real path ran | | proxy in front of the tool endpoint | the transport | strong, but framework result-handling is unobserved | | wrap the framework's tool dispatch | the agent process | weak -- an always-succeeding, zero-latency tool | | modify agent code or configuration | the target itself | weakest -- you are testing your own build | The agreement is ordinary engagement hygiene applied to new instrumentation: which environment, which accounts, which tools may be intercepted, the blast radius, the stop conditions, and who can pull the plug. Skipping it is how a red-team run becomes an incident -- particularly with planting, where the content lives in shared documents or mailboxes that other people and other systems can read. ### Design for the dispute you will get The predictable objection is "your harness did that", and it is often correct. Build the answer into the run rather than assembling it after the challenge: - a **matched control** for every reported behaviour -- same shim, benign content, same shape and length, same latency, same intercepted call; - **delivery verification**, so a null is a real null and a hit is a hit on content the model demonstrably read; - an explicit **changed-properties list** per finding: did the shim make a call succeed that would have failed, collapse latency, alter ordering, bypass result post-processing, or skip a side effect; - a **least-invasive reproduction** for the findings you expect the team to act on. ### What that discipline costs Controls double the inference bill and the wall clock of every scored pass. Delivery verification adds a context-capture path, storage for transcripts, and a redaction step before anything leaves the engagement. The expensive item is the least-invasive reproduction: planting means provisioning throwaway accounts, documents and mailboxes before the run and hunting the residue afterwards, which is routinely the largest single time cost of an agent engagement and a genuine source of cleanup risk. On a two-week engagement, expect the harness, controls and cleanup to consume more calendar time than the attacking. Price that in the scoping conversation, or you will silently trade it away for volume. ### Where the number misleads Two failures of arithmetic destroy interception reports. **Volume without controls:** a large hit count from a heavily modified stack is discounted wholesale, and correctly so -- once one hit is shown to be a shim artefact, the reader stops trusting the rest. **Confusing the tool's unit with the report's unit:** the harness counts attempts, and attempts duplicate enormously, since one underlying failure reproduces across every task variant that touches it. Deduplicate to distinct behaviours with distinct fixes; a hundred hits typically collapse to a handful of items, and quoting the hundred makes the agent look an order of magnitude worse than it is while telling the team nothing about what to change. ### Judgement about what to escalate Not every artefact is worthless. If the behaviour only appeared because your shim forced a call to succeed that would really have been refused, that is worth reporting -- as a **dependency**, documenting that the permission gate is load-bearing and what happens if it is ever relaxed or misconfigured. It is not a breach, and calling it one costs you the credibility you need for the findings that are. Conversely, a behaviour reproduced by planting content in a document that any colleague or any external sender could edit has a severity that does not depend on your instrumentation at all; that is the one worth the team's remediation budget, and the one to lead the report with. ### What you would check before handing it over For each item: is there a control run, is delivery verified, is the changed-properties list present, has it been reproduced at the least invasive placement that still demonstrates it, and is it a distinct behaviour with a distinct fix rather than one of forty transcripts of the same thing?
- The team says your finding is an artefact of the shim. What in your report should already answer that?The matched control run showing the behaviour disappears with benign content, the delivery verification, the list of properties the shim changed, and ideally a reproduction with the least invasive placement.
- When is a hit that only occurs because your shim forced a call to succeed still worth reporting?As a dependency rather than a breach: it documents that the real permission gate is load-bearing, which matters if that gate is ever relaxed or misconfigured.
- Why does the reporting unit differ from the harness's hit count?Hits are per-attempt and heavily duplicated; a report item should be a distinct behaviour with a distinct fix, so many hits usually collapse into a handful of items.
saying these in an interview costs you the question
- Modifies a live agent's code to intercept calls without written agreement or stop conditions.
- Ships an interception-derived hit count with no controls and no statement of what the harness changed.
- Treats every artefact as worthless instead of reporting the real gate it revealed.
- Optimises for the size of the finding list rather than for behaviours the team can act on.