skip to content

An agent red-team run that substitutes tool results through a shim reports a large number of policy violations. Before you write any of them up, how do you establish that the shim did not manufacture them?

level: seniorimportance: must knowfreq 52%

answer

  1. control = shim on, payload benign
  2. never-failing tool reaches gated actions
  3. instant return changes retries and budget
  4. injected first call = different position
  5. reproduce with least instrumentation

basics

~20 s

Re-run with the shim in place but a benign result substituted. If the behaviour still appears, the shim caused it, not the payload. Usual manufacturers: the shim answering a call that would have errored, returning instantly, changing call order, or replacing a result the framework would have truncated.

solid answer

~60 s

The control you need is not 'no shim' -- it is **shim present, payload absent**. Run the identical configuration with a benign, realistically shaped result substituted at the same point, and compare. Behaviour that survives that control is an artefact of interception, not of the injected content. The common manufacturers are structural, not textual: - **Success where reality errors.** The shim answers every call. The agent never hits the failure path that would have stopped it, so it proceeds to an action it would never have reached. - **Timing.** An instant return collapses the run, changes retry behaviour, and can change what fits in a turn budget. - **Ordering and position.** Delivering on the first call puts the content in a different context position than the real workflow would. - **Shape.** A hand-written result that skips the framework's truncation, escaping or schema validation gives the payload a cleaner path than any real tool output would have. Only after the control run separates these do you triage what is left.

code

python · 11 lines
python
# hold fixed: task, seed, intercepted call index, latency, result shape/length
for label, body in [("payload", ADVERSARIAL_TEXT), ("control", BENIGN_TEXT_SAME_SHAPE)]:
    shim.configure(
        intercept_call=2,          # not "the first call"
        succeed=REAL_TOOL_WOULD_SUCCEED,
        delay_ms=OBSERVED_REAL_LATENCY,
        body=body,
    )
    record(label, run_agent(task, seed=SEED))

# a violation present in BOTH runs is the shim's, not the payload's

go deeper

for a junior

Should at least suspect that a harness which fakes tool results can cause behaviour on its own, and know to check a run without the payload.

for a middle

Should describe the shim-present, payload-benign control and name two structural artefacts, such as never-failing calls and collapsed latency.

for a senior

Should design the control into the run, enumerate the properties the shim controls, and reproduce survivors with minimal instrumentation before writing anything up.

for a principal

Sets the standard that no interception-derived count ships without its matched control, and budgets for the roughly doubled run cost that implies.

A hit produced by an interception shim reads exactly like a real breach. The transcript shows the agent calling a tool, receiving content, and then doing something it should not. Nothing in that artefact distinguishes *the payload persuaded it* from *the harness put the agent in a state it could never have reached*. Because the two are indistinguishable after the fact, the separation has to be designed into the run. ### The control that actually isolates the payload The instinctive control -- run it again with the shim removed -- is the wrong one: it changes the instrumentation and the content at the same time, so a difference cannot be attributed to either. The control you need is **shim present, payload absent**. Hold fixed the task, the seed, which call is intercepted, whether that call succeeds, the injected latency, and the result's shape and approximate length; swap only the adversarial text for benign text. Any violation that still occurs is caused by the instrumentation, not by the injection. ### The structural artefacts, and why each one manufactures hits 1. **Never-failing tools.** The shim answers every call with success. If the real service would have returned a permission error, an empty result or a rate limit, the agent never meets the failure path that stops it in production, and walks on to actions the real environment gates. Everything downstream of that point is a statement about your shim's permissiveness. 2. **Collapsed latency and suppressed retries.** An instant return changes how many steps fit inside a turn budget or a wall-clock timeout, and removes retry loops that would otherwise have consumed the run. Agents behave differently near their step limit. 3. **Reordering and position.** Injecting on the first call rather than where the workflow would naturally return that content puts the payload earlier in the context, with less competing material after it. That is a different experiment with a different answer. 4. **Bypassed result handling.** A pre-baked string that skips the framework's truncation, escaping, schema validation or summarisation reaches the model in a condition no real tool output ever would. The payload gets a delivery quality no attacker could obtain. 5. **Missing side effects.** A short-circuited tool never writes the state a real call would have written, so later steps see a different world for reasons unrelated to your content. ### What the control costs, and why teams skip it The matched control roughly doubles the run: a 200-case suite at five samples and ~8 model turns per sample is on the order of 8,000 model calls per arm, so budget two arms, tens of dollars of inference, and an hour or more of wall clock per full pass -- plus the engineering to make the shim parameterised enough that latency, success and intercept index can be pinned identically across arms. That parameter plumbing is usually a day of work, and it is precisely the day teams skip, which is why so many interception reports cannot be defended when challenged. ### Where the number misleads "142 violations" from an interception run is not a measurement until you can say how many of the 142 survive the benign-content arm. Two specific misreadings follow from omitting it. First, a shim that always succeeds inflates the numerator with cases that are really about a missing permission gate, and those get filed as prompt-injection findings the scaffold team cannot fix. Second, the denominator is usually attempts, not distinct behaviours: the same underlying failure repeated across 40 task variants is one report item, and quoting 40 makes the agent look forty times worse than it is. A count without its control and without deduplication is an opinion with a number attached. ### What you would check Run the matched control and put its outcome next to the hit count in the write-up. For each surviving hit, ask whether it depended on a call succeeding that would really have failed -- if it did, report it as a dependency on that gate rather than as a breach. Reproduce the survivors with the least instrumentation that still shows the behaviour, ideally by planting the content in data a real tool returns, and treat that reproduction as the finding. Finally, keep an explicit list of every property the shim controls -- success, latency, ordering, shape, position, side effects -- and for each one either hold it at a realistic value or state in the report that you changed it.

  • Why is a run with the shim removed entirely a poor control?
    It changes both the instrumentation and the payload at once, so a difference cannot be attributed to either. The control must isolate the payload with the shim still in place.
  • Your control run shows the agent takes the sensitive action even with benign substituted content. What is the finding now?
    It is a finding about the harness, and possibly about the agent's willingness to act on any tool success -- but it is not an injection finding. Fix the shim's realism and re-run before reporting.
  • Name one shim property that must be held fixed between the payload run and the control run.
    Result length and shape, latency, which call is intercepted, and whether the call succeeds -- any of these; changing one alongside the text reintroduces the confound.

The benign-content run is the placebo arm. Without it you have a drug trial in which every patient also got surgery, and no way to say which one healed them.

saying these in an interview costs you the question

  • Uses a no-shim run as the control, which changes two things at once.
  • Reports a hit count with no control condition at all.
  • Assumes any behaviour that follows the payload was caused by the payload.
  • Lets the shim return success for calls the real environment would reject, and never mentions it.
  • Treats identical transcript appearance as evidence the breach is real.

context