In a packaged agent tool world (mock tools plus tasks, e.g. AgentDojo), you compare a clean replay of a task with a poisoned replay of the same task. What must be identical between the two replays for the behavioural difference to be attributable to the planted content?
answer
- one variable, one candidate cause
- freeze clock, ids, result order
- same container, same position
- length-matched control
- diff the two episodes' rendered inputs
basics
~20 sEverything except the planted bytes: same task, same starting world contents, same tool list and descriptions, same model, sampling settings and step limit. Hold the payload's container and position fixed too, so only its text differs. Drift in timestamps, result ordering or sampling shows up as a fake attack effect.
solid answer
~50 sThe paired replay is the fixture's whole argument: two episodes that differ in one variable, so a difference in behaviour has one candidate cause. That argument only holds if the pair is genuinely matched. Hold constant: the user task text, the world's starting contents, the tool set and every tool name and description, the model identity and sampling parameters, the system prompt, the step and tool-call limits, and the order in which results come back. Hold the *container* constant too — plant into the same message, at the same position, with a comparable length — because moving the payload to the front of a longer result changes two things at once. The common confounds are non-determinism (sampling temperature, timestamps and generated ids in mock data, unordered result sets) and length effects: the poisoned result is longer, so the episode hits a step or context limit the clean one did not. A single mismatched pair is an anecdote; matched pairs repeated a few times are evidence.
go deeper
Should say only the planted content may differ — same task, same tools, same model.
Adds sampling and world non-determinism, container and position matching, and that a single pair is not evidence.
Requires the clean replay to reproduce itself first, adds length-matched controls, and specifies the per-episode logging that makes attribution reviewable.
Sets the standard: no causal claim leaves the team without a reproducible control and the inputs diffable by a reviewer.
### The claim the pair is making A paired replay makes a causal claim: *this planted content changed what the agent did*. In experimental terms the clean replay is the control arm and the poisoned replay the treatment arm, and the claim is only as good as the matching. The list of things that must be identical is long and boring, which is why it is where teams actually fail: | Held constant | Why it breaks the pair if it moves | |---|---| | User task text, system prompt | Different instructions, different plan | | World seed contents | Different data reaches the model beyond the payload | | Tool set, and every tool name, description, argument shape | Those strings are in context on every turn; the model reacts to them | | Model identity and sampling settings (temperature, top-p, seed if offered) | Different distribution over trajectories | | Step limit, tool-call limit, context budget | Changes where episodes are allowed to end | | Result ordering, clock, generated ids | Silent input drift between arms | | Payload container and position, and result length | Moves two variables at once | ### The three ways pairs break in practice **Non-determinism inside the world.** Mocks that stamp the current time, mint random ids, or return an unordered set hand the two arms different inputs before the payload is even considered. Freeze the clock, pin ids, force a deterministic ordering — then verify by running the clean replay twice and diffing the rendered tool results byte for byte. Until the control reproduces itself, no diff against it means anything. **Non-determinism inside the model.** Sampling means identical input yields different trajectories, and an agent loop *amplifies* that: one different early tool call sends the whole episode down another branch. One clean run against one poisoned run is an anecdote. Run each arm several times and compare distributions of outcome, not single trajectories; the unit of evidence is a rate over repeats, with an interval, not a pair of transcripts. **Length and position confounds.** Planting text makes the tool result longer and shifts everything after it. If the poisoned episode ends because it exhausted its step budget or filled its context, you have measured a length effect — and the results table will record it as "no harmful action", i.e. as *safety*. Whenever the planted text is a large fraction of the result, or whenever poisoned episodes end early, add a **length-matched control**: the clean result padded with neutral filler of comparable size, planted at the same offset. ### What it costs Matching multiplies episodes. Repeats to see past sampling noise are a factor of five to ten; a length-matched control is a third arm. A pair that was 2 episodes becomes 30. That is the real price of a causal claim in this setting, and it is the reason attribution-grade runs are reserved for findings you intend to publish or act on, while the routine sweep runs one episode per pairing and claims only a trend. ### Where the number misleads The reported "attack success rate under injection" is a difference between two arms. If the arms differ in more than the payload, the difference absorbs everything else — a model version that rolled forward overnight, a mock that now returns three results instead of two, a longer poisoned body that trips the step cap. The direction of the bias is usually *towards looking safe*, because the most common confound (length) terminates episodes early and early termination scores as no harm. ### What to check Log per episode: the exact rendered tool results, a hash of the tool schema set, the model identity and sampling settings, the seed, the step count, which limit fired if any, and the terminal world state. The bar is that a reviewer can diff the two arms' inputs and see exactly one difference. If your logs cannot show that, the pair proves nothing, however large the reported gap.
- The clean replay does not reproduce itself twice. What do you do before running any poisoned pair?Fix the fixture first: freeze the clock, pin ids, force deterministic ordering, and lower sampling variance. Until the control reproduces, no diff against it is interpretable.
- When is a length-matched control worth the extra episodes?Whenever the planted text is a large fraction of the result, or whenever poisoned episodes end early. Padding the clean result with neutral filler separates 'the instruction worked' from 'the episode ran out of room'.
It is an A/B test where the B page also happens to load two seconds slower. Whatever the metric does, you cannot say whether the new copy or the extra latency did it — and here the extra latency is the longer poisoned tool result running the episode out of steps.
saying these in an interview costs you the question
- Treats one clean run versus one poisoned run as proof of an effect.
- Lets mock data carry live timestamps or random ids across the pair.
- Ignores that the poisoned result is longer and may trip a step or context limit.
- Cannot produce the exact rendered tool results for either episode.
- Changes the tool set or model between the clean and poisoned replay.