skip to content

Two engineers drive the same injection payload at the same agent with the same red-team harness and report attack success rates of 5% and 40%. Before concluding the target changed, which run-level settings do you line up between the two sweeps?

level: seniorimportance: should knowfreq 44%

answer

  1. denominator first, then criterion
  2. turn budget raises the rate mechanically
  3. seed narrows variance, does not remove it
  4. aborts: counted, dropped or retried
  5. run manifest, compare by diff

basics

~20 s

Compare trial counts first — 5% may be 1 of 20. Then decoding settings and any fixed seed, the turn budget per episode, which object decided a hit and on what threshold, the starting environment each trial began from, and whether aborted or retried episodes were counted as trials.

solid answer

~50 s

A rate is a property of the whole rig, so two rates are comparable only when the rig matches. Work down a short list. **Sample size.** 1/20 and 8/20 are 5% and 40%; those intervals overlap heavily, so there may be nothing to explain. **Decoding and seeding.** Temperature or top-p differences change how often the agent takes the compliant branch, and a fixed seed on one side only makes the two sweeps different experiments. **Turn budget.** More turns per episode is more chances for the payload to reach the surface it needs, which raises the rate mechanically. **The hit criterion.** A world-state-change criterion and a transcript-wording criterion measure different events under one label. **Starting world and content placement.** Both change reachability. **Bookkeeping.** Episodes killed by tool errors or rate limits, counted as misses on one side and dropped on the other, move the rate with no behaviour change.

go deeper

for a junior

Checks that both engineers ran the same number of trials and the same settings before assuming anything changed.

for a middle

Names the concrete knobs — trials, decoding, turn budget, hit criterion, starting world — and explains why each moves the rate.

for a senior

Orders the investigation cheapest-first, catches abort-accounting and retry bias, and knows a seed narrows rather than removes variance through a hosted endpoint.

for a principal

Institutes a run manifest so rates are only ever compared under matching configuration, making a genuine target-side change detectable at all.

**Attribute before you escalate.** "The vendor changed the model under us" is the most expensive explanation available and the last one to reach for: it is unfalsifiable from your side and it converts a measurement dispute into an external ticket nobody can close. Everything below is checkable in your own logs in about an hour, and in practice one of the first three explains the gap. **Order of suspicion, cheapest first.** 1. *The denominator.* Recount hits over completed trials. 1/20 is 5% and 8/20 is 40%; those intervals overlap heavily and there may be nothing at all to explain. Percentages are where a tiny n goes to hide. 2. *The hit criterion.* Ask each engineer which event they counted. This is the most common genuine cause. A criterion that matches wording in the transcript and a criterion that diffs the world the agent acted on disagree precisely on the interesting cases — the agent narrates a refusal and acts anyway, or narrates compliance and does nothing. Two numbers that scored different events are not two measurements of one thing. 3. *Turn budget and available tools.* A longer horizon, or one extra tool in the agent's toolset, enlarges the reachable state space and raises the rate mechanically with no change in the target's disposition. 4. *Decoding settings and seed.* Temperature and top-p change how often the compliant branch is taken, and a seed fixed on one side only makes the two sweeps different experiments. Note also that a seed does not buy determinism through a hosted endpoint: server-side batching, floating-point non-associativity across batch shapes, tool round-trips whose latency reorders things, and a stateful environment all remain. A seed narrows variance; it does not remove it. 5. *Starting world and placement.* Whether each trial began from an equivalent snapshot, and where the injected content sat relative to what the agent actually reads. Same payload, different placement, different reachability. 6. *Abort accounting.* Rate-limited, timed-out and tool-errored episodes: counted as misses, dropped, or retried. Retrying only the failures is a quiet bias, because failures are not a random subset of trials. This alone can move a rate several-fold with zero behavioural difference. 7. *Time and load.* Sweeps run days apart against a hosted endpoint can differ for reasons neither engineer controls — which is exactly why the six cheaper checks come first. **What the investigation costs.** An afternoon of log reading if both sweeps kept per-trial records, and impossible if either kept only an aggregate — which is the real argument for per-trial logging. The decisive test is re-running one sweep under the other's configuration, and that costs a full sweep: payloads x trials x turns x tokens, plus the wall-clock the endpoint's rate limit imposes, plus a fresh environment reset per trial. Budget it as a line item, because the alternative is a standing disagreement blocking a release decision. **Where the numbers mislead once you start comparing.** The dominant illusion is that a percentage measures the target. It measures the rig: harness version, scoring object, turn budget, tool set, environment snapshot, decoding settings, abort policy. Change any one and the same target honestly yields a different number. Second, treating a fixed seed as a determinism guarantee leads teams to conclude a target changed when nothing did — the seed made them expect identical episodes, so any difference looks like evidence. Third, and the trap that survives all the others: a comparison across a scorer upgrade looks matched. The same label, attack success rate, computed by a stricter judge, is a different quantity, and nothing in the number announces the change. Fourth, overlapping intervals get read as disagreement; at n = 20 a five-point difference is not a signal at all. **The durable fix.** Make the rate carry its configuration. Emit a run manifest with every sweep — trial count, turn budget, tool set, decoding settings and seed, hit-criterion identifier and version, environment snapshot id, abort policy — and refuse to compare two rates whose manifests differ on anything that moves the number. Then a disagreement becomes a diff rather than an argument, and a genuine target-side change finally becomes detectable, because it is the only explanation left once the manifests match and the intervals still fail to overlap.

  • One engineer retried every episode that hit a rate limit; the other counted those as misses. Which rate is right?
    Neither as reported. Aborted episodes are not trials — exclude them from the denominator and disclose how many there were, so the reader sees the evidence that never ran.
  • Both sweeps used a fixed seed and still disagree. What does that tell you?
    That a seed does not buy determinism through a hosted endpoint with batching, tool round-trips and a stateful environment. Keep attributing on the other axes rather than assuming a target change.

saying these in an interview costs you the question

  • Blaming a silent target-side change before checking the trial counts.
  • Comparing a transcript-substring hit criterion against a world-state-diff criterion as if they measured the same event.
  • Assuming a fixed seed makes a hosted agent's episodes deterministic.
  • Retrying only the failed episodes and reporting the resulting rate.
  • Sweeps with no recorded configuration, leaving the disagreement unresolvable.

context