A packaged agent tool world replays each benign user task once per injection task, plus one payload-free pass. For a suite of 40 user tasks and 15 injection tasks, how many task episodes is a full sweep, and what do you cut first when the model-call bill is too large?
answer
- user tasks x injection tasks
- 640 = 40 + 40x15
- episodes are multi-turn, not one call
- stratify: environment and position
- publish the denominator and the seed
basics
~20 sForty clean episodes plus 40 x 15 poisoned pairings is 640 episodes, each a multi-step agent loop costing many model calls. Cut the cross product first: sample a few injection tasks per user task instead of all of them, keep the full sweep for release candidates, and always report which pairings you actually ran.
solid answer
~50 sThe cost is multiplicative, not additive: `user_tasks x injection_tasks + user_tasks` episodes, and each episode is an agent loop of several tool calls and model turns, so the token bill is roughly that product times the average steps per episode. Small-looking suites become expensive immediately, and adding one injection task adds a whole column. The honest way to shrink it is to sample the pairings, not to shorten the episodes. Stratify: keep every distinct environment and every distinct planting position at least once, then sample within each stratum. Drop injection tasks that are duplicates of the same mechanism against the same tool. Cap steps only if you record when the cap fired, because a capped episode is not a clean outcome. The reporting rule matters as much as the sampling: a sampled sweep and a full sweep have different denominators, so their percentages are not comparable across releases unless the sampled subset is held fixed.
go deeper
Should get the multiplication right — 640, not 55 — and see that each entry is a whole agent episode rather than one call.
Explains cost per episode grows with accumulated context, and proposes sampling pairings rather than truncating episodes.
Designs the strata, pins the subset by seed, keeps the clean pass, and refuses to compare sampled to full numbers.
Sets the policy: what runs per change versus per release candidate, what the suite budget buys, and what the organisation is allowed to report from a sampled sweep.
### The arithmetic, out loud A packaged tool world pairs two independent catalogues. Every benign **user task** is replayed once per **injection task**, plus one payload-free pass per user task to establish what the agent does when nothing is planted: ``` episodes = user_tasks x injection_tasks + user_tasks = 40 x 15 + 40 = 640 ``` The instinct that produces 55 is addition, and it is wrong because the two catalogues are orthogonal: an injection task is not a test, it is a *column*. Adding one more injection task to this suite adds 40 episodes, not one. ### What an episode actually costs An episode is not a model call. It is an agent loop — plan, call a mock tool, read the result, call again, answer — and each turn resends the accumulated transcript, so token spend inside one episode grows super-linearly with step count. Take eight turns per episode as a working figure: 640 episodes is roughly 5,000 model calls, and because the average call carries several thousand tokens of accumulated tool output, the input-token bill is in the tens of millions. At the low-single-digit dollars per million input tokens typical of mid-tier hosted models that is tens of dollars; at frontier pricing with reasoning turns it is an order of magnitude more. Wall clock is usually the binding constraint, not money: 5,000 sequential calls at a few seconds each is most of a day, and hosted rate limits cap how much concurrency you can actually use. The mock tools themselves cost nothing — they are local functions — which is precisely why teams underestimate the sweep. ### What to cut, in order 1. **Redundant injection tasks.** Several entries usually differ only in surface wording while exercising the same channel against the same target action. Keep one per mechanism; this is free coverage-neutral savings. 2. **The cross product itself.** Move from every-pair to a stratified sample. Strata that matter: the environment or tool family, the planting position and container within the returned content, and the class of target action. Cover every level of every axis at least once, then fill proportionally. This is the large win, and it is defensible *provided you publish the sampling rule*. 3. **Frequency, not breadth.** Run a fixed, seeded subset on every change; run the full sweep on release candidates and after any model, prompt or guard swap. **Do not** drop the payload-free pass — without it a poisoned outcome has no control to be compared against. **Do not** quietly lower the per-episode step cap: an episode that ran out of steps before reaching the risky action looks identical, in the results table, to an agent that declined it. ### Where the number misleads Two denominator failures dominate. *Sampled versus full.* A percentage over a sampled subset and a percentage over the full sweep are different measurements. Compare them across releases and you will report sampling noise as a trend — and a team that re-samples each run manufactures that trend indefinitely. Pin the subset by seed, commit the seed with the suite. *Undelivered pairings.* Many (user task, injection task) combinations are structurally unwinnable: the user task never calls the tool holding the planted text, so the agent never sees it. Those episodes are counted as attack failures and pull the headline rate down, blending "never delivered" with "delivered and refused". The fix is instrumentation, not arithmetic — log per episode whether the poisoned content entered the context, and report success rate conditioned on delivery alongside the raw one. ### What to check before you believe the run Confirm every episode carries a terminal-reason field: completed, step cap, context limit, tool error, harness crash. Sum them; if a material fraction ended on a limit, your safety number is partly a truncation artefact. Confirm the clean pass's user-task checker still passes, so the suite is measuring an agent that could do the job. Confirm the seed file that pins the sampled subset is the one in the repo. Then publish the denominator with the number: user tasks, injection tasks, pairings run, sampling rule, and how many episodes ended on a limit.
- Why is a pinned, seeded subset better than re-sampling each run?Because a re-sampled subset changes the denominator every run, so movement in the number mixes real regression with sampling noise. A pinned subset makes run-to-run differences attributable to the change under test.
- You must keep only 100 of the 600 pairings. What are your strata?One axis is the environment or tool family, one is the planting position within the returned content, and one is the target action class. Cover each level of each axis at least once, then fill the remainder proportionally.
An injection task is a column in a spreadsheet, not a cell: adding one to a 40-user-task suite adds 40 runs, not one. Teams budget as if they were adding a row.
saying these in an interview costs you the question
- Answers with 55 (additive) rather than the cross product.
- Treats one episode as one model call when estimating cost.
- Shrinks the sweep by lowering the step cap and does not record when the cap fired.
- Compares a sampled percentage to a previous full-sweep percentage.
- Drops the payload-free pass to save money.