In an agent injection benchmark, "task completion under attack" is the share of tasks the agent still finishes while an injection is present. Should that share be computed over every task in the suite, or only over the tasks the agent finishes when no injection is present — and why does the choice change what the number means?
answer
- denominator moves, ratio lies
- whole suite vs conditioned on clean pass
- report the pair, read the drop
- per-task pairing shows flips
- task mix is a choice
basics
~20 sCompute it over the same fixed task set as the no-injection run, and read the drop between the two. Conditioning only on tasks the agent already passes cleanly isolates the defence's cost but shrinks the sample. Either way, report the benign number beside it; a bare under-attack percentage mixes agent capability with attack damage.
solid answer
~50 sThe two denominators answer different questions. Over the **whole suite**, the number is "how much of the intended work gets done in a hostile environment". It is the honest product-level figure, but it is dominated by the agent's raw capability: an agent that fails half the tasks anyway will look badly damaged by any attack. Conditioned on the **tasks the agent passes cleanly**, the number is "of the work this agent can do, how much survives". That isolates the attack's and the defence's cost, but it silently changes the denominator per agent, so two agents' conditioned numbers are not comparable to each other, and a weak agent gets an easy denominator. The usable practice is to report both raw rates on the identical fixed task list and treat the difference as the cost, plus a per-task pairing so you can see *which* tasks flipped. Aggregates hide the case where a defence loses ten tasks and gains eight different ones.
go deeper
Should at least say the under-attack number is meaningless on its own and needs the no-injection number from the same tasks.
Explains both denominators, what each answers, and that the drop between matched conditions is the cost figure.
Adds per-task pairing, checks whether the injection was delivered, and refuses cross-agent comparison of conditioned rates.
Treats the suite composition as a controlled asset — pinned task and injection sets — because both rates are averages over a mix someone chose.
### The quantity, stated precisely "Task completion under attack" is a ratio: episodes in which the environment's end state shows the user's own goal was reached, divided by some set of attacked episodes. The numerator is rarely the argument. The **denominator** is, because three different sets are in circulation and each turns the same numerator into a different claim. **1. The whole fixed suite.** Every user task in the environment, whether or not this agent can do it. Benign completion 71%, under-attack 38%; the 33-point drop is the visible cost of the attack plus the defence. Because the denominator is a property of the suite and not of the agent, two agents' numbers sit on the same footing. This is the pair you publish. **2. Conditioned on clean success.** Only the tasks this agent completed in the no-injection run. It answers the question a defence owner actually asks — *of the work this agent can do, how much survives?* — but the denominator is now agent-specific. A weaker agent is scored on its own smaller, easier subset, so its conditioned number flatters it, and cross-agent comparison quietly stops being valid. Use it inside one agent's before/after, never across agents. **3. Conditioned on the injection firing.** Only episodes where the payload demonstrably reached the model's context. This is a diagnostic denominator: comparing it against (1) tells you what fraction of your "resilience" was really non-delivery. ### What each denominator costs you in precision Conditioning is not free. A suite of 200 tasks with a 71% clean pass rate leaves 142 tasks in denominator (2). At an observed conditioned survival near 55%, the 95% interval is roughly +/- 8 points; two defences that differ by 5 points are indistinguishable. Halving that interval means **four times the episodes** — and each episode is 5-15 model calls plus tool round-trips, so a suite that took two hours and tens of dollars now takes eight hours and hundreds. That arithmetic is why teams reach for a conditioned denominator in the first place (it looks cleaner) and why the honest move is usually to keep the full suite and report the pair rather than to chase a tighter conditioned number. ### Where the number misleads - **A bare under-attack rate mixes two things.** 38% conflates "the attack broke it" with "the agent could never do it". Only the drop from a matched benign run separates them, and the drop is the figure that answers the interview question. - **A high under-attack rate is suspicious, not reassuring.** Under-attack completion at or above benign completion almost always means the injection was not delivered, the two conditions ran different task lists, or the checker is lenient. Treat it as a harness bug until disproved. - **Aggregates hide swaps.** Two runs at 38% can be different sets of 38%. Without per-task, per-seed pairing you cannot see that a defence lost ten tasks and gained eight others, and the losses usually cluster in one capability — long tool chains, multi-hop retrieval — which may be exactly the flagship use case. - **The suite's mix is a choice.** Both rates are averages over whatever tasks and injections the environment happens to hold. Adding ten easy single-tool tasks lifts both numbers with no change to model or defence, so a quarter-over-quarter comparison against an unpinned suite measures the suite. ### What you would check Ask which denominator, in writing. Confirm the benign and attacked arms ran the same task list, tools, model version and decoding settings. Confirm aborted episodes were excluded and counted separately rather than silently scored as failures — otherwise the denominator quietly includes runs that never tested anything. Confirm the undefended baseline shows real attacker actions, proving delivery. Ask for per-task outcomes, not just the two percentages: pair them and look at the flips. And pin the task and injection lists by version, because both rates are only comparable over time if the mix behind them is frozen. If only an aggregate survived the run, the number can tell you a defence was expensive; it can never tell you what it was expensive at.
- You see under-attack completion higher than benign completion. What do you check first?Whether the two conditions really ran the same tasks and configuration, and whether the injection was delivered at all — that ordering usually means a harness bug, not a helpful attack.
- Why keep per-task outcomes rather than just the two rates?Because equal aggregates can cover different task sets; pairing per task shows which capabilities the defence actually cost you.
- Can you compare this quarter's under-attack rate to last quarter's?Only if the task list, the injection set and the scoring checks are unchanged. Any composition change moves the rate on its own.
Two shops both report converting 60% of visitors, but one counts everyone who walked past the window and the other only the people who already had the product in hand. Same numerator, different denominator, incomparable claims.
saying these in an interview costs you the question
- Quoting an under-attack completion rate without the benign rate from the same suite.
- Comparing conditioned rates across agents with different baseline capability.
- Keeping only aggregates, so pass-to-fail flips are invisible.
- Not noticing that a very high under-attack completion rate can mean the injection never reached the model.