skip to content

Two candidate defences are measured on the same agent injection suite. With defence A the attacker-action rate falls from 24% to 3% and task completion under attack falls from 71% to 38%. With defence B they are 11% and 66%. A platform team asks you to fold these into one "security score" so the two can be ranked. How do you respond?

level: principalimportance: should knowfreq 34%

answer

  1. pair, not a scalar
  2. exchange rate is a product decision
  3. residual in events, not percent
  4. check refusal and abort rate first
  5. split posture by irreversibility

basics

~20 s

Give them the pair, not a scalar. The exchange rate between one blocked injection and one failed user task is a product decision, not a benchmark output. B keeps most of the utility; A buys a lower attacker-action rate by breaking a third of the work. Ask what a residual success actually costs.

solid answer

~50 s

A composite hides the only judgment that matters. Any single score encodes a weight — how many lost user tasks one prevented attacker action is worth — and that weight belongs to whoever owns the product's risk, not to the harness. So the answer is a decision frame, not an arithmetic one. What does a single successful injection cost in this deployment: an irreversible external action, or a bad paragraph a user can ignore? At what volume do attempts arrive? A 3% residual against a high-volume, high-blast-radius surface may still be unacceptable, in which case A is not the winner either; against a low-volume internal tool, B's 11% may be fine while A's 33-point utility loss is not. If a single number is genuinely mandated for a dashboard, define it with visible weights, publish the pair beside it, and re-derive it whenever the suite's task or injection mix changes — because both inputs are averages over a mix someone chose.

go deeper

for a junior

Should notice that A's security gain came with a large utility loss and that both numbers must be shown.

for a middle

Explains that combining them needs a weight, and that the weight is not something the benchmark supplies.

for a senior

Converts residual rates into expected events using volume and blast radius, and breaks the utility loss down by task family before recommending.

for a principal

Refuses an unowned implicit weight, proposes a mixed posture split by irreversibility of the action, and sets the governance around any tracked index.

### Why the scalar is the wrong artefact The two inputs are not commensurable. The attacker-action rate counts adversary-caused actions in a simulated environment; under-attack completion counts user work delivered. Folding them into one score requires an **exchange rate** — how many lost user tasks one prevented attacker action is worth — and that rate is a function of deployment context the harness never observes: blast radius per successful injection, whether the action is reversible, how often attempts arrive, and who absorbs a failed task. A composite does not compute that judgment; it *assumes* one, hides it inside a number, and makes it nearly impossible to revisit later because nobody remembers the weights. That is the objection to give the platform team: not that the arithmetic is invalid, but that it launders an unowned business decision into something that looks like a measurement. ### What the numbers here actually say Defence A removes about seven eighths of the attacker actions (24% -> 3%) and roughly a third of the delivered work (71% -> 38%). Defence B removes just over half the attacker actions (24% -> 11%) and about a twelfth of the work (71% -> 66%). They are not points on one axis. A is a candidate only where a single success is catastrophic and irreversible — an outbound payment, a destructive API call. B is the candidate where the agent's job carries real value and its failures are recoverable. The honest recommendation is usually a third option the composite would have foreclosed by construction: **A's posture on the narrow set of irreversible tool calls, B's on everything else.** ### Turning rates into something rankable - **Residual in expected events, not percent.** 3% versus 11% means nothing until multiplied by attempt volume and by the cost of one success. At 50 injection attempts a week, A leaves ~1.5 and B ~5.5 expected successful actions weekly; if one success is an irreversible transfer, the eight-point gap is the whole decision, and if it is a wrong paragraph a user can ignore, it is not. - **Utility loss by task family.** A 33-point aggregate drop is rarely uniform. Defences cost most on long tool chains and multi-hop tasks, which may be your flagship workflow or a rarity. Break it down before you price it. - **Hygiene figures first.** Refusal rate and abort rate per arm. A low attacker-action rate bought by refusing, or by running out of context, is not a defence; a composite would reward it enthusiastically. ### What the comparison cost, and what that buys in precision Four sweeps, not two: each defence needs a benign and an attacked arm on the identical suite. At 80 user tasks and 30 injection tasks that is ~2,400 attacked plus 80 clean episodes per arm, each 5-15 model calls, so a full A-vs-B decision is on the order of 10^5 calls, a working day of wall-clock at sane concurrency, and a bill in the hundreds of dollars — plus a day of engineer time keeping the arms configuration-identical, which is where these comparisons usually break. It is worth pricing because it bounds what the numbers can support: a 3%-versus-11% gap estimated on a few thousand episodes is solid, but the 66%-versus-71% utility difference sits inside a few points of sampling noise, so "B costs you nothing" is a stronger claim than the run paid for. ### Where these numbers mislead Both rates are averages over a task and injection mix somebody chose; changing the mix moves the ranking with neither defence changing. Both are simulated — a mock banking tool cannot tell you what the real payment API's idempotency behaviour does to a partially executed attack. And a residual rate measured against the suite's fixed injection set says nothing about an adaptive attacker who gets to see your defence; it is a floor on risk, never a ceiling. ### What you would check, and what you hand back Verify the four arms shared task list, tools, model version and decoding settings; pull the per-arm termination and refusal histograms; confirm the undefended baseline's 24% proves delivery. Then hand back the pair per defence with the benign baseline attached, the residual converted to expected events, the utility loss split by task family, and the suite version stated. If leadership still mandates one tracked number, make it an explicitly weighted index whose weights are reviewed like any other risk parameter, published beside the raw pair, and invalidated when the suite changes. What you refuse is not arithmetic; it is an implicit weight nobody owns.

  • Leadership insists on one number anyway. What is the least bad version?
    An explicitly weighted index with published weights, shown next to the raw pair and the abort and refusal rates, and re-derived whenever the task or injection mix changes.
  • What single extra measurement would most change the ranking?
    The blast radius and reversibility of one successful injection in the real deployment, combined with attempt volume — that converts residual percentages into expected harm.

saying these in an interview costs you the question

  • Averaging the two rates and declaring a winner.
  • Ranking on residual attacker-action rate without asking what one success costs.
  • Ignoring that the utility loss may fall entirely on the flagship task family.
  • Accepting a defence's low attacker-action rate without checking its refusal and abort rates.
  • Treating the suite's task mix as a fixed property of the world rather than a choice.

context