skip to content

If an AI red-team programme's headline number is confirmed findings per quarter, which part of its automated testing gets more investment over time, and why is that the wrong part?

level: middleimportance: must knowfreq 55%

answer

  1. count = production target
  2. cheapest yield wins the budget
  3. prolific family re-reports known classes
  4. new behaviour classes, not volume
  5. hard surface returning nothing = still work

basics

~20 s

Whatever produces findings cheapest wins the budget. That is usually a high-yield, low-severity probe family against an easy surface. The hard work — a slow multi-turn attack on a tool-using agent — yields few items and looks unproductive. The programme drifts toward whichever instrument is noisiest, not toward where the risk is.

solid answer

~60 s

A finding count is a production target, and people optimise production targets. The cheapest way to raise it is to run more of the probe family that already hits most often, against the surface that is easiest to point a tool at — typically a single-turn chat endpoint with a permissive category list. The expensive work scores badly under exactly the same rule. A multi-turn attack against an agent with tool access may burn a large query budget and return one item, but that item is the one with real blast radius. So the metric quietly reallocates effort away from it. The second-order damage is worse: once the number is a target, nobody wants to disable a noisy probe family even after its findings are known-and-accepted, because disabling it makes the programme look less productive. The fix is to stop making a single production count the headline. Report severity-banded confirmed issues, and pair them with what was tested and what was not, so effort spent on a hard surface that returned nothing still shows up as work done.

go deeper

for a junior

Recognises that counting findings rewards whatever produces the most findings, and that easy surfaces produce more than hard ones.

for a middle

Explains cost-per-finding across single-turn probes versus multi-turn agent work, and that a prolific family re-reports already-accepted classes.

for a senior

Proposes a replacement set — severity bands, new behaviour classes, surface tested, budget consumed — and can say what each pays for.

for a principal

Owns the framing that every funded metric will be optimised, and accepts the cost of a metric that can report an honest bad quarter.

**The economics that do the reallocating.** Cost per confirmed finding is not roughly equal across the things a red team does; it differs by one to two orders of magnitude, and a production target ranks by exactly that ratio. | activity | attempts | model spend | engineer time | typical confirmed items | |---|---|---|---|---| | single-turn probe sweep, permissive content categories, chat endpoint | 2,000-10,000 | low tens of dollars | ~1 day of triage | 5-20 | | multi-turn escalation against an agent with tool access, human in the loop | 100-400 | higher per call (long context, tool round-trips) | 1-3 weeks | 0-2 | Divide items by engineer-week and the first row wins by a wide margin, every quarter, without anyone deciding anything. Nobody games it; the ranking does the work. Budget follows the ratio, the sweep gets broadened because broadening is cheap, and the agent engagement gets deferred to next quarter because it looks unproductive on the sheet. **Why the prolific family never gets switched off.** Most hits from a high-yield probe family are variants of a behaviour the organisation has already seen, rated and formally accepted. Under a count metric those still score. Two consequences follow. The team has a standing reason not to disable the family even once its output stops carrying information, and the same already-accepted class is re-triaged every cycle, consuming the triage hours that the hard work needed. The cost is paid twice: once in engineer time spent confirming known things, once in the harder engagement that never got staffed. **Where the number misleads.** The base hit rate of a probe family is a property of how the family was *written*, not of how risky your system is. Catalogue authors build probes to be reliably detectable, because a probe whose failures a detector cannot recognise is useless to them; that design choice makes some families intrinsically prolific. So a findings count, and every ratio built on it, ranks your activities by the catalogue's authoring decisions. Two further readings go wrong in the same direction. Severity-weighting the count looks like a fix, but a family with a small weight and enormous volume still dominates the weighted sum, and the weights are usually assigned by the team the number scores. And a rising count is routinely read as a rising risk, when the honest reading is that the instrument got wider — while a genuinely productive quarter on a hard surface, which reaches somewhere new and returns two items, reads as collapse. **What to put on the sheet instead.** A small set that cannot be raised by running more of the same: - *Confirmed issues banded by severity*, with the low band shown but never summed into the headline, so volume in the cheap band cannot carry the number. - *New behaviour classes* — issues whose cause was not already on the risk register. A prolific family drives this to zero because its causes are already there; reaching a new surface is the only way it moves. - *Surface tested this period*, so an agent engagement that honestly found nothing still appears as work done rather than as an absence. - *Attempts run, query budget consumed and cost per attempt*, published as spend and context, never as score. This is what makes expensive work legible as expensive rather than as unproductive, and it is the number that turns "we could not reach the tool-calling path" into a funded ask. **What to check.** Take this quarter's confirmed items and reconcile each against the risk register: if more than about four in five map to a cause that was already recorded, the programme is paying to re-discover its own backlog and the headline is measuring catalogue volume. Then check whether the noisy family's hits all trace to a single detector — if so, its yield is one detector's sensitivity, and its contribution to the count should be reported as one line, not as dozens. Finally, price the deferred work explicitly: state the attempts and engineer-weeks the agent engagement would need, so the decision to defer it is a decision someone makes rather than an artefact of a ratio. **The judgement to say out loud.** Every metric a team is funded on gets optimised, so the design question is never "which number is most accurate" but "what behaviour does this number pay for". A count pays for volume. A new-behaviour-class metric pays for reaching somewhere new — at the honest cost that it can report a bad quarter when the truthful answer is that nothing new was found.

  • Is severity-weighting the count enough to fix the incentive?
    It helps but does not fix it, because a probe family with a small weight and a huge volume still dominates, and whoever assigns weights is usually the team being scored. Band and show the bands rather than collapsing to one weighted scalar.
  • How do you report a quarter where a hard agent engagement found nothing?
    As surface tested with the attack classes attempted and the budget spent, plus an explicit statement that no new behaviour class was reached. A credible null result is a product of the programme, not an absence of one.

Paying a red team by the finding is like paying a fishing crew by the fish: they will work the pond where the small ones bite constantly, not the lake that might hold the one big enough to matter.

saying these in an interview costs you the question

  • Defending the count on the grounds that the team would never game it — the metric reallocates effort without anyone deciding to game anything.
  • Adding severity weights to the same count and calling the incentive fixed, while the weights are set by the same team the number scores.
  • Keeping a noisy probe family enabled purely to protect the quarterly number.
  • No place on the report for an engagement that found nothing.

context