You get three numbers on the slide that decides whether an AI red-team programme is funded next year. Which three do you choose, and how do you defend each against the objection that the team producing it can move it?
answer
- choose by who can move it, not by accuracy
- new-cause issues vs the risk register
- untested surfaces with reason codes = the ask
- unowned items: only others can lower it
- no metric measures residual risk
basics
~20 sPick numbers the team cannot inflate alone. One: issues whose cause was new to the risk register, severity-banded. Two: named untested surfaces with reason codes, especially budget-blocked ones. Three: count and maximum age of open items with no owner and no acceptance. Defend each by naming who verifies it outside the team.
solid answer
~1 minThe selection rule is not accuracy, it is who can move the number. A funding metric that its own producer can raise unilaterally will be discounted the moment a sceptic notices, so choose numbers with an external check built in. **New-cause issues, banded by severity.** Counts volume-proof work: re-running a prolific probe family produces nothing here because the causes are already on the register. External check — the register is owned by risk, not by the red team, and the band is assigned in triage with the owning team present. **Named untested surfaces with reason codes.** This is the honest half of the story and, counter-intuitively, the strongest funding argument: budget-blocked entries convert straight into an ask with a measured cost per attempt behind them. External check — the denominator comes from the threat model's surface enumeration, which the architecture owners maintain. **Unowned open items, count and maximum age.** Measures the escalation the programme exists to force. The red team cannot lower this by working harder; only a decision elsewhere can. That asymmetry is what makes it credible. What I deliberately leave off: raw hit counts, total findings, and a remediation mean whose clock runs on teams the programme does not control.
go deeper
Can say raw counts are weak and that severity and scope belong on the slide.
Picks defensible metrics and explains why counts and remediation means are poor, but tends not to consider who sets each number.
Selects metrics with external verification and explains the incentive each one creates, including its cost to the programme.
Treats metric choice as an organisational contract: negotiates the framing with the owners the numbers touch, states plainly that none of them measures residual risk, and accepts a metric that can report an honest bad quarter.
## The design rule Every metric a programme is funded on gets optimised, so the selection criterion is not accuracy — it is ***who can move the number***. A figure its own producer can raise unilaterally will be discounted the moment a sceptic notices, and the discount is permanent. Pick three whose value is set at least partly outside the team, and say so on the slide. ## Metric one: new causes **1. Confirmed issues whose cause was new to the risk register, banded by severity.** - **Pays for:** reaching a surface or technique that has not been reached before. - **Gaming resistance:** the novelty denominator is the **risk register**, maintained by risk rather than by the red team, and re-running a prolific probe family scores zero here because its causes are already recorded. - **What it costs:** reconciling each confirmed item against the register is roughly 15-30 minutes with a risk owner, so 40 items a quarter is a couple of engineer-days of pure bookkeeping, and it must be funded or it will be skipped. - **How it can still mislead:** the number is set by the register's **granularity** more than by the testing. A register with five coarse causes makes almost everything a repeat and drives the metric to zero; a register with five hundred fine-grained ones makes almost everything new. Publish the register's cause count and its revision alongside the metric, or the trend is uninterpretable. It can also be legitimately near zero in a good quarter on a mature product, which is why it must never appear without the second number. ## Metric two: untested surfaces **2. Named untested surfaces with reason codes, against a dated enumeration.** - **Pays for:** honesty about scope, and it is the ask itself. The **reason codes** split the list into: - a funding request (budget exhausted, with a measured cost per attempt behind it), - an engineering request (no harness), - a third-party negotiation (not authorised) - and a recorded decision (out of scope by design). - **Gaming resistance:** shrinking the list requires actually testing something, and the **enumeration** is owned by architecture. - **What it costs:** a standing review with architecture per product per quarter to keep the enumeration current. - **How it can still mislead:** the list shrinks if the enumeration shrinks, and a surface nobody wrote down is invisible in both directions — so publish the enumeration's revision date and its diff, not just the ratio. ## Metric three: unowned items **3. Unowned open items — count and maximum age.** - **Pays for:** forcing decisions. An item nobody is fixing and nobody has formally accepted is the **escalation** that the reporting exists to produce. - **Gaming resistance:** the red team cannot lower it by working harder; only an owner fixing, mitigating or explicitly accepting moves it, and that **asymmetry** is precisely what makes it credible under challenge. - **What it costs:** every item needs an ownership field and a decision SLA, which is process work someone must maintain. - **How it can still mislead:** it also falls when someone bulk-accepts a batch of risks, so an improvement by paperwork looks identical to an improvement by engineering. Report acceptances as their own line so the two are distinguishable. It is also politically loaded — it makes other leaders look bad in a shared room — so it has to be introduced as an **estate measure**, agreed with those owners in advance, or the programme loses the cooperation it depends on. ## What goes underneath, never on top Attempts run, query budget consumed and cost per attempt belong on the slide as **spend and context**. They explain the shape of the work — why a single-turn sweep is cheap per attempt and an agent engagement is not — without ever becoming the target, because as a target they pay for volume, and volume is the one thing the programme most needs not to be paid for. ## What I deliberately leave off, and why - **Raw scanner hit counts**, because they move with catalogue and sampling configuration rather than with the system. - **Total confirmed findings**, because it is a production count and production counts reward the noisiest instrument. - **A remediation mean**, because its clock runs on teams the programme does not control and it improves when the team files fewer hard findings. ## The caveat to say out loud, before someone else does None of the three measures **residual risk**, because nothing a red team runs can. They measure: - where the instrument was pointed, - where it was never pointed, - and what nobody has decided about. Claiming more than that is the fastest way to lose the argument the first time a real incident lands on a surface the sheet implied was covered — and that argument, once lost, takes the funding with it.
- New-cause issues came out at zero this quarter. How do you present that?As a result, paired with what was tested, what was newly reached, and the budget spent. Zero on a mature surface after real work is meaningful; zero because nothing new was attempted is a different story and the surface list shows which it was.
- Why not put attempts run or query budget consumed on the slide as one of the three?They are context, not score. As targets they pay for volume, and volume is the thing the programme most needs to not be paid for. Show them under the three, labelled as spend.
- What is the strongest objection to the unowned-items metric?That it scores other teams from a testing org's slide. The answer is to frame it as an estate state with no individual attribution, agreed with those owners before it appears, and to report acceptance as an equally valid resolution.
saying these in an interview costs you the question
- Choosing three numbers the team can all raise on its own.
- Claiming the metrics show the product is safe rather than showing where the instrument was pointed.
- Dropping the untested list because it looks bad in front of leadership.
- Introducing an unowned-items count in a shared review without agreeing the framing beforehand.
- A remediation mean presented as the programme's performance.