An AI red-team programme's dashboard leads with mean time to remediate, but the fixes are made by application teams and by a model provider the programme does not control. What breaks about that metric, and what would you put on the dashboard instead?
answer
- clock runs on someone else's backlog
- probabilistic behaviour has no closure event
- define closure as measured rate on rerun
- own: triage, notify, verify
- publish the unowned bucket
basics
~20 sIt scores teams you do not manage, so it measures their backlog, not your work. Worse, many AI findings have no fix you own: the behaviour lives in a hosted model or a third-party guardrail. Report what you control — time to triage, time to notify, and the age and acceptance status of open items.
solid answer
~60 sMean time to remediate is a remediation-org metric borrowed into a testing org. Three things break. First, ownership: the clock runs on application teams and on a provider whose roadmap you cannot see, so a bad quarter can mean one team was reorganised, not that anything about the testing changed. Second, definition: for a probabilistic behaviour there is often no discrete fix event. The team adds a filter and the success rate drops from a high rate to a low but non-zero one. Is that closed? Without a stated exit criterion — a measured rate below an agreed threshold on a rerun — the clock has no stop. Third, the mean hides the shape. A handful of ancient accepted-risk items and a pile of same-day trivia produce a comfortable average that describes neither. Replace it with metrics you own end to end: time from hit to triage decision, time from triage to notification with a reproduction, and an ageing view of open items split by status — fixed, mitigated behind a compensating control, formally accepted, or unowned. Publish the unowned bucket loudly; it is the one that actually needs a decision from the room you are reporting into.
go deeper
Notices the red team does not do the fixing, so the metric is measuring another team's backlog.
Adds that a probabilistic behaviour has no clean closure event and that the clock needs a defined stop condition.
Splits the dashboard into programme-owned timings and estate state, defines closure as a measured rate on a rerun, and surfaces the unowned bucket as the escalation path.
Negotiates the reporting contract itself — which numbers score whom — and accepts that verification reruns cost budget that must be funded explicitly.
## What the borrowed metric assumes **Mean time to remediate** comes from vulnerability management, where it works because three things hold, with the clock starting at disclosure and stopping at deploy: - each finding has **one owner**, - **one discrete fix**, - and an **unambiguous closed state**. Applied to AI red-team findings, none of the three holds, and each break moves the number in a way that has nothing to do with the testing team's performance. ## Break one: the clock runs on someone else's queue A behaviour can be fixed: - in the application's own system prompt or retrieval layer, - in a guardrail's threshold or category configuration, - in the choice of model version pinned by the client, - or nowhere the organisation controls at all — because the behaviour is a property of a hosted model whose vendor roadmap is invisible. The programme owns **discovery** and nothing after it. A quarter where the mean doubles may mean one application team was reorganised, or that the vendor released nothing that quarter. ## Break two: there is no closure event A **probabilistic control** does not go to zero. A team ships a filter and the reproduction rate falls from roughly 40% of attempts to something small but non-zero. Is that closed? Without a stated **exit criterion** the ticket closes on an opinion and reopens the next time anyone tries the prompt. The defensible criterion is a measured reproduction rate below a threshold agreed *before* the rerun, on the same attack set at the same sample count, recorded as **mitigated with its residual rate** rather than as fixed. ## Break three: verification costs the budget again, and rare things are expensive to measure If the original attack set was 300 prompts at 10 samples, every "is it fixed yet?" is another 3,000 calls, and a three-iteration fix cycle spends it three times. Worse, the statistics of rare events bite exactly where mitigations land: - At a true 2% reproduction rate, a 20-attempt rerun returns zero hits about two-thirds of the time, so "we could not reproduce it" is the expected outcome of an **under-powered check** even when the behaviour is entirely present. - Distinguishing 2% from 0% with any confidence takes hundreds of attempts. - Estimating a 2% rate to within about a percentage point takes on the order of 750. That cost lands on the testing team's calendar while the metric it feeds scores somebody else. ## How the mean itself misleads The distribution is **bimodal**: a pile of trivial items closed same-day and a long tail of accepted or unowned ones, so the mean describes no real item and moves mostly with the mix. Three specific misreadings follow: 1. A **silent provider model swap** can make a batch of behaviours stop reproducing, closing items and improving the mean with nobody having fixed anything — and it can as easily reintroduce them next month. 2. **Accepted risks** sitting in the same pool either inflate the mean forever or get quietly excluded, and which one is happening is rarely written down. 3. And the ugliest incentive of all: the mean improves when the team files fewer hard findings, so a metric meant to reward remediation quietly rewards **shallower discovery**. ## What to put up instead — two groups, clearly separated *Owned by the programme, reported as performance:* - median time from hit to **triage decision**; - median time from triage to a notification carrying a working reproduction and its measured rate; - **verification turnaround** once a team declares a fix ready; - and the **precision of the automated stage**, because it explains the triage load behind the first two. *Owned elsewhere, reported as estate state rather than as score:* an ageing view of open items by status — - fixed and verified by rerun, - mitigated behind a compensating control with a residual rate, - formally accepted with a named accepter, - and **unowned**. The unowned bucket is the point of the whole view. An item nobody is fixing and nobody has accepted is an **escalation**, and attaching a count and a maximum age to it is how a testing team forces that decision without pretending to manage another team's backlog. ## What to check 1. Pull the ten most recent closures and ask for the **rerun evidence** on each: attack set, sample count, measured residual rate, date. Closures justified by "no longer reproduces" with no mitigation shipped are sampling luck or a vendor-side change, and both will come back. 2. Then look at the **histogram** rather than the mean. 3. And confirm accepted-risk items are excluded from the timing pool and reported separately. A remediation clock with no defined stop is not a metric; it is a negotiation, and it will be renegotiated in the room where it hurts most.
- A team ships a filter and the reproduction rate drops from frequent to rare but non-zero. Do you close the item?Only against a pre-agreed threshold, measured on a rerun with the same attack set and sample count, and recorded as mitigated with a residual rate rather than fixed. Otherwise closure is an opinion and the item reopens.
- Leadership insists on a single remediation number. What do you give them?The ageing view summarised, split by owner and by status, with the unowned count called out. If they still want one number, make it the count and maximum age of unowned items, since that is the one they can act on.
- Why is a mean particularly bad here?The distribution is bimodal — trivial items closed same-day and a long tail of accepted or unowned ones — so the mean describes no real item. Show the ageing distribution instead.
Scoring a testing team on mean time to remediate is like grading a smoke detector on how fast the fire brigade arrives. It measures a queue the detector does not own, and it improves if the detector simply stops going off.
saying these in an interview costs you the question
- Reporting a remediation mean as the red team's own performance number.
- Closing an item because a mitigation shipped, with no rerun measuring the new reproduction rate.
- Hiding accepted-risk items inside the same mean as ordinary open items.
- No 'unowned' state at all, so items with no accepter quietly age as 'open'.