Two promptfoo red-team runs over the same saved adversarial suite report 12% and 10% failures for your app. Why is "two points better" a weak summary of what the change did, and what view do you look at instead?
answer
- paired join, not two averages
- improved vs regressed off-diagonal
- net two points can hide sixteen moves
- weight by severity and category
- read the regressed transcripts
basics
~20 sA net delta hides composition. Fixing nine cases while breaking seven also reads as two points better, and the seven new failures may be worse than the nine fixed ones. Join the runs case by case, split the moves into newly-failing and newly-passing, and group them by plugin category and severity.
solid answer
~50 sThe suite is heterogeneous, so one rate over it is an average of unrelated things. Because the cases are identical across the two runs, you can do far better than compare averages: join on case identity and classify every case as pass-pass, fail-fail, fail-to-pass or pass-to-fail. That four-way split answers the question the percentage cannot. A change that hardens one attack family while opening another shows up as a small net move but a large churn, and the newly-failing set is where the risk is — a fixed prompt-injection case does not compensate for a new data-leakage one. Grouping the moves by plugin category tells you whether the effect is broad or concentrated; weighting by severity tells you whether the trade was worth it. Whether two points is more than the natural spread of repeated runs is a separate question from what moved; establish the composition first, because a clean two-point move made of two churned cases is a very different result from one made of sixteen.
go deeper
Should at least see that an average can hide individual cases getting worse, and that you can look at which cases changed verdict.
Describes the pass/fail before-after join and looks at the newly-failing cases, not just the total.
Weights the off-diagonal by severity and plugin category, recognises that some churn is grader relabelling, and makes the churn table the reviewed artefact.
Defines what evidence a change needs before it ships: which regressions are disqualifying regardless of the net number, and what per-case data must be retained to prove it.
### Why a paired analysis is available here, and why that matters Comparing two aggregate rates is what you fall back on when the two samples are different populations. That is not the situation. Both runs replayed the *same* saved promptfoo suite, so every case has a before-verdict and an after-verdict and can be joined on its own identity. A paired view is strictly more informative than the two marginal rates, and it is the view that answers the question a reviewer is actually asking: not "is the average better" but "what got worse". ```text after: pass after: fail before: pass stable REGRESSED before: fail IMPROVED stable ``` Only the off-diagonal is news. The headline delta is nothing more than the difference between the two off-diagonal counts, and it collapses very different worlds into one number. Over 200 cases, 12% to 10% is 24 failures to 20 — a net of four. That could be 4 improved and 0 regressed, or 13 improved and 9 regressed. The first is a small clean win; the second churned 22 cases and introduced nine new failures, and it prints the identical headline. ### Where the number misleads, concretely Two readings of "two points better" are wrong in different ways. **It hides composition.** Cases are not fungible. Trading several low-severity tone-of-policy failures for one new high-severity data-exposure failure is a net loss no matter what the count says. The percentage weights every case equally because it has no other information; you do. **It survives statistical scrutiny far less well than it looks.** The discordant pairs are the evidence, and they are what a McNemar-style check reads. Thirteen improved against nine regressed is the kind of split repeated runs of an unchanged configuration will produce on their own — it is not distinguishable from noise at any sensible confidence. Thirteen improved against zero regressed, from the same net delta, is. If you only ever saw the two percentages, you could not tell those two situations apart, and you would ship the first one believing it was the second. **A share of the off-diagonal is not behaviour at all.** The grader is a model reading free-form output; on borderline transcripts it will label the same behaviour differently on two passes. Some fraction of any churn is relabelling, and the only way to find out what fraction is to read the transcripts of cases that changed verdict. ### What to attach to each off-diagonal case - **Plugin category.** Moves concentrated in one category usually mean the change genuinely addressed that behaviour. Churn scattered evenly across every category more often means sampling or grader variation, and a repeated baseline will show similar churn with nothing changed. - **Severity.** Rank the regressed set by severity and look at the worst one first. That single case, not the average, is usually what decides whether the change ships. - **Strategy.** A fix that only holds under one attack strategy has narrowed the surface, not repaired the behaviour; the same underlying weakness will return under a strategy you have not configured. - **The transcript.** Read it. This is where grader relabelling gets separated from real regressions. ### What it costs to have this at all Almost nothing in compute — the join is over data the run already produced — but it costs *retention discipline*. You must keep per-case verdicts, not just the summary rate, and case identity must be stable across runs (join on the case's variable payload, or a hash of it, if the runner's ids are not stable across regenerations). A few megabytes of JSON per run, or promptfoo's local eval store, is the entire storage bill; promptfoo's own viewer will diff two recorded evals for you. The real spend is human: budget twenty to forty minutes to read a churn set of twenty cases, and treat that reading as mandatory rather than optional. Discard per-case results and keep only the percentage, and no later analysis can recover the pairing — the information is simply gone. ### The artefact to review Make the churn table the thing that goes in the review, not the rate: improved and regressed counts, broken out by plugin category, with the worst regression by severity quoted in full alongside its transcript, and a line naming the suite version and the pinned grader both runs used. The single rate remains useful as a trend line across many changes. It is a poor basis for a decision about one.
- What has to be true of the suite for the paired view to be possible?Stable case identity across runs and stored per-case verdicts. If cases are regenerated or only the summary is kept, you are back to comparing two averages.
- The churn is large but scattered evenly across every plugin category. What does that suggest?More likely variance — sampling, or a grader that is borderline on many cases — than a targeted behavioural change. Repeat the baseline and see whether similar churn appears with nothing changed.
A net two-point improvement built from thirteen fixes and nine breakages is like a portfolio that is up 2% because thirteen holdings rose and nine collapsed. The average is true and tells you nothing about what you now own.
saying these in an interview costs you the question
- Reporting only the net rate when per-case verdicts were available.
- Treating all cases as equally severe when trading improvements against regressions.
- Never reading a regressed transcript, so grader relabelling is mistaken for a behaviour change.
- Discarding per-case results and keeping only the summary, which makes any later pairing impossible.