A promptfoo red-team run reports 96% of cases passing, and almost every safety check in the config is a model-graded llm-rubric assertion. What has that 96% actually measured, and what do you check before quoting it to anyone?
answer
- 96% = grader did not object
- denominator is generated cases
- read passes, not failures
- grader errors counted as passes?
- known-bad fixtures must still fail
basics
~20 sIt measured how often a grading model, reading your rubric, declined to object. That is not the same as safe. Before quoting it, read a sample of passing transcripts, check what the rubric text literally asks, confirm errored grader calls are not counted as passes, and state the denominator: cases this suite generated.
solid answer
~50 sThe number is a measurement made by a second model, and every property of that measurement is baked into it. **What it is.** The share of generated cases where the grader, given your criterion, returned a pass. A loose criterion ("the response should be safe") produces a high number on almost any target. **What it is not.** It says nothing about your attack surface. The denominator is the cases this suite generated, so a high rate over a thin set is thin evidence. **Before quoting it**, I do four things. Hand-read ten to twenty passing transcripts on the highest-risk cases, because refuse-then-comply outputs are obvious to a human and invisible to a lenient grader. Read the rubric text and rewrite anything vague into a specific question. Confirm how the run treats grader errors and timeouts, since counting an errored check as a pass inflates the rate. And check that known-bad fixtures in the suite still fail — the cheapest proof the checks still discriminate.
go deeper
Recognises that a graded pass reflects a grader's opinion rather than a proof, and that the number needs context.
Names rubric looseness and the generated-case denominator, and knows to sample passing transcripts by hand.
Adds error accounting, comparability across config changes, a deterministic floor, and fixtures that prove the checks still discriminate; reports the number with its scope attached.
Decides what such a number is allowed to authorise organisationally, and what independent evidence must accompany it before it appears in a release decision.
## The only sentence the number supports **96% of the cases this suite generated produced an output that a grading model, reading this rubric text, did not object to.** That is the whole claim. It contains three qualifiers that people drop when they quote it: *the cases this suite generated* (the denominator), *a grading model* (the instrument), and *this rubric text* (the question the instrument was asked). Every stronger sentence - "the system is 96% safe", "we fixed 96% of the risk" - requires evidence the run did not produce. ## How the number inflates without anyone lying **Rubric looseness.** In promptfoo an `llm-rubric` assertion puts your criterion in `value:` and hands it, with the target output, to the grading provider. A criterion like `the response should be safe` invites agreement; a criterion like `does the response supply operational detail that advances the requested task` does not. Nothing in the tool warns you that a criterion is loose - the run looks identical either way, and the looser one returns a better number. **Error accounting.** Each graded assertion is a network call to a metered service. Timeouts, rate limits and the grader's own content filters produce assertions with no verdict. If whatever summarises the run folds those into the pass bucket, then a run that was throttled reads as a safer run than one that completed. Under load the missing verdicts are not random either: they cluster in the longest, most adversarial outputs, which are exactly the cases you needed graded. **Denominator drift.** The rate is over cases the suite generated. In promptfoo's red-team mode that set is a function of which plugins and strategies are enabled and how many cases each was asked to produce. Enable another plugin and the denominator changes, which means yesterday's 96% and today's 88% are not the same measurement, and the lower number can represent the better-tested system. A fourth, structural one: if the config uses assertion `weight:` or an `assert-set` with a `threshold:`, a case's verdict is an aggregate score rather than an AND over its checks, so a case can be counted as passing while its safety rubric failed. ## What it cost you to obtain, and why that matters here Nearly every case in this config carries a graded assertion, so the run made roughly one grader call per case on top of each target call - doubling the calls, roughly doubling the wall clock, and adding a second rate-limited dependency. That expenditure buys semantic reach that no regex has. What it does not buy is reproducibility: rerunning the same config with `promptfoo eval --repeat 3` will flip a fraction of verdicts, and that flip rate is the noise band. If nobody has measured it, nobody knows whether a move from 94% to 96% is a change or a coin toss. ## What I do before it leaves my desk - **Hand-read a sample of passes**, twenty or so, concentrated on the highest-risk categories. Sampling failures feels natural and teaches nothing new - failures are already queued for triage. Passes are where a lenient grader hides an unsafe output, and no field in the report points at them. Read the grader's `reason` strings alongside the transcripts, and treat a reason that merely restates the criterion as a sign the grader did no work. - **Read the rubric text literally** and rewrite anything vague into a question two reviewers would answer identically. - **Establish the error accounting.** Confirm the summary distinguishes passed, failed and not-measured, and report the not-measured count explicitly. - **Run the fixtures.** A handful of stored outputs known to be unsafe must fail, and known-good outputs must pass, on this same run. If the known-bad fixtures passed, the 96% is not a safety result at all - the checks have stopped discriminating. - **Put a deterministic floor underneath.** Canary tokens, forbidden artefacts, schema violations and disallowed link hosts should be checked by `contains`, `regex`, `is-json` or a `javascript:` assertion, so some part of the result is reproducible and independent of grader mood. ## How I would phrase it to a stakeholder Not "we are 96% safe", but: of the N cases this suite generated, 96% were judged non-violating by an automated grader running this criterion; here are the M failures and their transcripts; here is the sample of passes we read by hand and what we found; K cases were not measured because grading errored; and here is what the suite did not attempt at all. The last clause is the one that changes decisions, because it is the only place the reader learns what the denominator excluded.
- Why sample the passes rather than the failures?Failures already get triaged by whoever fixes them. Passes are where a lenient grader hides an unsafe output, and nothing in the report will point you at them.
- The suite grew after someone enabled more generated cases and the rate dropped to 88%. Did the system get worse?Not necessarily. The denominator changed, so the two numbers are not comparable. The broader suite is probing behaviour the old one never attempted, and finding more is the expected outcome.
- How do grader errors distort the headline number?If an errored or timed-out check is scored as a pass, a run that hits rate limits looks safer than one that completed. You need the run to distinguish 'passed', 'failed', and 'not measured'.
It is a school reporting a 96% pass rate on an exam the school wrote itself and marked itself. Widening the syllabus makes the percentage drop even though the teaching improved, so a lower number on a broader paper can be the better school.
saying these in an interview costs you the question
- Reports it as '96% safe' with no scope or denominator.
- Only reviews the failing cases and never reads a passing transcript.
- Cannot say how errored grader calls are counted.
- Compares the rate against a previous run whose suite generated a different set of cases.