skip to content

A promptfoo red-team run over an internal support assistant comes back with zero failures, and the configuration enabled three plugins. What can and cannot you conclude from that report?

level: middleimportance: must knowfreq 61%

answer

  1. rate over enabled classes only
  2. absence of evidence
  3. name what was not tested
  4. stochastic generation, one run is weak
  5. scan describes one build

basics

~20 s

You can conclude that on this run, for those three harm classes, at that suite size, nothing failed. You cannot conclude the assistant is safe. Harm classes with no plugin enabled generated no cases, so the report is silent about them. Silence is a missing denominator, not a pass.

solid answer

~60 s

The only defensible reading is scoped: *no failures were observed for three harm classes, at the suite size configured, against this build, on this date.* Everything beyond that is invention. The trap is that a promptfoo red-team report renders a pass rate, and a rate implies a denominator that feels like "all the ways this could go wrong". It is not. The denominator is the cases the enabled plugins generated. Three plugins out of a substantial catalogue means most harm classes were never attempted, and an untested class is indistinguishable in the report from a class the application handles perfectly. Two secondary caveats belong in the same breath. Generation is stochastic, so a clean run is weaker evidence than a run that found and then stopped finding failures. And a small suite per class can miss a rare behaviour purely by not asking enough times. What I would actually write in the summary: the enabled class list, the suite size, the build, and the phrase "untested" beside every class not enabled.

go deeper

for a junior

Should at least say the result covers only the harm classes that were enabled, and that untested classes are not passes.

for a middle

Gives the scoped sentence — classes, suite size, build, date — and explains that the denominator is a configuration choice rather than the space of possible harms.

for a senior

Adds stochasticity and sensitivity: one clean run is weak evidence, small per-class counts cannot catch rare behaviours, and a scan describes the exact build it ran against.

for a principal

Focuses on how the number gets consumed — what a launch review is allowed to infer, what must be published with the result, and how to stop a narrow scan being cited as broad clearance.

### What the number is made of A promptfoo red-team pass rate is failing cases over graded cases. Graded cases come only from the plugins listed in `redteam.plugins`, at the `numTests` you configured, expanded by whatever is in `redteam.strategies`. So the universe of the metric is a configuration choice you made in advance — usually in a hurry, often by copying a starter config — and a narrow choice produces a flattering number by construction. Three enabled classes with clean results and thirty unenabled classes are indistinguishable on the page: both contribute zero failures. One of those zeros is evidence and the rest are absence of evidence. ### The defensible sentence *No failures were observed across three harm classes, at N cases per class, against build X, on this date, with this promptfoo version.* Everything beyond that is invention, and every clause in it is load-bearing. ### What the run cost, and why that is itself a warning Do the arithmetic on the run described. Three plugins at the default five cases each is fifteen base cases. Fifteen target calls plus fifteen grader calls is a couple of minutes of wall clock and small change on the bill. A result that cheap cannot possibly be a safety clearance for an assistant that reads customer records — the cost of the evidence is a decent proxy for its weight, and anyone quoting a clean run into a launch review should be asked what it cost to produce. ### Where the number misleads **Sensitivity.** Five cases per class is a very blunt instrument. A behaviour that shows up in roughly one reply in ten survives five attempts undetected about 59% of the time (0.9 to the fifth). Absence at that sample size is close to uninformative for anything rare, and rare-but-severe is exactly the profile of the failures worth catching. **Stochasticity.** Both halves of the pipeline are sampled. Regenerating writes different cases; re-running the same cases samples different responses. There is no seed that makes a red-team result reproducible in the way a unit test is, so a single clean run is the weakest form of this evidence. **The judge.** Every verdict is an LLM grader's opinion. Graders are more prone to false negatives than false positives on partial compliance — a reply that refuses in its first sentence and then supplies part of what was asked frequently scores as a pass. A clean report is partly a statement about the judge. **Severity blending.** The headline mixes classes of wildly different consequence. A pass rate dominated by low-severity classes reads the same as one dominated by the class that could disclose another customer's record. **Build identity.** The system prompt, the retrieval corpus and the tool wiring all move independently of the code. The scan describes the exact build it ran against, not the service. ### What I would check before repeating the number anywhere - Read the per-plugin case counts, not the headline. Any plugin that generated fewer cases than `numTests` asked for is a green card over a tiny denominator. - List the plausible harms for *this* application that no enabled plugin covers, and put that list in the same document as the result. For a support assistant with access to customer records, a run that never attempted cross-customer disclosure is not a safety statement about it. - Re-generate and re-run at least once. Two independent clean runs are meaningfully stronger than one; a difference between them tells you how noisy your suite is. - Prove the pipeline can detect something. A clean run that follows a run which found and fixed real failures is far better evidence than a clean first run, because a first clean run is equally consistent with a well-behaved application and with a suite pointed at the wrong harms. If you have never seen this configuration produce a failure, you have not yet established that it can. - Check what changed since the last run: a plugin dropped from the list, a lowered `numTests`, or a promptfoo upgrade that resolved a collection alias differently all improve the headline without improving the application.

  • How would you word the result so a launch reviewer cannot over-read it?
    State the harm classes tested, the per-class case count, the build identifier and the date, then list the plausible harm classes for this application that were not attempted.
  • Why is a clean run that follows a run which found real failures stronger evidence than a clean first run?
    Because you have demonstrated the configuration can detect something. A first clean run is equally consistent with a well-behaved application and with a suite pointed at the wrong harms.
  • Someone proposes tracking the pass rate over time as a safety trend. What must be held constant?
    The enabled plugin set and the per-class case count. If either moves, consecutive numbers have different denominators and the trend line is meaningless.

saying these in an interview costs you the question

  • Reporting the result as the assistant being safe or as a launch clearance.
  • Not asking which harm classes were enabled before interpreting the number.
  • Treating one clean run as reproducible when generation and sampling are stochastic.
  • Failing to name the plausible harms for this specific application that the run never attempted.
  • Confusing an untested class with a class the application handles well.

context