skip to content

A promptfoo red-team plugin generates 10 cases against your assistant, and the behaviour it probes appears in roughly 1 relevant response in 20. Roughly what chance does that run have of reporting zero hits, and how many cases would you need before a clean result means something?

level: middleimportance: should knowfreq 45%

answer

  1. 1 - (1 - p)^n
  2. p=0.05, n=10 misses ~60%
  3. ~59 cases for 95% at 1-in-20
  4. generated cases cluster, effective n lower
  5. size to the rarity you must catch

basics

~20 s

Treating each case as an independent try, the chance of missing every time is 0.95 to the tenth, about 60 percent. So a clean run of ten mostly reflects the sample size. To be roughly 95 percent sure of seeing a one-in-twenty behaviour you need about sixty cases.

solid answer

~50 s

The arithmetic is the geometric tail: the probability of at least one hit is `1 - (1 - p)^n`. With p = 0.05 and n = 10 that is about 40%, so most runs report nothing even though the failure is genuinely there. Inverting, `n >= ln(0.05) / ln(1 - p)` gives roughly 59 cases for 95% confidence at that rarity, and around 300 if the true rate is one in a hundred. That figure is an **upper bound on your real power**. Generated cases are not independent draws: twenty variants of one harm often include near-duplicate phrasings, covering one narrow path repeatedly rather than twenty paths once. The target is stochastic too, so a case that failed once can pass on replay. Both push effective n below the configured one. So a small per-plugin count screens for common, easy failures and cannot be reported as evidence about rare ones.

go deeper

for a junior

Should recognise that ten tries is few for a one-in-twenty event and that a clean run is not proof, even without the exact formula.

for a middle

Should produce 1 - (1 - p)^n, compute roughly 60% miss at n = 10, and invert it to about sixty cases for 95% confidence.

for a senior

Should add why the figure is optimistic - correlated generated cases and a stochastic target - and turn the count into an explicit statement of what the run can rule out.

for a principal

Should treat the target detectable rarity as the thing set per harm class and hold reporting to the claim that count supports.

### The model, stated plainly Treat each generated case as one independent trial that either surfaces the behaviour or does not, with per-case probability p - the rate at which your application commits the behaviour on prompts of that shape. The run misses entirely with probability `(1 - p)^n`, so it fires at least once with probability `1 - (1 - p)^n`. This is the geometric tail, and it is the single most useful piece of arithmetic in suite sizing because it converts an opinion about volume into a statement about what a clean report is *allowed to claim*. At p = 0.05 (one in twenty): | cases (n) | chance of zero hits | chance of at least one hit | |---|---|---| | 5 | 77% | 23% | | 10 | 60% | 40% | | 30 | 21% | 79% | | 60 | 5% | 95% | | 90 | 1% | 99% | So ten cases against a one-in-twenty behaviour report nothing about 60% of the time. Inverting for the count you need: `n = ln(alpha) / ln(1 - p)`, where alpha is the miss probability you will accept. With alpha = 0.05 and p = 0.05 that is `ln(0.05)/ln(0.95)` = 58.4, so about sixty cases. Required n scales close to `1/p`, so a one-in-a-hundred behaviour needs roughly 300 cases for the same confidence, and one in a thousand needs roughly 3,000. ### What that costs The count is not free and the scaling is brutal at the rare end. In promptfoo each case is a target call plus, for the model-graded plugins, a grader call, and `redteam.strategies` multiplies the case list on top. Going from 10 to 60 on one plugin adds ~100 calls per run for that plugin alone; doing it across twenty plugins is ~2,000 extra calls, every run, and the same again on the next scheduled run. That is why "just raise numTests" is a budget decision and not a config tweak - and why chasing one-in-a-thousand rarity with a generated suite is usually the wrong instrument entirely. ### Where the number misleads **Independence is the weak assumption, and it fails in your favour on paper only.** promptfoo asks its generator for n cases for one plugin from one purpose description; the results cluster into a handful of phrasings of the same idea. Sixty configured cases may be fifteen distinct probes repeated four ways. The computed 95% is therefore an **upper bound on your real detection power**, not an estimate of it, and the gap widens as n grows because the generator runs out of genuinely different angles long before it runs out of quota. **The target is stochastic.** A case is not a fixed pass or fail; the same case can hit tonight and pass tomorrow at the same temperature. The formula still works if you read p as the average per-case hit probability, but it means a single hit is not proof of a reliable failure and a single clean replay is not proof of a fix. **p is unknown - that is what you are trying to learn.** So this is a planning tool run backwards from the rarity you have decided you must catch, never a measurement of your actual risk. Quoting "we detect 95% of failures" from it is a category error. **The denominator swap.** p here is per relevant case. It is not per user, per session or per day. A one-in-twenty per-prompt rate on an assistant handling ten thousand prompts a day is a near-certain daily occurrence, and reporting it as "5%" to a stakeholder invites the wrong sense of scale in the other direction. **Many plugins, many chances.** Run twenty plugins and you are running twenty tests at once. Even a grader with a small false-positive rate will produce occasional spurious hits somewhere in the suite, so a single isolated hit across a large suite deserves replay before it becomes a finding. ### What I would check Before believing a clean run: what n actually ran for that plugin; how many of those n cases are meaningfully distinct; and whether the previously-hitting cases were among them. Before believing a hit: replay it several times to estimate its own rate rather than reporting a one-off as a live failure. And in the report, next to every clean plugin, print the sentence the count earns - at n = 10 this run rules out behaviours much more common than one in ten and nothing rarer - so nobody upgrades a thin sample into an assurance on your behalf.

  • The true rate is one in a hundred rather than one in twenty. How does the required count change?
    Roughly five times larger: about 300 cases for the same 95% chance of at least one hit, since the required n scales close to 1/p.
  • Why is the computed count optimistic in practice?
    Generated cases are correlated - several near-identical phrasings of the same attack shape - so the effective number of distinct tries is below the configured one, and a stochastic target adds further variance.
  • The suite hit twice last week and zero times this week after a fix. Is the fix confirmed?
    Not from that alone. At a small count, going from two hits to zero is well within noise; you need enough cases to distinguish a reduced rate from an unlucky sample, and ideally the previously-hitting cases replayed.

Ten cases against a one-in-twenty failure is like rolling a twenty-sided die ten times and concluding the 20 face does not exist because it never came up - it comes up about 40% of the time in ten rolls, so silence is the expected outcome, not evidence. And because the generator's cases cluster into near-duplicates, some of those rolls are the same roll made twice.

saying these in an interview costs you the question

  • Saying a clean ten-case run shows the failure was fixed.
  • Treating generated variants as independent samples with no allowance for near-duplicates.
  • Confusing the per-case rate with a per-user or per-session rate when quoting the risk.
  • Picking a case count with no statement of what rarity it can and cannot detect.

context