A promptfoo red-team plugin generates 10 cases against your assistant, and the behaviour it probes appears in roughly 1 relevant response in 20. Roughly what chance does that run have of reporting zero hits, and how many cases would you need before a clean result means something?
answer
- 1 - (1 - p)^n
- p=0.05, n=10 misses ~60%
- ~59 cases for 95% at 1-in-20
- generated cases cluster, effective n lower
- size to the rarity you must catch
basics
~20 sTreating each case as an independent try, the chance of missing every time is 0.95 to the tenth, about 60 percent. So a clean run of ten mostly reflects the sample size. To be roughly 95 percent sure of seeing a one-in-twenty behaviour you need about sixty cases.
solid answer
~50 sThe arithmetic is the geometric tail: the probability of at least one hit is `1 - (1 - p)^n`. With p = 0.05 and n = 10 that is about 40%, so most runs report nothing even though the failure is genuinely there. Inverting, `n >= ln(0.05) / ln(1 - p)` gives roughly 59 cases for 95% confidence at that rarity, and around 300 if the true rate is one in a hundred. That figure is an **upper bound on your real power**. Generated cases are not independent draws: twenty variants of one harm often include near-duplicate phrasings, covering one narrow path repeatedly rather than twenty paths once. The target is stochastic too, so a case that failed once can pass on replay. Both push effective n below the configured one. So a small per-plugin count screens for common, easy failures and cannot be reported as evidence about rare ones.
go deeper
Should recognise that ten tries is few for a one-in-twenty event and that a clean run is not proof, even without the exact formula.
Should produce 1 - (1 - p)^n, compute roughly 60% miss at n = 10, and invert it to about sixty cases for 95% confidence.
Should add why the figure is optimistic - correlated generated cases and a stochastic target - and turn the count into an explicit statement of what the run can rule out.
Should treat the target detectable rarity as the thing set per harm class and hold reporting to the claim that count supports.
### The model, stated plainly Treat each generated case as one independent trial that either surfaces the behaviour or does not, with per-case probability p - the rate at which your application commits the behaviour on prompts of that shape. The run misses entirely with probability `(1 - p)^n`, so it fires at least once with probability `1 - (1 - p)^n`. This is the geometric tail, and it is the single most useful piece of arithmetic in suite sizing because it converts an opinion about volume into a statement about what a clean report is *allowed to claim*. At p = 0.05 (one in twenty): | cases (n) | chance of zero hits | chance of at least one hit | |---|---|---| | 5 | 77% | 23% | | 10 | 60% | 40% | | 30 | 21% | 79% | | 60 | 5% | 95% | | 90 | 1% | 99% | So ten cases against a one-in-twenty behaviour report nothing about 60% of the time. Inverting for the count you need: `n = ln(alpha) / ln(1 - p)`, where alpha is the miss probability you will accept. With alpha = 0.05 and p = 0.05 that is `ln(0.05)/ln(0.95)` = 58.4, so about sixty cases. Required n scales close to `1/p`, so a one-in-a-hundred behaviour needs roughly 300 cases for the same confidence, and one in a thousand needs roughly 3,000. ### What that costs The count is not free and the scaling is brutal at the rare end. In promptfoo each case is a target call plus, for the model-graded plugins, a grader call, and `redteam.strategies` multiplies the case list on top. Going from 10 to 60 on one plugin adds ~100 calls per run for that plugin alone; doing it across twenty plugins is ~2,000 extra calls, every run, and the same again on the next scheduled run. That is why "just raise numTests" is a budget decision and not a config tweak - and why chasing one-in-a-thousand rarity with a generated suite is usually the wrong instrument entirely. ### Where the number misleads **Independence is the weak assumption, and it fails in your favour on paper only.** promptfoo asks its generator for n cases for one plugin from one purpose description; the results cluster into a handful of phrasings of the same idea. Sixty configured cases may be fifteen distinct probes repeated four ways. The computed 95% is therefore an **upper bound on your real detection power**, not an estimate of it, and the gap widens as n grows because the generator runs out of genuinely different angles long before it runs out of quota. **The target is stochastic.** A case is not a fixed pass or fail; the same case can hit tonight and pass tomorrow at the same temperature. The formula still works if you read p as the average per-case hit probability, but it means a single hit is not proof of a reliable failure and a single clean replay is not proof of a fix. **p is unknown - that is what you are trying to learn.** So this is a planning tool run backwards from the rarity you have decided you must catch, never a measurement of your actual risk. Quoting "we detect 95% of failures" from it is a category error. **The denominator swap.** p here is per relevant case. It is not per user, per session or per day. A one-in-twenty per-prompt rate on an assistant handling ten thousand prompts a day is a near-certain daily occurrence, and reporting it as "5%" to a stakeholder invites the wrong sense of scale in the other direction. **Many plugins, many chances.** Run twenty plugins and you are running twenty tests at once. Even a grader with a small false-positive rate will produce occasional spurious hits somewhere in the suite, so a single isolated hit across a large suite deserves replay before it becomes a finding. ### What I would check Before believing a clean run: what n actually ran for that plugin; how many of those n cases are meaningfully distinct; and whether the previously-hitting cases were among them. Before believing a hit: replay it several times to estimate its own rate rather than reporting a one-off as a live failure. And in the report, next to every clean plugin, print the sentence the count earns - at n = 10 this run rules out behaviours much more common than one in ten and nothing rarer - so nobody upgrades a thin sample into an assurance on your behalf.
- The true rate is one in a hundred rather than one in twenty. How does the required count change?Roughly five times larger: about 300 cases for the same 95% chance of at least one hit, since the required n scales close to 1/p.
- Why is the computed count optimistic in practice?Generated cases are correlated - several near-identical phrasings of the same attack shape - so the effective number of distinct tries is below the configured one, and a stochastic target adds further variance.
- The suite hit twice last week and zero times this week after a fix. Is the fix confirmed?Not from that alone. At a small count, going from two hits to zero is well within noise; you need enough cases to distinguish a reduced rate from an unlucky sample, and ideally the previously-hitting cases replayed.
Ten cases against a one-in-twenty failure is like rolling a twenty-sided die ten times and concluding the 20 face does not exist because it never came up - it comes up about 40% of the time in ten rolls, so silence is the expected outcome, not evidence. And because the generator's cases cluster into near-duplicates, some of those rolls are the same roll made twice.
saying these in an interview costs you the question
- Saying a clean ten-case run shows the failure was fixed.
- Treating generated variants as independent samples with no allowance for near-duplicates.
- Confusing the per-case rate with a per-user or per-session rate when quoting the risk.
- Picking a case count with no statement of what rarity it can and cannot detect.