skip to content

In a promptfoo red-team configuration, what does the per-plugin case count (numTests) actually control, and what does leaving it at a small default mean for a failure your app commits only rarely?

level: juniorimportance: must knowfreq 62%

answer

  1. cases per plugin, not per suite
  2. plugins x count = volume
  3. each case = target call + grading call
  4. zero hits at low n = not measured
  5. paid again every scheduled run

basics

~20 s

It sets how many adversarial cases promptfoo generates for each enabled plugin, so total volume scales with plugins times that count. A small default gives each harm class only a handful of tries, so a failure the app commits rarely can easily never fire, and the run then reports a clean pass.

solid answer

~50 s

promptfoo's red-team generation produces adversarial cases **per enabled plugin**, and the per-plugin count is the multiplier on everything else in the configuration. Rough volume is plugins x cases, and each case costs at least one call to the target under test plus a call to whatever grades the response, so the number sets both your detection power and your bill. The important consequence is what a low count does to interpretation. With a few cases per plugin, a behaviour the app exhibits maybe one response in twenty will usually not fire at all, and the report shows that plugin with no hits. That is *not measured*, not *safe*. The honest reading of a zero-hit plugin at a low count is that the sample was too thin to conclude anything. Raise the count on the harm classes that would actually hurt you, and accept a thin sample only where you are explicitly buying breadth rather than evidence.

go deeper

for a junior

Should say it is the number of generated cases per plugin, and that a low number means a rare failure may simply never be tried.

for a middle

Should add the multiplication against the plugin set and the per-case cost of target plus grading, and read a zero-hit plugin as under-sampled rather than clean.

for a senior

Should talk about the recurring cost of a scheduled suite and about sizing per harm class by consequence instead of setting one global number.

for a principal

Should frame it as allocating a fixed recurring budget across harm classes, and insist the per-plugin denominator appears wherever the results are reported.

### Where the number lives In a promptfoo red-team configuration the case count is `redteam.numTests` in the config file that `promptfoo redteam init` scaffolds and `promptfoo redteam generate` reads, and any individual entry under `redteam.plugins` may override it with its own `numTests`. The thing to internalise first is that this is a **per-plugin quota, not a suite total**. promptfoo asks its generator model for that many adversarial cases for *each* enabled plugin, so the generated suite is roughly `plugins x numTests` cases before anything else touches it. Everything listed under `redteam.strategies` then takes those base cases and rewrites them into further variants, so a strategy list multiplies the same number again rather than adding a fixed amount to it. Twenty plugins at ten cases with three strategies enabled is not thirty-something of anything; it is on the order of six hundred cases. ### What one case costs A case is not one API call. `promptfoo redteam run` does two separately billed things in sequence. **Generation** calls a model to write the cases, and it is paid again on every regeneration, which is exactly why splitting the flow into `promptfoo redteam generate` once and `promptfoo eval` many times matters when you intend to reuse a case set. **Execution** then sends each case to the provider under test - one target call - and hands the response to that plugin's grader, which for the harm-shaped plugins is itself a model call rather than a regex. So the run-cost arithmetic is roughly: ``` calls ~= plugins x numTests x strategy_multiplier x (1 target + 1 grader) + generation ``` Six hundred cases is therefore something like twelve hundred metered calls for one run. On a nightly schedule that is thirty-odd thousand calls a month, at whatever your target and grader models charge, plus wall-clock governed by your concurrency setting and the target's rate limits, plus the cost nobody budgets: every case that hits is a finding a human has to read, reproduce and triage. Raising `numTests` raises the triage load in the same proportion as the bill. ### What the count actually buys Detection probability, and nothing else. If the application commits the probed behaviour on some fraction p of prompts of that shape, a run of n cases sees it at least once with probability `1 - (1 - p)^n`. At p = 0.05 that is: five cases miss 77% of the time, ten miss about 60%, thirty miss 21%, sixty miss roughly 5%. The count is therefore the sample size sitting behind every per-plugin verdict in the report, whether or not the report shows it. ### Where the number misleads Three misreadings, in descending order of how often they happen. **Zero hits read as safe.** The report renders a plugin with no findings identically whether it ran four cases or four hundred. The denominator is not in the headline, so "no issues found" for a harm class sampled five times is presented with the same visual weight as one sampled two hundred times. The honest reading of a zero-hit plugin at a low count is *not measured*, not *clean*. **Effective n below configured n.** The generator writes all n cases for one plugin from one description of your application, and they cluster. Sixty configured cases can be twenty distinct probes phrased three ways each, which is not sixty independent tries at anything. The formula above is an upper bound on your real power, never an estimate of it. **The inflated total.** "600 tests passed" reads like six hundred independent probes. It is twenty harm classes crossed with mechanical rewrites. Strategy expansion inflates the summary number far faster than it broadens the risk surface, and that number is the one that ends up in a slide. ### What I would check before believing a clean run Open the generated case file and count cases per plugin, because a per-plugin `numTests` override is easy to set and easier to forget. Skim the case texts for near-duplicates to sanity-check effective n. Confirm the graders actually ran, since a grader error can register as a non-hit rather than as an error in some setups. Re-run the cases that hit on a previous run - they are the only cases in your possession proven capable of firing. Then write down the claim the count supports: at n cases, a clean plugin rules out behaviours more common than roughly one in n, and says nothing whatsoever about anything rarer.

  • Why is total volume not simply the per-plugin count?
    It is that count multiplied by the number of enabled plugins, and further multiplied by anything layered on top of the generated cases. Doubling the count doubles the whole suite, not one plugin.
  • A plugin reports zero hits. What single number do you ask for before believing it?
    How many cases that plugin actually ran. Zero hits out of four says nothing; zero out of sixty is weak but real evidence.
  • Where does the money actually go in one run?
    Generating the cases, calling the target once per case, and calling the grader once per response. Cutting the case count cuts all three at once, which is why it is the first knob people reach for.

numTests is not a line item, it is a multiplier applied to every plugin and then multiplied again by every strategy - like changing the serving size on a recipe rather than adding one dish. Nudging it from 5 to 20 does not add fifteen cases; it adds fifteen per plugin, per strategy, on every future run.

saying these in an interview costs you the question

  • Reading a zero-hit plugin as proof the app is safe for that harm class, with no reference to how many cases ran.
  • Assuming one case equals one API call, so the cost estimate ignores generation and grading traffic.
  • Treating the shipped default as a tested, principled number rather than a starting point.
  • Raising the count uniformly across every plugin because 'more coverage is better', with no budget arithmetic.

context