You own promptfoo red-team suites for a dozen applications under one fixed monthly spend. How do you decide the per-plugin case counts across that portfolio, and why is a single uniform number the wrong answer?
answer
- fixed recurring budget = allocation problem
- rarity sets count, consequence sets worth
- uniform overspends and under-samples at once
- denominator published with every result
- re-allocate after each finding
basics
~20 sTreat it as allocating a fixed recurring budget, not picking a setting. Give high counts to the harm classes where a miss is expensive and to apps with real exposure, and thin counts elsewhere. A uniform number overspends on harms an app cannot commit while under-sampling the ones it can, and hides both behind one green score.
solid answer
~50 sThe budget is fixed and recurring, so every case added somewhere is a case removed somewhere else. That makes this an allocation problem with two inputs per harm class: how rare a failure you would still need to catch, and what a miss would cost. The case count follows from those, not from a house default. A uniform number is wrong in both directions at once. On harm classes an application barely touches you buy cases that never fire; on the ones tied to its real exposure you buy too few to support any claim. Uniformity also yields one comparable-looking score per app, inviting a portfolio ranking that is really a ranking of sample thinness. So size per harm class per app, publish the per-plugin denominator with every result, permanently fund the cases that have ever hit, and re-allocate after each finding.
go deeper
Should recognise that different applications carry different risk and that one number for all of them is a simplification.
Should connect the count to the rarity it can detect and propose spending more on higher-risk harm classes.
Should design the allocation with published denominators, permanently funded previously-hitting cases, and re-allocation driven by findings.
Should own it as recurring budget policy: what claims each count entitles a team to make, how cuts are absorbed without uniform thinning, and why a portfolio ranking over unequal samples is misleading.
### Reframe it: this is spend, not configuration A dozen scheduled promptfoo suites under a fixed monthly ceiling is a subscription, not a settings file. Every case added to one plugin on one application is a case removed from somewhere else in the portfolio. So the question is never "what should `redteam.numTests` be" but "given this recurring budget, where does the next thousand cases buy the most detection?" That reframing is most of the answer, and it is the thing that separates someone who owns a programme from someone tuning a YAML file. Ground it in the actual cost model first, because the allocation is meaningless without it. Per application, one run costs generation plus, for every case, a target call and (for model-graded plugins) a grader call - with the case list being `plugins x numTests` expanded again by `redteam.strategies`. Twelve applications at 20 plugins, `numTests: 10` and three strategies is roughly 7,200 cases and ~14,400 metered calls per cycle; nightly, that is comfortably six figures of calls a month plus the triage hours behind every hit. Those hours are the scarcer budget and are the reason breadth is not free even when calls are cheap. ### The two inputs that set each count For each harm class on each application: 1. **The rarity you must be able to detect.** This sets the count arithmetically, from `1 - (1 - p)^n`: about 60 cases to catch a one-in-twenty behaviour with 95% confidence, about 300 for one in a hundred, since required n scales close to `1/p`. 2. **The consequence of missing it.** This decides whether that price is worth paying at all. The same plugin can therefore deserve 200 cases on the customer-facing agent that holds tool access and 5 on the internal drafting assistant - not because the harm differs, but because the exposure and the blast radius do. ### Why one uniform number fails, in both directions at once Uniformity optimises for looking fair rather than for detecting anything. On harm classes an application barely touches you buy cases that will never fire, month after month. On the classes tied to its real exposure you buy too few to support any claim, and the report says "no findings" in both situations with identical formatting. So the money lands where the risk is not, and the two errors cancel out visually into one clean dashboard. ### Where the portfolio numbers mislead **The comparability illusion.** Equal counts across applications produce per-application pass rates that *look* comparable and are not. The denominators match; the underlying failure rates, traffic volumes, plugin relevance and exposure do not. Ranking twelve applications on those rates ranks sample thinness and plugin selection at least as much as safety, and the app at the top is often the one whose enabled plugins least resemble what it actually does. **Aggregation over unequal samples.** A single portfolio-wide "safety score" averaged across suites with different counts is arithmetic performed on incommensurable quantities. It will move when someone changes a config, and it will be read as the risk moving. **Silent uniform thinning under budget pressure.** This is the failure mode to watch for by name: a cut arrives, every suite is trimmed evenly because that is the defensible-looking move, the reporting language does not change, and the portfolio stays green while measuring progressively less. **Breadth bought before depth is paid for.** Adding plugins and applications is visible progress; it is the wrong purchase while existing counts cannot support the claims teams are already making from them. ### The policy I would actually write Publish the executed per-plugin case count beside every result, so a clean plugin carries its own caveat and no one can quote it without quoting n. Permanently fund a small always-run set of previously-hitting cases on every application - cheap, proven able to fire, and the only direct evidence a fix held. Re-allocate on evidence after each cycle: a plugin that has hit twice earns more cases; one that has never fired in a year at a genuine count is a candidate for thinning, provided it still matches a real surface so the class keeps a visible denominator. State explicitly what claim each count entitles a team to make, and require the release gate to cite a depth run rather than a nightly smoke tier. ### What I would check each cycle Executed counts versus configured counts per application; which suites moved because their config changed rather than their behaviour; whether any harm class lost its denominator entirely to a budget cut; and whether the previously-hitting set is still running everywhere.
- A plugin has never fired in a year at a reasonable count. More cases or fewer?Usually fewer, and the saving moves to a harm class that has fired - provided the plugin still matches a surface the app has, so the class keeps a visible denominator rather than disappearing.
- Leadership wants one portfolio safety number. What do you give them?Findings by harm class with the case count behind each, plus what rarity each count can rule out. A single aggregate over unequal samples is a ranking of sample sizes.
- Budget is cut 30%. What is the first thing you protect?The permanently funded cases that have hit before, and the depth counts on the harm classes with real consequence; breadth on speculative classes is where the cut lands.
saying these in an interview costs you the question
- Picking one house-wide number so every team is treated the same, with no reference to exposure.
- Ranking applications on pass rates produced at incomparable sample sizes.
- Absorbing budget cuts by thinning every suite uniformly and keeping the same reporting language.
- Spending only on new breadth while existing counts are too small to support the claims already made.