skip to content

A promptfoo red-team suite regenerates and re-runs nightly against a metered chat endpoint, and its bill has become the argument for deleting it. Someone proposes cutting the per-plugin case count to two across the board. What do you do instead?

level: seniorimportance: should knowfreq 40%

answer

  1. uniform cut = pay for nothing
  2. regression tier vs depth tier
  3. reuse cases, stop nightly regeneration
  4. keep every case that ever hit
  5. print per-plugin denominators

basics

~20 s

Do not thin everything evenly - that leaves every harm class too small to conclude anything while still costing money. Split the schedule: a small nightly tier on the surfaces that change, and a deep sweep with high case counts before releases. Stop regenerating cases you could reuse, and label the thin tier as a smoke test.

solid answer

~50 s

A uniform cut is the worst option available: it keeps paying for a suite while destroying the property that made it worth paying for. Two cases per plugin detect nothing but a near-constant failure, so you buy a green dashboard that means nothing. Separate the two jobs the nightly run is doing. A **regression tier** repeats a small, mostly fixed case set each night - especially cases that have hit before - and is cheap because it is narrow and reused rather than regenerated. A **depth tier** runs high case counts on a slower cadence, tied to releases, where the arithmetic actually supports a claim about rare failures. Then attack cost per case: nightly regeneration is billed generation plus fresh graded traffic, and it makes night-to-night comparison noisier, so reusing a stored case set saves money and variance at once. Finally, print the per-plugin denominator so nobody reads the thin tier as coverage.

go deeper

for a junior

Should at least notice that two cases per plugin cannot detect much and that the cheap suite would still cost money.

for a middle

Should propose different cadences for cheap and deep runs and identify generation plus grading as parts of the bill.

for a senior

Should design the split concretely, keep previously-hitting cases as a permanent set, allocate counts by consequence, and change the reporting so the thin tier is not misread.

for a principal

Should reframe the argument as cost per unit of detection and set the policy for what a release is allowed to cite as evidence.

### First, name the cost structure you are being asked to cut A nightly `promptfoo redteam run` bills three distinct things. **Generation**: a model writes fresh cases for every enabled plugin, paid in full every night because `run` regenerates rather than reusing. **Target traffic**: one call per case against the metered chat endpoint, where the case list is `plugins x redteam.numTests` expanded again by everything under `redteam.strategies`. **Grading**: for the harm-shaped plugins, another model call per response to decide whether it counted as a hit. Whoever proposes "two per plugin" is almost always reasoning about the second bucket only, and will be surprised when the bill drops by less than they promised because generation is a per-plugin fixed-ish cost and grading tracks cases one-for-one. Work the numbers before arguing. Twenty plugins at `numTests: 10` with three strategies is roughly 600 cases, about 1,200 metered calls a night, ~36,000 a month. Cutting `numTests` to 2 makes it ~120 cases and ~240 calls - an 80% saving, and that is the honest part of the proposal. ### Why the uniform cut is still the worst option on the table Two cases per plugin detects essentially nothing. From `1 - (1 - p)^n`, two cases catch a one-in-twenty behaviour about 10% of the time and a one-in-five behaviour about 36% of the time. Only a near-constant failure reliably fires. So you keep paying every night - a fifth of a large bill is still a bill - and you buy a dashboard that is green because it cannot see, which is strictly worse than no dashboard, because a green board gets cited in release decisions. The correct framing to put in front of whoever proposed it is **cost per unit of detection**, and by that measure a uniform cut is the one change that makes the number worse rather than better. ### What I would do instead: split the two jobs the nightly run is conflating The suite is doing two different things that want opposite sizing. A **regression tier** answers "did last night's deploy break something we already know about?" Its value is comparability, not volume: a small, *stored, reused* case set - generate once with `promptfoo redteam generate`, then run it with `promptfoo eval` each night - so the questions stay fixed and only the answers move. Every case that has ever hit lives here permanently. These are the highest-yield cases you own: few, cheap, and proven capable of firing. A **depth tier** answers "does this rare harm occur at all?" That needs the sixty-plus cases the arithmetic demands, but it does not need daily cadence. Weekly, or gated on release, is usually right - and it is where regeneration earns its cost, because fresh phrasings are what stop the suite from over-fitting to one generator's habits. Then allocate rather than flatten: high counts on the harm classes where a miss is actually expensive for *this* application, token counts where the plugin is enabled for completeness so the class keeps a visible denominator. ### Where the numbers mislead after you change this **The green nightly.** Once thinned, the nightly tier must be named a smoke test in the report and in the release checklist, with the release claim pointing at the depth run. Otherwise the only visible artefact is a passing suite whose sample size nobody re-reads. **Regeneration variance charged to the model.** Nightly regeneration changes the questions as well as the answers, so a pass rate moving from 96% to 91% may be a harder batch of generated cases, not a regression. Teams routinely open incidents on that. Freezing the regression tier's case set removes the confound and the generation bill together. **The savings estimate.** Any projection built from target calls alone understates the cases-scale part (grading) and overstates the fixed part (generation), so it will not match the invoice. ### What I would not do I would not drop plugins to save money: that removes a harm class's denominator entirely, which reports worse than a small sample. I would not disable grading to halve the calls - nothing then decides whether a response was a hit, so you buy traffic and no findings. And I would not let the thinned suite keep its old reporting language. ### What I would check The invoice split by generation, target and grader; the per-plugin case counts as actually executed, not as configured; whether every previously-hitting case is in the nightly set; and whether the last release cited the nightly run or the depth run as its evidence.

  • Why does reusing a fixed case set for the nightly tier help beyond cost?
    It makes results comparable night to night. Regenerating changes the questions as well as the answers, so a moving pass rate could be new cases rather than new behaviour.
  • What is the cheapest set of cases to keep running forever?
    The ones that have hit before. They are few, they are proven capable of firing, and they are the direct check that a fix held.
  • How do you keep the thinned nightly tier from being read as assurance?
    Report the per-plugin case count next to every result and name the tier a smoke test, with the release claim pointing at the depth run instead.

Pressing the test button on a smoke alarm every morning is worth doing and costs nothing; it does not tell you the building would survive a fire. Cutting every plugin to two cases turns the whole suite into the test button while still charging you for the fire drill.

saying these in an interview costs you the question

  • Cutting the case count uniformly and continuing to report the nightly run as coverage.
  • Dropping whole plugins to save money, leaving a harm class with no denominator at all.
  • Estimating savings from target calls only, ignoring generation and grading traffic.
  • Keeping nightly regeneration while blaming the resulting score movement on the model.

context