skip to content

promptfoo's red-team mode generates adversarial test cases from its plugin and strategy catalogue before running them. What changes if you write that generated suite to a file and commit it, versus letting every CI run generate a fresh one?

level: middleimportance: must knowfreq 46%

answer

  1. generate then evaluate = two phases
  2. generation is a model call, non-deterministic
  3. pinned = stable denominator, diffable
  4. frozen suite invites overfitting
  5. pinned baseline + exploratory run, reported apart

basics

~20 s

Generation is itself a model call, so a fresh run produces different test cases. Committing the generated suite makes it a fixed artifact: runs become comparable and diffable, and a failure is reproducible. Regenerating every run measures a new suite each time, so month-over-month numbers move for reasons nobody can attribute.

solid answer

~50 s

Red-team generation in promptfoo is two phases: an LLM expands the configured plugins and strategies into concrete adversarial test cases, then those cases are executed against the target and graded. Only the second phase is what people think they are measuring. Commit the generated suite and you get a stable denominator: the same cases, in the same count, every run. A failure comes with the exact input that produced it, a reviewer can diff what changed when the suite is refreshed, and two runs differ only in the target. The cost is drift in the other direction — a frozen suite ages, stops reflecting new plugin coverage, and a target can be tuned until it passes those specific cases without being safer. The usual resolution is both: a pinned suite carries the trend line, and a separate regenerated run explores. They are never averaged into one number.

go deeper

for a junior

Knows promptfoo generates red-team test cases rather than shipping a fixed list, and that saving them to a file lets you rerun the same ones.

for a middle

Separates generation from evaluation, explains that generation is a model call and therefore variable, and names both the comparability gain and the staleness cost of pinning.

for a senior

Runs a pinned regression suite plus a separate exploratory generation, keeps refreshes as reviewed commits, and treats a rising pinned score with a flat exploratory score as an overfit signal.

for a principal

Sets the policy for who approves a suite refresh, how the two run types are reported, and how the repository handles a committed file of adversarial content.

### The two phases, and which one people think they are measuring promptfoo's red-team mode splits into generation and evaluation. `promptfoo redteam generate` reads your red-team configuration — the `purpose` or application description, the `plugins` list (harm categories: what to test for) and the `strategies` list (transformations: how to deliver it), with `numTests` controlling how many cases each plugin contributes — and calls a model to expand all of that into concrete adversarial test cases, writing them to a file (`redteam.yaml` by default, or whatever `--output` names). Evaluation then sends each case to the `providers` under test and applies the per-plugin graders. `promptfoo redteam run` does both in one command, which is exactly the shape that gets pasted into CI and causes the problem. Generation is a model call. Two runs of the same configuration produce different wording, different scenarios, and — because strategies expand base cases and some expansions are conditional — sometimes a different total count. The configuration is deterministic; the suite it produces is not. ### What pinning buys - **A stable denominator.** The same cases in the same number every run, so a rate means something across runs. - **Reproducibility.** A reported failure names a concrete case a developer can rerun locally in seconds. - **Reviewability.** A refresh becomes a diff someone approves, not an invisible change in the instrument. - **Attribution.** When the number moves, the suite is off the list of candidate causes, and the target is the only thing left. ### What pinning costs A frozen adversarial suite is a fixed target, and prompts get tuned until those exact strings stop working. The score improves, real robustness does not — the ordinary overfit-to-the-eval failure, and it is fast, because a system prompt edit that hard-codes a refusal for the pinned scenarios is a day's work. The suite also stops representing what the plugin and strategy catalogue has grown to cover since it was written, so its coverage claim silently expires. And it is a committed file of adversarial content: it needs the same access thinking as any other sensitive artefact, and reviewers need to know that a large diff of attack strings is expected rather than alarming. ### What each option costs Regenerating every run pays for generation calls on top of evaluation calls, every night, forever, plus the latency and the external dependency if generation runs through a hosted generation service rather than locally. Pinning removes the generation phase from the nightly job entirely, which is a real but modest saving — the evaluation calls dominate — and removes most of the run-to-run variance in wall clock and case count. The genuine cost of pinning is not money; it is the discipline of scheduling and reviewing refreshes, which is engineer time nobody budgets for. ### Where the number misleads With a regenerated suite, a month-over-month delta has two causes and no way to separate them. Worse, the suite-only noise is not small: plugins sample a space of possible attacks, so two generations of an identical configuration against an *unchanged* target can differ by several points of failure rate purely from which scenarios happened to be written. A three-point improvement quoted from such a job is comfortably inside the noise of the instrument, and reading it as progress is the specific mistake. With a pinned suite, the misleading reading is the opposite one: a pass rate on a frozen file is not a probability that a real attacker fails. It is the fraction of a few hundred model-written strings that one grader declined to flag. "We pass 100% of the red-team suite" means those strings no longer work, and nothing at all about strings nobody generated. ### What you check Before believing a delta, hash the suite file and confirm both runs used the same one, then confirm the case counts are identical; if you regenerated deliberately, run old and new suites the same day so the level shift is visible. Before believing a pinned improvement, run a freshly generated exploratory suite against the same target: if it still finds failures at roughly the old level while the pinned score climbs, you measured tuning, not safety. ### The operating pattern Keep a pinned suite as the regression baseline, refreshed as a reviewed commit on a deliberate cadence. Run a separate, freshly generated exploration job whose purpose is finding *new* categories of failure, and promote what it finds into the pinned suite as added cases rather than replacing it wholesale. Report the two apart: the pinned suite answers "did we regress", the exploratory run answers "did we miss something". One blended number answers neither.

  • The pinned promptfoo suite's pass rate has climbed for four months. What would make you distrust it?
    The suite has been fixed long enough for prompt changes to be tuned against those exact cases. Check against a freshly generated exploration run; if that one has not improved, the trend is overfitting, not robustness.
  • How do you add findings from an exploratory promptfoo run to a pinned suite without breaking the trend?
    Append the new cases as an explicit, reviewed change, record the suite version, and mark the point on the chart where the case count changed rather than back-fitting earlier numbers.

saying these in an interview costs you the question

  • Believing the generated suite is deterministic given the same configuration
  • Plotting a trend from runs that regenerate their own test cases
  • Replacing the pinned suite wholesale on every refresh and keeping the old chart
  • Treating a rising pass rate on a long-frozen suite as evidence of real robustness

context