skip to content

Suite as an Artifact

A generated suite is data, and data that quietly regenerates makes last month's number meaningless. Interviewers ask because trend lines get quoted long after anyone remembers how they were produced.

on this pageshow

explore

questions

4

promptfoo's red-team mode generates adversarial test cases from its plugin and strategy catalogue before running them. What changes if you write that generated suite to a file and commit it, versus letting every CI run generate a fresh one?

level: middleimportance: must knowfreq 46%

answer

  1. generate then evaluate = two phases
  2. generation is a model call, non-deterministic
  3. pinned = stable denominator, diffable
  4. frozen suite invites overfitting
  5. pinned baseline + exploratory run, reported apart

basics

~20 s

Generation is itself a model call, so a fresh run produces different test cases. Committing the generated suite makes it a fixed artifact: runs become comparable and diffable, and a failure is reproducible. Regenerating every run measures a new suite each time, so month-over-month numbers move for reasons nobody can attribute.

solid answer

~50 s

Red-team generation in promptfoo is two phases: an LLM expands the configured plugins and strategies into concrete adversarial test cases, then those cases are executed against the target and graded. Only the second phase is what people think they are measuring. Commit the generated suite and you get a stable denominator: the same cases, in the same count, every run. A failure comes with the exact input that produced it, a reviewer can diff what changed when the suite is refreshed, and two runs differ only in the target. The cost is drift in the other direction — a frozen suite ages, stops reflecting new plugin coverage, and a target can be tuned until it passes those specific cases without being safer. The usual resolution is both: a pinned suite carries the trend line, and a separate regenerated run explores. They are never averaged into one number.

go deeper

for a junior

Knows promptfoo generates red-team test cases rather than shipping a fixed list, and that saving them to a file lets you rerun the same ones.

for a middle

Separates generation from evaluation, explains that generation is a model call and therefore variable, and names both the comparability gain and the staleness cost of pinning.

for a senior

Runs a pinned regression suite plus a separate exploratory generation, keeps refreshes as reviewed commits, and treats a rising pinned score with a flat exploratory score as an overfit signal.

for a principal

Sets the policy for who approves a suite refresh, how the two run types are reported, and how the repository handles a committed file of adversarial content.

### The two phases, and which one people think they are measuring promptfoo's red-team mode splits into generation and evaluation. `promptfoo redteam generate` reads your red-team configuration — the `purpose` or application description, the `plugins` list (harm categories: what to test for) and the `strategies` list (transformations: how to deliver it), with `numTests` controlling how many cases each plugin contributes — and calls a model to expand all of that into concrete adversarial test cases, writing them to a file (`redteam.yaml` by default, or whatever `--output` names). Evaluation then sends each case to the `providers` under test and applies the per-plugin graders. `promptfoo redteam run` does both in one command, which is exactly the shape that gets pasted into CI and causes the problem. Generation is a model call. Two runs of the same configuration produce different wording, different scenarios, and — because strategies expand base cases and some expansions are conditional — sometimes a different total count. The configuration is deterministic; the suite it produces is not. ### What pinning buys - **A stable denominator.** The same cases in the same number every run, so a rate means something across runs. - **Reproducibility.** A reported failure names a concrete case a developer can rerun locally in seconds. - **Reviewability.** A refresh becomes a diff someone approves, not an invisible change in the instrument. - **Attribution.** When the number moves, the suite is off the list of candidate causes, and the target is the only thing left. ### What pinning costs A frozen adversarial suite is a fixed target, and prompts get tuned until those exact strings stop working. The score improves, real robustness does not — the ordinary overfit-to-the-eval failure, and it is fast, because a system prompt edit that hard-codes a refusal for the pinned scenarios is a day's work. The suite also stops representing what the plugin and strategy catalogue has grown to cover since it was written, so its coverage claim silently expires. And it is a committed file of adversarial content: it needs the same access thinking as any other sensitive artefact, and reviewers need to know that a large diff of attack strings is expected rather than alarming. ### What each option costs Regenerating every run pays for generation calls on top of evaluation calls, every night, forever, plus the latency and the external dependency if generation runs through a hosted generation service rather than locally. Pinning removes the generation phase from the nightly job entirely, which is a real but modest saving — the evaluation calls dominate — and removes most of the run-to-run variance in wall clock and case count. The genuine cost of pinning is not money; it is the discipline of scheduling and reviewing refreshes, which is engineer time nobody budgets for. ### Where the number misleads With a regenerated suite, a month-over-month delta has two causes and no way to separate them. Worse, the suite-only noise is not small: plugins sample a space of possible attacks, so two generations of an identical configuration against an *unchanged* target can differ by several points of failure rate purely from which scenarios happened to be written. A three-point improvement quoted from such a job is comfortably inside the noise of the instrument, and reading it as progress is the specific mistake. With a pinned suite, the misleading reading is the opposite one: a pass rate on a frozen file is not a probability that a real attacker fails. It is the fraction of a few hundred model-written strings that one grader declined to flag. "We pass 100% of the red-team suite" means those strings no longer work, and nothing at all about strings nobody generated. ### What you check Before believing a delta, hash the suite file and confirm both runs used the same one, then confirm the case counts are identical; if you regenerated deliberately, run old and new suites the same day so the level shift is visible. Before believing a pinned improvement, run a freshly generated exploratory suite against the same target: if it still finds failures at roughly the old level while the pinned score climbs, you measured tuning, not safety. ### The operating pattern Keep a pinned suite as the regression baseline, refreshed as a reviewed commit on a deliberate cadence. Run a separate, freshly generated exploration job whose purpose is finding *new* categories of failure, and promote what it finds into the pinned suite as added cases rather than replacing it wholesale. Report the two apart: the pinned suite answers "did we regress", the exploratory run answers "did we miss something". One blended number answers neither.

  • The pinned promptfoo suite's pass rate has climbed for four months. What would make you distrust it?
    The suite has been fixed long enough for prompt changes to be tuned against those exact cases. Check against a freshly generated exploration run; if that one has not improved, the trend is overfitting, not robustness.
  • How do you add findings from an exploratory promptfoo run to a pinned suite without breaking the trend?
    Append the new cases as an explicit, reviewed change, record the suite version, and mark the point on the chart where the case count changed rather than back-fitting earlier numbers.

saying these in an interview costs you the question

  • Believing the generated suite is deterministic given the same configuration
  • Plotting a trend from runs that regenerate their own test cases
  • Replacing the pinned suite wholesale on every refresh and keeping the old chart
  • Treating a rising pass rate on a long-frozen suite as evidence of real robustness

context

open as a page

A nightly promptfoo red-team job finishes in seconds and reports the same failure count as last month, even though the team shipped a new model behind the same endpoint. How does promptfoo's response cache produce that result, and when is leaving the cache on still the right call?

level: middleimportance: must knowfreq 52%

basics

~20 s

promptfoo caches responses keyed by the request it sent. If the test cases and the provider settings did not change, a rerun replays stored answers instead of calling the new model, so the numbers still describe the old deployment. Keep the cache while iterating on graders; clear or disable it whenever the system under test changed.

open as a page

A team quotes a monthly trend of failures from a scheduled promptfoo red-team job, and this month it dropped by half. The report does not say whether the suite was regenerated or whether responses were replayed from cache. What do you check to explain the drop, and what should each run record so the question is answerable next time?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Check whether the graded population changed: case count and case identity, then whether responses were served from cache, then whether the grader or its judge model changed. Only after those hold constant does a target change explain the drop. Each run should record suite identity, case count, target identity and cache mode alongside the number.

open as a page

You own a pinned, committed promptfoo red-team suite that several teams run against different applications. How do you decide how often it is regenerated, and what do you do so numbers reported before and after a regeneration remain meaningful?

level: principalimportance: should knowfreq 24%

basics

~20 s

Refresh on events, not on a calendar alone: a new application capability, a catalogue update, or evidence that scores are being tuned against the frozen cases. Treat each refresh as a reviewed commit, run old and new suites together once to bridge the series, and break the trend line at the boundary rather than interpolating across it.

open as a page