You ran promptfoo's red-team mode before and after editing your application's system prompt, and the failure rate moved. What has to have stayed identical between the two runs for that difference to be evidence about the prompt?
answer
- one variable, frozen suite
- cases are generated, not fixed
- grader is a model too
- same target decoding settings
- regenerated = new denominator
basics
~20 sThe same generated adversarial test cases, the same promptfoo plugins and strategies that produced them, the same grader deciding pass or fail, and the same target model and decoding settings. Only the system prompt may differ. If the cases were generated again, the two runs attacked different inputs and the numbers are not comparable.
solid answer
~50 spromptfoo's red-team mode does not carry a fixed list of attack cases: the plugins (what harm or failure is probed) and the strategies (how the probe is dressed up) **generate** the prompts. So two runs of the same configuration are not automatically the same suite. For a prompt A/B you need: the generated cases saved and replayed rather than produced fresh; the same grader, which is itself model-driven and is the thing that decides a response counted as a failure; and the same target settings on both sides — model, temperature, tools, context, output limits. Then the delta has exactly one candidate cause. If anything else moved, say so instead of reporting the delta. A number produced from two different populations of generated attacks mixes "the prompt got better" with "this batch of attacks was easier", and you cannot separate the two after the fact.
go deeper
Should say the test cases and the grading must be the same and only the prompt should differ, and should know promptfoo's red-team cases are generated rather than shipped as a fixed list.
Adds the target settings (temperature, output limits, tools, context) and notices that the grader is itself a model whose configuration is part of the control.
Talks about running both configurations over one saved suite in a single run, recording the suite and grader alongside the number, and refusing to report a delta when more than one variable moved.
Frames it as experimental hygiene for a recurring safety metric: what is versioned, who may change it, and how a comparison is invalidated and re-baselined when the suite has to change.
### What promptfoo's red-team mode actually produces promptfoo's red-team mode ships **no fixed corpus of attacks**. You declare intent in the `redteam` block of promptfoo's `promptfooconfig.yaml`: `redteam.purpose` describes what your application is for, `redteam.plugins` lists the harms and failure modes to probe (a PII-leak plugin, a harmful-content plugin, an out-of-scope/overreach plugin, and so on), `redteam.strategies` lists how each base case gets dressed up (an encoding wrapper, a multi-turn escalation, an attacker-model rewriting loop), and `redteam.numTests` sets how many base cases each plugin is given. promptfoo's `redteam generate` command then calls an attacker model and writes the concrete cases it invented into a file — `redteam.yaml` by default. promptfoo's `redteam eval` sends those saved cases to whatever is listed under the config's `providers` key and grades each response. promptfoo's `redteam run` does both in one command, and that convenience is exactly the trap in an A/B: run it twice and you have *generated* twice. ### Three of the moving parts are models, not fixtures 1. **The attack generator.** Sampling, not enumeration. Regenerate and you get different wordings, a different mix across plugin categories, and often a different total count — so the numerator *and* the denominator of your failure rate both moved. 2. **The grader.** Something must decide whether a given response counted as a failure. For open-ended harm that is a model reading the transcript against a rubric (promptfoo lets you pin it via `defaultTest.options.provider`; leave it unset and you inherit a platform default that can change under you). Change the grading model or the rubric and the *same* transcripts get relabelled. 3. **The target.** The `providers` entry is not just a model id: `temperature`, `max_tokens`, tool access, retrieved context and any system message all change how the same attack lands. Only one thing is allowed to differ between the two runs — your application's system prompt. Everything above is the control. ### What the run costs, so you know what a rerun buys Multiply it out. Twenty plugins at `numTests: 10` is 200 base cases; three strategies typically produce their own variant of each base case, so the saved suite is closer to 800. Each single-turn case is one target call plus one grader call; each multi-turn case is a target call *per turn* plus the attacker model's own call per turn, so a turn-capped conversational strategy costs five to ten times a single-turn case. A mid-sized suite is therefore a few thousand model calls, single-digit to low-double-digit dollars at commodity API prices, and ten to forty minutes of wall clock at the concurrency promptfoo defaults to. Generation itself is a separate bill on top. The reason to save `redteam.yaml` and replay it is not only rigour — it is that the generation step is the part you are paying for twice. ### Where the number misleads The specific wrong reading is: *"failure rate went 14% to 11% after the prompt edit, so the prompt helped."* If the second run regenerated, those two percentages have different denominators (say 802 cases and 764), different per-plugin composition, and disjoint case sets — nothing about them is paired. "Three points better" is then partly the prompt and partly which attacks the generator happened to invent this morning, in unknown proportion, and no post-hoc analysis can split them because the evidence was never collected. Two quieter versions of the same error. First, the aggregate is an average over a heterogeneous suite: one plugin drawing an easier batch moves the headline while nothing about your application changed. Second, promptfoo caches provider responses keyed by prompt plus provider; because you edited the system prompt, the new side is generated fresh while the unchanged side may be replaying cached responses from weeks ago — one side is a fresh sample, the other a frozen one, and a hosted model may have been silently updated in between. ### What you check before believing your own number ```text same saved cases? identical file, identical count and per-plugin mix same grader? grading model pinned by id + same rubric text same target settings? temperature, max tokens, tools, retrieved context cache? both sides fresh, or both sides replayed - not one each errors/timeouts? similar counts, and counted the same way on both sides only one thing moved? name it in the report, in one sentence ``` If more than one of those moved, the honest output is "incomparable, rerunning" — not a percentage with a caveat stapled to it. Where promptfoo can drive two providers over the same saved cases inside a single `eval`, prefer that: one run, one suite, one grading pass removes the entire class of "was it really the same suite?" doubt.
- Why is saving the generated cases better than rerunning the same plugin and strategy configuration?The configuration only fixes what kind of attacks are produced, not which ones. Rerunning generation gives a fresh sample: different wordings, different mix, sometimes a different count. Replaying saved cases makes the two runs compare the same inputs.
- You changed the prompt and the temperature at the same time. What can you report?Only that the pair scored differently on this suite. To attribute it you rerun from a common baseline moving one at a time; temperature especially changes variance, so the same prompt can score differently at a different setting.
Regenerating the adversarial cases between the two runs is like weighing yourself on a different scale in a different room and calling the difference weight loss. The reading is real; the comparison is not.
saying these in an interview costs you the question
- Assuming two runs of the same promptfoo red-team configuration produce the same test cases.
- Treating the grader as a fixed oracle rather than a model-driven component that is part of the control.
- Reporting a percentage delta when the model, prompt and settings all moved, with the confounds only mentioned verbally.
- Comparing failure counts across runs with different case counts without normalising.