skip to content

You changed only the target model in a promptfoo red-team run — same saved adversarial cases, same plugins and strategies, same grader — and the failure rate moved several points. What could have changed besides the model, and how would you rule each out before reporting the model as the cause?

level: seniorimportance: should knowfreq 44%

answer

  1. seeds pinned, sent prompts not
  2. adaptive strategies branch on replies
  3. grader is style-sensitive
  4. cache replay on one side
  5. rerun baseline twice for spread

basics

~20 s

Multi-turn and adaptive strategies build follow-ups from the target's own replies, so the prompts actually sent differ per model even from identical seeds. The grader is a model and may score a new refusal style differently. Caching, truncation and errored cases also differ. Rule them out by diffing sent transcripts and re-running the baseline twice.

solid answer

~50 s

"Same suite" pins the seeds, not the conversation. Any strategy that adapts — multi-turn escalation, iterative rephrasing driven by an attacker model — generates its later turns conditioned on what the target just said. Point it at a different model and the actual text sent diverges after turn one, so the two runs did not present the same attack. Three more suspects. The grader is a model reading free-form output: a model that refuses in a different register can be scored differently for the same underlying behaviour, and if the grading model changed with the platform it changed silently. Caching can replay one side's responses while the other side is generated fresh. And errors, rate limits and truncated outputs are not neutral — check they occurred at similar rates and that both runs counted them the same way. To separate real effect from run-to-run spread, rerun the baseline configuration unchanged and see how much the number moves on its own before you attribute anything to the swap.

go deeper

for a junior

Expected to say the model output changed and maybe that responses vary run to run; unlikely to reach the adaptive-strategy point.

for a middle

Names the grader as a model and mentions caching and sampling variance, and knows to rerun the baseline for comparison.

for a senior

Explains that adaptive and multi-turn strategies branch on the target's replies so identical seeds do not mean identical attacks, and separates single-turn from adaptive categories in the report.

for a principal

Sets the standard for what may be claimed from a target swap at all, and requires transcripts, grader pinning and a repeated baseline before any safety comparison is published.

### The distinction that carries the whole answer: seeds versus sent prompts A saved promptfoo red-team suite (`redteam.yaml`, produced by promptfoo's `redteam generate` and replayed by `redteam eval`) fixes the **starting points** of each attack. For a single-turn plugin case, the starting point *is* the whole attack: one prompt goes to the target, one response comes back, the grader labels it. Two targets receiving that same string is a genuinely like-for-like comparison. Adaptive strategies break that equivalence by design. promptfoo's multi-turn strategies — crescendo-style gradual escalation, GOAT-style attacker-driven dialogue, and the iterative jailbreak strategy that has an attacker model rewrite its own prompt after each rejection — compute every turn after the first from *what the target just said*, up to a configured turn cap. Point the identical seed at a different model and the conversation forks at turn two. From then on the two targets faced different attacks, and neither the transcript nor the verdict is a property of the model alone. So there are two honest report shapes, and you should use both: single-turn plugin categories as a like-for-like model comparison, and adaptive categories as "this strategy converted at X against target A and Y against target B" — a statement about the strategy-and-target *pair*. ### The grader is the second suspect Grading is a model reading free-form output against a rubric. Two distinct failure modes follow. It is **style-sensitive**: a terse refusal and a long, hedged partial compliance can land on opposite sides of the same rubric even when the underlying behaviour is comparable, and different model families refuse in noticeably different registers — so a rubric can systematically favour one of the two targets. And it **drifts**: if the grading provider was left at a platform default rather than pinned by id (promptfoo's `defaultTest.options.provider`), a vendor-side update relabels transcripts with no change on your side at all. Mitigations, in order of value: pin the grading model explicitly; hand-label a stratified sample of both runs and measure agreement with the automated labels; keep those hand labels as a standing regression check on the grader itself. ### Mechanics that quietly differ ```text cache promptfoo keys cached responses by provider + prompt, so swapping the target invalidates the new side while the old side may replay months-old responses - one fresh sample versus one frozen one limits a lower max_tokens truncating long compliance into an apparent refusal errors 429s, timeouts and provider-side content-filter rejections, and whether each run counted them as passes, failures or exclusions turns an adaptive strategy stopping short because turns timed out, so target B simply got fewer attempts than target A ``` ### What it costs to do this properly The verification is not free. A repeated baseline is one extra full run — for an 800-case suite with multi-turn strategies, several thousand target calls plus a grader call per case, tens of minutes to a couple of hours of wall clock, and a bill that scales with whichever target is more expensive. Two repeats is better than one if you want any sense of the spread's shape. Transcript diffing is engineer time rather than API spend: budget an afternoon to skim the cases that changed verdict, more if the off-diagonal is large. That cost is the point — the alternative is publishing a safety comparison you cannot defend. ### Where the number misleads The wrong reading is *"the new model is four points worse on adversarial robustness."* Three specific ways that fails. The delta may sit entirely in adaptive categories, where the two models were never given the same attacks — you measured an interaction between the attacker strategy and the target, and a model that argues back more can *invite* a longer, more effective escalation while being no more compliant. The delta may be grader style preference rather than behaviour. And it may simply be inside the run-to-run spread: rerun the unchanged baseline and if the number moves as much on its own, the swap has explained nothing yet. ### The procedure Rerun the unchanged baseline twice back to back to establish the natural spread. Diff the recorded *sent* prompts between the two runs and find the turn at which they diverge. Split the delta by plugin category and by single-turn versus adaptive. Hand-read the newly-failing and newly-passing cases. If the movement lives entirely in adaptive categories and the transcripts fork at turn two, you have measured a strategy-times-target interaction, and the sentence you write says so.

  • How do you tell whether the divergence is real or run-to-run variance?
    Rerun the unchanged baseline twice on the same saved cases and measure the spread. If a repeat of the same configuration moves as much as the swap did, the swap explains nothing yet.
  • What do you report when the whole delta sits in multi-turn categories?
    That the adaptive strategies converted at different rates against the two targets — a strategy-times-target result. Quote the single-turn categories separately as the like-for-like part of the comparison.
  • Why is a shared grading model between two compared targets still a risk?
    It grades free-form text, so it can systematically favour one model's refusal or answer style, and any change to that grading model relabels historical transcripts. Pin it and spot-check its labels by hand.

Replaying saved seed prompts against a new model is like replaying the opening move of a chess game against a different opponent: the first move is identical and nothing after it is, so the final score says as much about the game that unfolded as about either player.

saying these in an interview costs you the question

  • Treating a saved suite as guaranteeing identical prompts even for multi-turn strategies.
  • Reporting a delta without ever rerunning the unchanged baseline to see the natural spread.
  • Letting the grading model float on a platform default while comparing targets.
  • Never inspecting the recorded transcripts of the cases that changed verdict.
  • Ignoring errored, rate-limited or truncated cases instead of checking how each run counted them.

context