skip to content

A release swapped the backing model, rewrote the system prompt and added an input guardrail, and the promptfoo red-team failure rate for the app dropped afterwards. Why is "we got safer" not a supportable conclusion, and how would you rerun to attribute the change?

level: middleimportance: must knowfreq 68%

answer

  1. three causes, one number
  2. one-at-a-time ladder from a baseline
  3. changes can cancel each other
  4. interaction term is not zero
  5. guard: blocks, and false positives unmeasured

basics

~20 s

Three variables moved, so the drop has three possible causes and any of them could be masking a regression in another. Rerun the same saved adversarial suite from one baseline, adding one change at a time, with the grader and target settings fixed. Only then does a delta name a cause.

solid answer

~60 s

The number is real; the attribution is not. A single aggregate over a suite cannot tell you whether the guardrail blocked the attacks, the new prompt refused them, or the new model was simply less compliant — and the three can cancel: a guardrail that stops half the jailbreak strategies can hide a model that regressed on data leakage. The rerun is a small matrix over **one frozen suite**: baseline, baseline+model, baseline+prompt, baseline+guardrail, then the shipped combination. Grader, plugins, strategies and target decoding settings stay pinned throughout; the only thing that moves is the labelled variable. Two caveats worth saying out loud. The single-variable deltas will not add up to the combined delta, because prompt and guardrail interact — a guardrail that removes a strategy's phrasing changes what the prompt ever sees. And a guardrail is measured differently from the other two: it produces blocks, which you must decide to count as passes, and it carries a false-positive cost on real traffic that this suite does not measure at all.

go deeper

for a junior

Should recognise that three simultaneous changes mean the number cannot be attributed, and that the fix is to change one thing at a time over the same cases.

for a middle

Lays out the baseline-plus-one-change ladder, keeps the suite and grader pinned, and knows the guardrail has to be measured against a stated convention for blocked requests.

for a senior

Adds interaction between mitigations, stratified subsetting when the budget is short, per-plugin-category breakdowns instead of a single rate, and the false-positive check the attack-only suite cannot give.

for a principal

Decides what the organisation is allowed to ship without attribution, and who owns the frozen suite and grading convention that make the ladder meaningful across teams.

### Why the confound is structural, not sloppiness A promptfoo red-team run reports one failure rate over a deliberately heterogeneous suite: the cases come from many `redteam.plugins` entries (PII leakage, harmful content, scope overreach, prompt injection, and so on), each wrapped by the `redteam.strategies` you configured. Different fixes move different parts of that mixture. A backing-model swap shifts refusal behaviour broadly and unevenly. A system-prompt rewrite mostly moves the categories the new text explicitly mentions. An input guardrail moves whatever its classifier recognises in the *incoming* text, and nothing else. Collapsing all of that into one percentage destroys precisely the structure you would need to attribute the move — and it does so before you ever get to look. Worse, the three can cancel. A guardrail that intercepts half the jailbreak-strategy phrasings can comfortably hide a new model that regressed on data leakage: the categories move in opposite directions and the aggregate looks like a clean win. "We got safer" is then not merely unproven, it may be false in the dimension you care most about. ### The ladder From one common baseline, over one identical saved suite (`redteam.yaml` generated once, replayed by promptfoo's `redteam eval`): ```text run 0 old model, old prompt, no guard <- baseline run 1 new model, old prompt, no guard <- model effect run 2 old model, new prompt, no guard <- prompt effect run 3 old model, old prompt, guard on <- guard effect run 4 new model, new prompt, guard on <- the shipped combination ``` Run 4 minus the sum of runs 1-3 is the interaction term, and it is almost never zero: two mitigations that catch the same attacks overlap rather than stack, so the shipped combination usually improves less than the parts suggest. Report the combination as measured, not as a sum. Held constant across all five: the saved cases, the grading model and rubric (pin it via promptfoo's `defaultTest.options.provider` rather than inheriting a default), the target's decoding settings, and the counting convention for blocks and errors. ### What the ladder costs, and what to cut when you cannot afford it Five runs over an 800-case suite, at one target call plus one grader call per single-turn case and five to ten target calls per multi-turn case, is on the order of ten to twenty thousand model calls — hours of wall clock and a real, if modest, API bill; a frontier target model turns "modest" into "needs approval". Add engineer time: standing up four extra configurations, wiring the guardrail so it can be toggled without touching anything else, and reading transcripts afterwards. When the budget is short, **cut cases, not configurations.** Take a stratified subset — a fixed number of cases per plugin category, so the mixture is preserved — and run the whole ladder on it. Fewer cases only widens the uncertainty band on each delta, and you can say by how much. Fewer runs destroys attribution outright, and no amount of later analysis recovers it. ### Where the guardrail's number specifically misleads Position changes the measurement. With the guard in front, a blocked attack never reaches the model, so that case is now scoring the *classifier's recall on generated attacks* rather than your application's behaviour. Two consequences follow. First, the counting convention is load-bearing: decide and write down whether a block counts as a pass, because that choice alone can move the headline several points without any code changing. Second — and this is the trap that ships bad guardrails — an attack-only suite contains no benign traffic, so it can never show the guard's false-positive cost. A classifier that blocks every request scores a perfect zero-failure run on promptfoo and is unshippable. The guardrail therefore needs a *second* experiment on ordinary requests before its promptfoo delta means anything about production. A third, smaller trap: the generated attacks were authored by promptfoo's attacker model against your `redteam.purpose`, not by anyone probing your specific guard. A guard tuned on the same public phrasings the generator favours will look better here than against a motivated human. ### What you actually report Per-variable deltas broken down by plugin category, the measured interaction term, the suite identifier every run in the ladder used, the pinned grader configuration, the block/error counting convention, and one line on what the suite did not test (benign traffic, latency, cost, any behaviour no configured plugin generates). "Safer" without those is an opinion with a decimal point attached.

  • You cannot afford five full runs. What do you cut?
    Cut cases, not configurations. Take a stratified subset — a fixed number per plugin category — and run the whole ladder on it. Fewer runs destroys attribution; fewer cases only widens the uncertainty on each delta.
  • The single-variable deltas add up to more improvement than the combined run shows. Is one of them wrong?
    Not necessarily — that is interaction. Two mitigations often catch the same attacks, so their effects overlap rather than stack. Report the combination measured directly instead of the sum.
  • Why is a guardrail delta not comparable to a prompt delta?
    It changes the measurement surface: blocked requests never reach the model, so you are scoring a classifier's recall on generated attacks, plus a counting convention for blocks. And its real cost, false positives on benign traffic, is invisible to an attack-only suite.

saying these in an interview costs you the question

  • Claiming the guardrail is responsible because it was the most visible change.
  • Assuming the single-variable effects add up to the combined effect.
  • Dropping runs rather than cases when the query budget is tight, and still reporting attribution.
  • Not stating whether a blocked request counted as a pass.
  • Treating an attack-only suite as evidence the guardrail is safe to enable for all traffic.

context