skip to content

Your product's promptfoo red-team results are unacceptable and you can move three things: the backing model, the application's system prompt, or add a guardrail in front. How do you sequence the experiments over one frozen generated suite, and what do you refuse to change while the campaign runs?

level: principalimportance: should knowfreq 28%

answer

  1. order by test cost and revert cost
  2. prompt, then guard, then model
  3. freeze suite, grader, counting convention
  4. held-out slice against overfitting
  5. guard needs a benign-traffic run too

basics

~20 s

Sequence by cost to test and cost to reverse: prompt first, guardrail second, model last, one variable per comparison over the same saved suite. Freeze the suite, the grader and the counting convention for the whole campaign, and hold back cases the prompt is never tuned against so improvement is not just fitting the suite.

solid answer

~60 s

Three axes decide the order: cost to run the experiment, cost to ship the result, and how much of the measurement surface it disturbs. **Prompt first** — cheapest to iterate and to revert, and its effect is directly visible in the same measurement. **Guardrail second** — it moves the surface (blocked requests never reach the model) and adds a false-positive cost on benign traffic that this suite cannot measure, so it needs a second experiment on ordinary requests before it can be judged. **Model last** — the most expensive to test and to ship, and it re-baselines everything downstream, including any prompt you tuned against the previous one. What is frozen for the whole campaign: the generated suite, the grader and its rubric, the target decoding settings, and the convention for counting blocks and errors. Changing any of those mid-campaign makes every earlier run incomparable and quietly discards the work already done. The real principal-level risk is overfitting: iterate a prompt against one fixed suite for long enough and the score improves on those cases specifically. Keep a held-out slice you never tune against, and confirm the winner there.

go deeper

for a junior

Should say to change one thing at a time and try the cheapest change first, over the same test cases.

for a middle

Justifies the ordering by cost and reversibility and names what must stay frozen across the campaign — suite, grader, target settings, counting convention.

for a senior

Adds the guardrail's separate false-positive experiment, the interaction between prompt and guard, and per-category reporting rather than a single rate.

for a principal

Owns the campaign as an experiment programme: discovery versus confirmation, a held-out slice against overfitting a frozen suite, a planned refresh cadence with re-baselining, and an explicit statement of what the suite never tested.

### The tension, and how to resolve it Attribution wants one variable at a time; the deadline wants the biggest fix first. Resolve it by splitting the campaign into **discovery** and **confirmation**. Discovery is cheap and permissive: try several system prompts, look at movement per plugin category, form a hypothesis, change several things at once if it helps you think. Confirmation is strict: one clean single-variable comparison from a stated baseline over the frozen suite. Only the confirmation goes in the report. Conflating the two is how organisations end up with a number nobody can reproduce. ### Ordering, and the costs that drive it ```text change test cost ship / revert cost surface disturbed prompt low low (edit and redeploy) none guardrail medium medium (false-positive blocks before the model risk on real users) model high high (re-baselines every- everything downstream thing tuned before it) ``` **Prompt first.** Cheapest to iterate, cheapest to revert, and its effect is directly visible in the same measurement — no new component, no new counting convention. Each iteration costs one full replay of the suite: for an 800-case suite with multi-turn strategies, a few thousand target calls plus a grader call per case, tens of minutes of wall clock, single-digit dollars against a commodity model and considerably more against a frontier one. Ten prompt iterations is a real bill; budget it before you start, because it is the line item that quietly grows. **Guardrail second.** It moves the measurement surface — blocked requests never reach the model, so those cases now score a classifier rather than your application — and it carries a cost this suite structurally cannot see. Budget a *second* experiment on benign traffic (a few thousand real or realistic ordinary requests, measured for wrongly-blocked rate) before the guardrail's promptfoo delta means anything about production. **Model last.** Most expensive to test, most expensive to ship, and it invalidates work: any prompt you tuned against the old model has to be re-confirmed against the new one, because prompt effects do not transfer cleanly across model families. ### What is frozen for the whole campaign The generated suite file, the grading model and its rubric (pinned explicitly, not inherited from a default that a vendor can update), the target decoding settings, and the convention for counting blocked and errored cases. Change any of those mid-campaign and every earlier run silently becomes incomparable — you have not just lost the next comparison, you have discarded the work already paid for. ### The failure mode that defines this level: overfitting the frozen suite A frozen suite is the price of attribution and the source of its worst failure. Iterate a prompt against the same cases long enough and you are tuning to those cases specifically: the score climbs while nothing generalises. This is not hypothetical — it is the expected outcome of any optimisation loop pointed at a fixed objective. Countermeasures, all three needed. Hold out a stratified slice at the start (a fixed fraction per plugin category) that you never look at during tuning, and confirm the winning configuration there — the gap between tuning-slice gain and held-out gain *is* the overfitting measurement. Refresh the tuning slice on an announced cadence, accepting that a refresh is a new denominator: the trend line breaks and every configuration you still care about must be re-measured on the new suite, which is exactly why refreshes are planned and budgeted rather than incidental. And read the wins: if they concentrate in the exact phrasings the prompt now names, you have memorised the exam, not learned the subject. ### What the suite never tells you An attack-only, generator-authored suite is silent on false positives, latency, cost per request, behaviours no configured plugin produces, and any attack a motivated human would invent that the attacker model did not. A guardrail can top the campaign and still be unshippable; a prompt can win here and lose on task quality, which this suite does not measure either. ### The deliverable A per-variable, per-category effect table on a named suite version, with a named grader, a stated counting convention, the held-out confirmation, the interaction between prompt and guard measured rather than assumed, and an explicit list of what was not tested. That is a decision record, and it survives the question "how do you know?". A single improved percentage does not.

  • When is it right to move the model first anyway?
    When the failures are capability-level rather than instruction-level — the current model cannot follow the policy at all — or when the model swap is already committed for other reasons, in which case you re-baseline first and tune the prompt afterwards.
  • How do you know the prompt is being overfitted to the suite?
    The score improves across iterations on the tuning slice while the held-out slice stays flat, and the wins concentrate in the exact wordings the prompt now mentions. That gap is the overfitting measurement.
  • Why must a suite refresh be planned rather than opportunistic?
    A new suite is a new denominator: the trend line breaks and every configuration you care about has to be re-measured on it. Planned refreshes budget that re-baselining; incidental ones silently invalidate comparisons.

A frozen suite you tune against every day stops being a thermometer and becomes an exam whose paper you have already seen. The held-out slice is the second paper you never opened, and it is the only one whose score still means anything.

saying these in an interview costs you the question

  • Iterating a prompt against one frozen suite for weeks with no held-out slice and calling the improvement robustness.
  • Swapping the grader or regenerating the suite mid-campaign and continuing to compare against earlier runs.
  • Judging a guardrail purely on the attack suite, with no benign-traffic false-positive measurement.
  • Changing the model first because it feels decisive, then tuning a prompt whose earlier results no longer apply.
  • Reporting one improved percentage with no statement of what the suite did not cover.

context