skip to content

Your team runs a promptfoo red-team suite on every release, and every attack strategy you enable is re-paid in full on each run. How would you decide which strategies stay in the per-release run and which move to a periodic deeper run?

level: principalimportance: should knowfreq 34%

answer

  1. gate arm vs deep arm
  2. fixed cases, no regeneration in the gate
  3. deterministic cheap framings gate, iterative ones schedule
  4. diff against baseline, not absolute threshold
  5. promote confirmed findings into pinned cases

basics

~20 s

Split by what each strategy is for. The per-release run should be a cheap, stable regression arm: deterministic framings, fixed cases, fast enough to block a release. Expensive iterative and conversational framings go to a scheduled deeper run, where a slow, stochastic result is investigated by a person rather than gating a deploy.

solid answer

~50 s

Two different jobs are being conflated. A per-release run answers *did we regress?* — it needs to be fast, reproducible and comparable to last week, so it should hold a **fixed** set of cases under **deterministic** framings, with the pass condition being no new failures rather than an absolute threshold. A deep run answers *what can we still be broken by?* — it wants breadth, expensive iterative and conversational framings, and the fresh case generation that makes results non-comparable but informative. Decide per strategy on three axes: cost per case (single rewrite versus many turns), reproducibility (does the same input produce the same trajectory), and historical yield (has this framing ever produced a triaged finding here). Cheap plus deterministic plus proven belongs in the gate. Expensive or stochastic belongs in the periodic run, plus any framing that once produced a real finding, pinned as a fixed regression case so the specific fix stays fixed.

go deeper

for a junior

Knows the full suite is too slow for every release and that some subset must run more often than the rest.

for a middle

Separates cheap deterministic framings from expensive iterative ones and puts the cheap ones in the frequent run.

for a senior

Designs the gate as a regression diff over fixed cases, promotes confirmed findings into pinned cases, and budgets grader calls and rate limits alongside target calls.

for a principal

Owns the two-arm design as a standing cost against release frequency, anticipates that a slow or flaky gate gets disabled, and fixes the reporting contract so configuration changes are never read as safety changes.

### Two different questions are being asked of one suite A per-release run answers *did we regress?* A deep run answers *what can we still be broken by?* Those want opposite properties. Regression detection wants a fixed, deterministic, fast suite whose result is comparable with last week's. Discovery wants breadth, fresh cases, and the expensive iterative framings whose results are informative precisely because they are not repeatable. Trying to serve both from one configuration produces a suite that is too slow to gate and too repetitive to discover anything. ### The arithmetic that forces the split Take a realistic shape: ten plugins at `numTests: 10` is 100 base cases. Enable four single-shot framings and two multi-turn framings with a five-turn cap. The single-shot side is 500 cases at roughly one target call and one grader call each; the multi-turn side is 200 cases at up to five target calls plus an attacker-model call per turn, plus grading. That is on the order of two to three thousand model calls per run — tens of minutes to hours of wall-clock depending on concurrency and the provider's rate limit, and real money. Now multiply by your release frequency. At twenty releases a month that suite is not a check, it is a standing service with a budget line. The failure modes are organisational, not technical, and both are predictable: **a slow gate gets bypassed, and a flaky gate gets ignored and then disabled.** Either outcome leaves less safety than a small honest gate plus a real periodic engagement. ### The gate arm Fixed, versioned cases — commit the generated test file and replay it; do not regenerate per run. Deterministic framings only, so the same input produces the same request. Sized to what the pipeline will actually tolerate, which is usually minutes. The pass condition is a **diff against a recorded baseline**: a case that used to pass and now fails blocks the release; the absolute rate is reported but does not gate. Every previously-confirmed finding is promoted into this arm as a pinned case with its framing and its expected outcome, which is how a fix stays fixed. ### The deep arm Scheduled, off the release path, sized by budget rather than by pipeline patience. Fresh `redteam generate` each time, the full strategy catalogue including the iterative and multi-turn framings, and a named owner whose job is triage rather than a red-or-green light. Its output is findings, and the durable ones are promoted into the gate arm as pinned cases. ### Ranking strategies between the arms Rank by findings-per-unit-cost **measured on your own target**, not by reputation. Cheap, deterministic, and historically productive belongs in the gate. Expensive or stochastic belongs in the deep run. Retire a framing from the gate when it has produced nothing across many runs and its cost is real — but keep it in the deep arm, and revisit the decision after any material change to the guard, the system prompt or the model, because a framing that stopped finding things may have stopped only because of a defence that has just been rewritten. ### Where these numbers mislead Three ways, all of which get quoted upward. First, an absolute pass-rate threshold computed over freshly generated cases is a coin flip: a fresh draw changes both which harms were sampled and how they were phrased, so run-to-run movement is sampling noise, not target change, and the gate blocks releases at random until someone disables it. Second, "we red-team every release" is heard as full coverage while the gate arm deliberately holds only the cheap deterministic slice — say plainly what the gate does and does not cover. Third, the gate's own rate *improves* when you retire an expensive strategy, which reads exactly like a safety improvement and is nothing of the kind; pin the configuration to every number. ### What to check Measure the gate arm's p95 wall-clock over a few weeks, not its best case, and compare it with what developers will actually wait for. Confirm the red team's provider quota is not shared with production traffic, or a scheduled deep run becomes an availability incident. Budget grader calls explicitly — they roughly double the naive target-call estimate. And verify your pinned regression cases actually work: a pinned case that has never once gone red has not been shown to detect anything, so exercise it against a build with the fix removed before trusting it to guard that fix.

  • Why should the gate arm not regenerate its cases each run?
    Because a fresh sample makes runs incomparable: the rate moves with which cases happened to be drawn, so a threshold on it flakes. Fixed cases make the pass condition a genuine regression diff.
  • A framing has found nothing in the gate for months. Drop it?
    Drop it from the gate if its cost is meaningful, but keep it in the periodic deep run, and revisit after any change to the guard, system prompt or model — silence may reflect a defence rather than an irrelevant framing.
  • How do you keep a fixed confirmed finding from silently regressing?
    Pin that exact case, with its framing, as a permanent gate case with a recorded expected outcome, so the fix is verified on every release rather than only at the next deep run.

saying these in an interview costs you the question

  • Puts the entire strategy catalogue in the release gate and accepts a very long pipeline.
  • Gates on an absolute pass-rate threshold computed from freshly generated cases each run.
  • Does not pin previously-confirmed findings as permanent regression cases.
  • Ignores grader calls and shared provider rate limits when sizing the recurring run.
  • Treats a strategy that stopped finding things as permanently irrelevant without revisiting after defence changes.

context