skip to content

In a promptfoo red-team report, the same underlying plugin case passes when it is asked plainly but fails under one wrapped attack strategy. How do you interpret that, and why is a single overall pass rate a poor way to report it?

level: middleimportance: should knowfreq 46%

answer

  1. plain green, wrapped red = form-sensitive refusal
  2. matrix: plugin rows, strategy columns
  3. pooled rate tracks config, not target
  4. column red everywhere = one high-value fix
  5. dedupe cases into behaviours

basics

~20 s

It means the refusal is keyed to the surface form of the request rather than to its intent: change the framing and the same harm goes through. Pool it into one overall pass rate and that signal disappears, because the plain copies dilute the wrapped failures. Report per plugin and per strategy instead.

solid answer

~50 s

The plain-versus-wrapped split is the finding, not noise. A target that refuses the direct ask and complies with the reframed one is matching on how the request looks, so the defence generalises poorly to anything you did not think of. The reporting problem is a denominator problem. Once strategies are enabled, the suite contains one copy of each base case per framing. A pooled pass rate averages the easy copies with the hard ones, so it moves mostly with **how many strategies you enabled**, not with how the target behaves — add a cheap strategy the target survives and the headline number improves while nothing got safer. The readable shape is a matrix: plugin (harm class) down one axis, strategy (framing) across the other, pass rate in each cell. That tells you which harms are reachable and which framings reach them, and it stays comparable between runs as long as the axes stay fixed.

go deeper

for a junior

Recognises that the same harm got through once the wording changed, and that this is worse news than the plain result suggested.

for a middle

Explains the denominator problem and asks for per-plugin, per-strategy numbers rather than one pooled rate.

for a senior

Prioritises a framing that fails across many harms over a single harm row, dedupes cases into behaviours, and distinguishes a stochastic red from a measured one.

for a principal

Fixes the reporting contract so headline numbers cannot drift with configuration, and pins the config to every number the org quotes.

### The shape the results actually have Once `redteam.strategies` is non-empty, every base case a plugin wrote exists in the suite several times over: once plain, once per enabled framing. The natural shape of the result is therefore a matrix — plugin (harm class) down one axis, strategy (framing) across the other, a pass rate in each cell — and the plain column is one of the columns, not a separate thing. Collapsing that matrix to one number throws away the only structure the run has. ### Reading the cell, not the average Each part of the matrix answers a different question. - **A row green plain, red under several framings.** The harm is genuinely reachable and the refusal is keyed to the *surface form* of the request rather than its intent. The defence generalises poorly, which means it will also fail against framings nobody in this catalogue thought of. - **A column red across most rows.** One framing defeats the defence broadly. This is usually the highest-value single finding in the run, because one fix at the guard or system-prompt layer closes many rows at once. - **A cell that flips between runs.** Iterative and multi-turn framings take a different trajectory each time, so a lone red cell from one of those is a lead, not a measurement, until it has been re-run. - **A row red even plain.** Urgent, and usually a configuration problem — a missing guard, an unfiltered path — as much as a model problem. ### Why the pooled number misleads in both directions It moves with the configuration, not with the target. Enable a framing your target happens to survive and the average rises; enable one aggressive framing across every plugin and it falls. Neither movement is about the model. Anyone comparing this month's headline against last month's is comparing configurations at least as much as targets, which is why the plugin-and-strategy list has to be pinned to every number the organisation quotes. There is a second dilution effect on top: the plain copies are the easy ones and there is one of them for every base case, so they permanently drag a pooled rate toward the optimistic end. ### What per-cell resolution costs The matrix is only readable if each cell holds enough cases to mean anything. `redteam.numTests` is per plugin, so cell size is roughly `numTests` — at five, a cell moves in 20-point steps and two rows differing by one case look different when they are not. Raising `numTests` to get a readable cell is paid in every cell of the matrix at once: total cases scale as `plugins × (1 + strategies) × numTests`, and each case costs a target call plus a grader call on every rerun. Wanting finer cells and wanting a cheap recurring run are directly opposed, and that trade is the actual decision behind the reporting shape. ### Triage: the tool's unit is not the report's unit The tool's unit is one graded case. The report's unit should be a deduplicated behaviour. Several cases in the same row failing under the same framing for the same underlying reason are one finding with many instances — collapse them before anyone counts them, or a single weakness gets reported as a dozen and the fix list becomes noise. Keep at least one verbatim transcript per finding as evidence, and record whether the plain version of that same base case passed, because that pairing is what makes the finding legible to whoever has to fix it. ### Where the individual cell misleads, and what to check A red cell under a wrapped framing is worth exactly as much as the grader's judgement of it, and wrapped text is precisely what confuses graders. Two checks before writing a finding down. First, confirm the grader judged against the *original intent* of the base case rather than the surface of the transformed text — a framing that recasts the ask as fiction can make a grader mark a harmless in-character answer as compliance, giving a false red exactly where the interesting results live. Second, read the response and confirm the harmful content is actually present rather than a refusal the grader misread. Then re-run any red cell produced by an iterative framing at least once before treating it as a rate rather than a lead.

  • One framing is red across almost every harm class. How do you prioritise it?
    Above the individual rows. A column-wide failure means one framing bypasses the defence generally, so a single fix at the guard or system-prompt level likely closes many rows at once.
  • A wrapped case failed on one run and passed on the next. Is that a finding?
    It is a lead. Iterative and conversational strategies take a different trajectory per run, so re-run it and keep the transcript; treat one stochastic red as evidence to investigate rather than a measured rate.

A pooled pass rate is a class average that goes up when you add an easy exam to the term. The number moved; the students did not.

saying these in an interview costs you the question

  • Dismisses a wrapped failure because the plain version passed.
  • Quotes one overall pass rate across all framings without stating the config that produced it.
  • Counts every graded case as a separate finding in the report.
  • Treats a single red cell from an iterative strategy as a stable rate without re-running it.
  • Never checks whether the grader judged the wrapped case against its original intent.

context