skip to content

Attack Strategies

Strategies wrap each plugin's cases in a harder framing, so enabling one is a multiplier on case count rather than an addition. Interviewers ask because plain single-turn asks measure the easy half.

on this pageshow

explore

questions

5

A promptfoo red-team run was configured with several attack plugins but no attack strategies enabled, and every generated case passed. What does that result support, and what should you refuse to claim from it?

level: juniorimportance: must knowfreq 66%

answer

  1. no strategies = plain single-turn asks
  2. easy half of the attack surface
  3. baseline, not safety
  4. green sweep may mean nothing reached the model
  5. delta between plain and wrapped is the finding

basics

~20 s

It supports one narrow claim: the target refused those harms when asked plainly, in a single turn, in the generator's own phrasing. It says nothing about the same asks under an adversarial framing or across several turns, which is where real failures live. Report it as passing the plain baseline, never as safe.

solid answer

~50 s

With no strategies enabled, promptfoo sends each plugin's case as written: a direct, single-turn request for the harmful thing. That is the easiest form of every attack, and it is the form most tuned targets already refuse. A green run therefore establishes a **floor**, not safety. What you can say: the target does not comply with a plain request in these harm categories, at this sample size, for the purpose text you supplied. What you cannot say: that it holds under framing, encoding, role-play, gradual escalation, or a multi-turn conversation — none of that was generated, so none of it has a denominator. The right next move is to re-run with strategies enabled and compare against this baseline; the gap between the two is the actual finding. Also note the score depends on the purpose text and the plugin selection, so a green baseline with a thin purpose is weaker still.

go deeper

for a junior

Says that with no strategies the cases are plain asks, so a clean result is a baseline and not proof of safety.

for a middle

Names the three things the pass rate is a denominator over — harm classes enabled, cases per plugin, and the single plain framing.

for a senior

Also sanity-checks the plumbing before believing a clean sweep, and frames the plain-versus-wrapped delta as the reportable result.

for a principal

Sets the org's reporting language so a baseline run can never be quoted as a safety sign-off, and decides what evidence a release actually requires.

### What a strategy-free run actually sends When `redteam.strategies` is empty or absent, promptfoo evaluates exactly the base set that `redteam.plugins` produced and nothing else. Each case is a direct, single-turn request for the harmful thing, phrased by the attacker model in ordinary language and delivered to the target cold — no wrapper, no role-play, no prior turn, no encoding. That is the simplest possible form of every attack in the catalogue you selected. ### Why that form passes so easily Safety tuning is trained largely on directly-phrased harmful requests, so the plain surface is the exact distribution a tuned model has seen most of. A base case reads like the textbook example of the thing it should refuse, and any target with real alignment work behind it turns those down. The practical consequence is that a strategy-free suite behaves much closer to a smoke test than to an assessment: it confirms the target is not catastrophically open, and it confirms your wiring works end to end — the provider is reachable, responses are being extracted, the grader is running and returning verdicts. Those are genuinely worth confirming. None of them is a safety result. ### What the run costs — and why the cheapness is part of the trap This is the cheapest arm you will ever run: one target call plus one grader call per base case, no attacker loop, no extra copies. Eight plugins at ten tests each is 80 cases, roughly 160 model calls, minutes of wall-clock and small change in spend. That is exactly why it gets run first, run often, and quoted — and why a claim built on it costs the organisation far more than the run did. The strategies you did not enable are where the money and the time would have gone, and their absence is invisible in the output. ### What the pass rate is a denominator over Four things, all of which you chose: 1. **The harm classes you enabled** as plugins. A harm you did not select has no denominator at all, not a passing score. 2. **`numTests` per plugin** — how many cases were drawn for each class. 3. **One framing**: plain, single turn. 4. **The `redteam.purpose` text**, which is what made the cases specific to your app. Thin purpose text yields generic cases the target refuses easily, so a weak description inflates the result. The sample size deserves a second look. Ten cases per plugin, all passing, is not evidence of a low failure rate — a behaviour that genuinely fails one time in twenty will show zero failures in ten draws most of the time. Zero-out-of-ten is consistent with a true failure rate of up to roughly a quarter. "No failures" and "a low rate" are different claims, and only the first one was measured. ### How a green sweep can be entirely fake Before believing a clean result, rule out the failure modes that produce one for free. Response extraction that returns the whole JSON envelope rather than the assistant's text makes the grader score a wrapper, not an answer. A provider misconfiguration, a rate limit or a timeout turns calls into errors or empty strings, which a grader will happily score as non-compliance — errors read as refusals, and the report goes green. An authentication failure at the endpoint does the same thing on every case at once. All of these produce a perfect score from a suite that never reached the model. ### What to check Open a handful of raw outputs and read them. A real pass looks like the model's own voice declining, varied case to case. Identical strings on every case, empty bodies, or an error object mean the harness, not the target, produced your result. Confirm the non-2xx count is zero. Reproduce one case by hand against the same endpoint. Then enable one cheap deterministic strategy and rerun: if literally nothing moves across framings, be suspicious rather than pleased. ### How to report it One sentence, with the qualifiers attached: the target refused every plain, single-turn request in the harm classes we enabled, at this sample size, under the purpose text we supplied — and we have not yet tested adversarial framing or multi-turn escalation, so this is a baseline, not a safety verdict. The reportable finding arrives later, as the gap between this baseline and the same cases wrapped.

  • Your stakeholder asks for one sentence. What is it?
    The target refused every plain, single-turn request in the harm classes we enabled; we have not yet tested it under adversarial framing or multi-turn escalation, so this is a baseline rather than a safety verdict.
  • How would you tell a genuine clean pass from a run that never reached the model?
    Inspect a sample of raw responses. Non-empty model text that visibly refuses is a real pass; empty strings, error envelopes, or an identical wrapper on every case means the provider or response extraction is broken and the grades are meaningless.

saying these in an interview costs you the question

  • Reports a strategy-free green run as evidence the model is safe.
  • Does not know that omitting strategies leaves every case as a plain single-turn ask.
  • Never inspects raw responses to confirm the run reached the target.
  • Treats the pass rate as an absolute property of the model rather than a number over the harms and framings that were generated.

context

open as a page

In a promptfoo red-team configuration, attack plugins and attack strategies are two separate lists. If you enable one more strategy, how does the generated test-case count change, and why does that matter for a suite that reruns on every release?

level: middleimportance: must knowfreq 72%

basics

~20 s

Strategies do not add cases, they re-express them. Each strategy takes the cases the plugins generated and emits a transformed variant of every one, so the suite scales with plugins times strategies rather than growing by a fixed amount. Every scheduled rerun re-pays that multiplied count in target calls, grader calls and wall-clock time.

open as a page

In a promptfoo red-team report, the same underlying plugin case passes when it is asked plainly but fails under one wrapped attack strategy. How do you interpret that, and why is a single overall pass rate a poor way to report it?

level: middleimportance: should knowfreq 46%

basics

~20 s

It means the refusal is keyed to the surface form of the request rather than to its intent: change the framing and the same harm goes through. Pool it into one overall pass rate and that signal disappears, because the plain copies dilute the wrapped failures. Report per plugin and per strategy instead.

open as a page

You enable a multi-turn attack strategy in a promptfoo red team against an HTTP chat endpoint you wired up yourself. It reports almost no failures, while a single-shot strategy against the same endpoint finds several. What target-side configuration do you check first, and why?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Check that the target definition carries conversation state between turns. If every request hits the endpoint as a fresh conversation, a multi-turn strategy's gradual escalation is thrown away and each turn lands as an isolated plain ask, so it under-reports by construction rather than because the target is strong.

open as a page

Your team runs a promptfoo red-team suite on every release, and every attack strategy you enable is re-paid in full on each run. How would you decide which strategies stay in the per-release run and which move to a periodic deeper run?

level: principalimportance: should knowfreq 34%

basics

~20 s

Split by what each strategy is for. The per-release run should be a cheap, stable regression arm: deterministic framings, fixed cases, fast enough to block a release. Expensive iterative and conversational framings go to a scheduled deeper run, where a slow, stochastic result is investigated by a person rather than gating a deploy.

open as a page