skip to content

Red-Team Generation

Plugins choose which harms get attempted and strategies how hard they push, so the suite is generated from a description rather than written by hand. Interviewers probe what that generation left out.

on this pageshow

explore

questions

18

In promptfoo's red-team mode, what does adding a plugin to the red-team configuration actually change about the suite that gets generated?

level: juniorimportance: must knowfreq 68%

answer

  1. plugin = harm class
  2. generator plus its grader
  3. unselected means untested
  4. plugin list is the taxonomy
  5. record the enabled set

basics

~20 s

Each promptfoo red-team plugin stands for one harm class. Adding one makes the generator write adversarial test cases aimed at eliciting that harm, and it supplies the grading criteria used to judge the replies. A plugin you leave out generates nothing, so that harm is never attempted and never shows up in the report.

solid answer

~50 s

A plugin in promptfoo's red-team mode is a **harm class**, not an attack trick. Enabling one does two things: it makes the generator synthesise test cases whose goal is to make your application commit that specific harm, and it attaches the grading logic that decides whether a given response counts as a failure of that class. So the plugin list is really a **risk taxonomy** — a declaration of which harms you are testing for at all. The consequence that matters at review time is asymmetric: enabling a plugin can produce failures *or* passes, but leaving one off produces neither. There is no row, no denominator, no evidence. A report is only ever a statement about the classes you selected, and it says exactly nothing about the rest of the catalogue. That is also why the plugin list belongs in the artefact you hand to a reviewer, not just in a config file nobody reads.

go deeper

for a junior

Should say a plugin corresponds to a category of harm, that enabling it causes adversarial cases for that category to be generated, and that unselected categories are simply absent from the results.

for a middle

Adds that the plugin also determines how responses in that class are graded, and frames the plugin list as the risk taxonomy the report is scoped to.

for a senior

Stresses the asymmetry — enabling yields evidence either way, omitting yields none — and insists the enabled set be recorded with each result so pass rates stay comparable over time.

for a principal

Treats plugin selection as a governed decision mapped to the organisation's risk register, with per-application deviations recorded and reviewed rather than left to whoever wrote the config.

### Where the decision physically lives In a promptfoo red-team project the whole thing sits in one YAML block, `redteam:`, inside `promptfooconfig.yaml`. The key is `redteam.plugins`, and each entry is either a bare plugin id — `pii:direct`, `bola`, `excessive-agency`, `hallucination`, `policy` — or an object carrying its own settings, `{ id: pii:direct, numTests: 20, config: { ... } }`. Some entries are **collection aliases** rather than single plugins: `harmful` expands to the whole family of harm sub-classes, `pii` to the personal-data ones, and a framework alias such as `owasp:llm` expands to whatever plugin set your installed version maps that framework to. One line of YAML can therefore be twenty plugins, which matters for both cost and for knowing what you actually ran. ### What generation does with each entry Generation and evaluation are separate steps: `promptfoo redteam generate` writes a `redteam.yaml`, `promptfoo redteam eval` runs it against the target, and `promptfoo redteam run` does both in one go. For every enabled plugin the generator takes that plugin's own template, splices in `redteam.purpose` — your prose description of what the application is and is allowed to do — plus any per-plugin `config`, and asks the attack-generation provider for `numTests` cases (default 5 per plugin, overridable globally with `redteam.numTests` or per entry). That generation call goes to promptfoo's hosted generation service by default; `PROMPTFOO_DISABLE_REDTEAM_REMOTE_GENERATION=true` keeps it on your own provider at the price of weaker or unavailable generation for some plugins. What lands in `redteam.yaml` is ordinary promptfoo test cases, each tagged with `metadata.pluginId` naming the plugin that produced it, and each carrying an assertion of the form `promptfoo:redteam:<pluginId>`. ### The half people forget: the grader travels with the plugin That assertion **is** the plugin's grader, and it is specific to the harm. The `pii:direct` grader asks whether personal data about a third party was disclosed. The `excessive-agency` grader asks whether the application claimed to perform an action it has no tool for. The `bola` grader asks whether it returned an object belonging to a different user. These questions have nothing structural in common, which is why there is no single global "was that bad?" judge to fall back on. Enabling a plugin buys two things bound together — cases aimed at one harm class, and the rubric that decides whether a reply committed it. Delete the entry and both halves go; no generic safety check remains behind. ### What a plugin costs The arithmetic is worth carrying in your head. The `harmful` alias alone is north of twenty sub-plugins; at the default five cases each that is well over a hundred base cases before the strategy axis multiplies anything. Every case costs at least one call to the target application and one to the grading model, so a hundred cases is a couple of hundred model calls — typically single-digit dollars on a mid-tier hosted judge and single-digit minutes of wall clock at default concurrency. Money is almost never the constraint. The real budget is human: every failure a plugin produces has to be read by someone, and triage time scales with enabled classes far more painfully than the bill does. ### Where the report's number misleads Two specific misreadings. The first is the asymmetry: a plugin you never listed generates no cases, so the report has no card, no row and no zero for it. Silence renders identically to safety. The second is subtler and catches experienced people — `numTests` is a **request, not a guarantee**. Deduplication, a refusal from the generation provider, or a purpose too thin for that plugin to write against can all return fewer cases than you asked for. A plugin card showing "2 of 2 passed" is coloured exactly like one showing "50 of 50 passed", and the first is nearly worthless evidence. ### What I would check before believing a run - Count the generated cases per plugin, from `redteam.yaml` or the report's per-plugin counts, and compare against the `numTests` you configured. Any plugin far under quota needs re-generating, not reading. - Record the enabled plugin list — the resolved list, not the alias — beside the result, so the next reader does not have to reconstruct it from version control. - Diff that list against the previous run before comparing any two numbers.

  • If a promptfoo red-team run reports a pass rate, what is the denominator?
    The cases generated from the plugins you enabled, at the suite size you set. It is a set you chose, not the space of harms your application could commit.
  • Why does each promptfoo red-team plugin bring its own grading criteria instead of sharing one?
    Different harm classes fail in different ways. Disclosing another customer's record and answering outside the assistant's remit have nothing structural in common, so one universal rule would mislabel both.
  • Two promptfoo runs a month apart both show zero failures. What must you compare before calling that stable?
    The enabled plugin sets. If one was dropped in between, the second run is measuring less, and the identical headline hides a shrunken denominator.

saying these in an interview costs you the question

  • Describing a promptfoo plugin as an attack technique or an obfuscation trick, which confuses it with the other configuration axis.
  • Claiming a clean run means the application is safe, with no reference to which plugins were enabled.
  • Assuming plugins ship only generated prompts and that a single global grader judges every harm class.
  • Treating the enabled plugin set as a config detail not worth reporting alongside the result.

context

open as a page

A promptfoo red-team run was configured with several attack plugins but no attack strategies enabled, and every generated case passed. What does that result support, and what should you refuse to claim from it?

level: juniorimportance: must knowfreq 66%

basics

~20 s

It supports one narrow claim: the target refused those harms when asked plainly, in a single turn, in the generator's own phrasing. It says nothing about the same asks under an adversarial framing or across several turns, which is where real failures live. Report it as passing the plain baseline, never as safe.

open as a page

In promptfoo's red-team mode you write a free-text description of the application under test. What is that description used for, and why does a one-line version weaken the results?

level: juniorimportance: must knowfreq 70%

basics

~20 s

promptfoo's red-team mode reads that description twice: the generator writes attack cases from it, and the graders judge replies against it. A one-line description yields generic attacks that miss your app's real surfaces, and gives the graders no stated rule to fail an answer against, so weak replies get marked pass.

open as a page

In a promptfoo red-team configuration, what does the per-plugin case count (numTests) actually control, and what does leaving it at a small default mean for a failure your app commits only rarely?

level: juniorimportance: must knowfreq 62%

basics

~20 s

It sets how many adversarial cases promptfoo generates for each enabled plugin, so total volume scales with plugins times that count. A small default gives each harm class only a handful of tries, so a failure the app commits rarely can easily never fire, and the run then reports a clean pass.

open as a page

A promptfoo red-team run over an internal support assistant comes back with zero failures, and the configuration enabled three plugins. What can and cannot you conclude from that report?

level: middleimportance: must knowfreq 61%

basics

~20 s

You can conclude that on this run, for those three harm classes, at that suite size, nothing failed. You cannot conclude the assistant is safe. Harm classes with no plugin enabled generated no cases, so the report is silent about them. Silence is a missing denominator, not a pass.

open as a page

In a promptfoo red-team configuration, attack plugins and attack strategies are two separate lists. If you enable one more strategy, how does the generated test-case count change, and why does that matter for a suite that reruns on every release?

level: middleimportance: must knowfreq 72%

basics

~20 s

Strategies do not add cases, they re-express them. Each strategy takes the cases the plugins generated and emits a transformed variant of every one, so the suite scales with plugins times strategies rather than growing by a fixed amount. Every scheduled rerun re-pays that multiplied count in target calls, grader calls and wall-clock time.

open as a page

Where does the application description you write for a promptfoo red-team run go by default, and what does that mean for what you are allowed to put in it?

level: middleimportance: must knowfreq 60%

basics

~20 s

By default promptfoo generates red-team cases through its hosted generation service, so the description of your application leaves your machine. Treat it as text you would hand a vendor: no secrets, no customer records, no internal hostnames. There is a documented switch that turns remote generation off and generates locally instead.

open as a page

A promptfoo red-team run against your internal HR assistant reports almost no failures. The application description in the config reads, in full: 'an HR chatbot'. Is that result evidence the assistant is safe, and what do you do next?

level: seniorimportance: must knowfreq 55%

basics

~20 s

No. promptfoo's graders judge each reply against the description you supplied. If it never says what the assistant must refuse, who may ask, or what records it reaches, an off-policy answer breaks no declared rule and is scored a pass. Rewrite the description with roles, data and refusals, then re-run.

open as a page

How do you decide which promptfoo red-team plugins apply to a given application, and what goes wrong if you simply enable the whole catalogue?

level: middleimportance: should knowfreq 50%

basics

~20 s

Work from what the application can actually do: its tools, the data it can reach, who talks to it, and what it must never say. Enable the harm classes it could genuinely commit. Enabling everything costs generation and inference on irrelevant classes and buries real failures in noise nobody triages.

open as a page

In a promptfoo red-team report, the same underlying plugin case passes when it is asked plainly but fails under one wrapped attack strategy. How do you interpret that, and why is a single overall pass rate a poor way to report it?

level: middleimportance: should knowfreq 46%

basics

~20 s

It means the refusal is keyed to the surface form of the request rather than to its intent: change the framing and the same harm goes through. Pool it into one overall pass rate and that signal disappears, because the plain copies dilute the wrapped failures. Report per plugin and per strategy instead.

open as a page

A promptfoo red-team plugin generates 10 cases against your assistant, and the behaviour it probes appears in roughly 1 relevant response in 20. Roughly what chance does that run have of reporting zero hits, and how many cases would you need before a clean result means something?

level: middleimportance: should knowfreq 45%

basics

~20 s

Treating each case as an independent try, the chance of missing every time is 0.95 to the tenth, about 60 percent. So a clean run of ten mostly reflects the sample size. To be roughly 95 percent sure of seeing a one-in-twenty behaviour you need about sixty cases.

open as a page

In a promptfoo red-team run, one enabled harm-class plugin produces a large batch of failures that triage decides are not real problems for this application. Do you disable that plugin, and what does disabling it cost you?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Not before separating two causes: the application genuinely cannot commit that harm, or it can and the failures are being judged against the wrong expectation. Only the first justifies disabling. Disabling deletes the class from every future run too, so the next report is silently narrower while the headline number improves.

open as a page

You enable a multi-turn attack strategy in a promptfoo red team against an HTTP chat endpoint you wired up yourself. It reports almost no failures, while a single-shot strategy against the same endpoint finds several. What target-side configuration do you check first, and why?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Check that the target definition carries conversation state between turns. If every request hits the endpoint as a fresh conversation, a multi-turn strategy's gradual escalation is thrown away and each turn lands as an isolated plain ask, so it under-reports by construction rather than because the target is strong.

open as a page

A promptfoo red-team suite regenerates and re-runs nightly against a metered chat endpoint, and its bill has become the argument for deleting it. Someone proposes cutting the per-plugin case count to two across the board. What do you do instead?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Do not thin everything evenly - that leaves every harm class too small to conclude anything while still costing money. Split the schedule: a small nightly tier on the surfaces that change, and a deep sweep with high case counts before releases. Stop regenerating cases you could reuse, and label the thin tier as a smoke test.

open as a page

Your team runs a promptfoo red-team suite on every release, and every attack strategy you enable is re-paid in full on each run. How would you decide which strategies stay in the per-release run and which move to a periodic deeper run?

level: principalimportance: should knowfreq 34%

basics

~20 s

Split by what each strategy is for. The per-release run should be a cheap, stable regression arm: deterministic framings, fixed cases, fast enough to block a release. Expensive iterative and conversational framings go to a scheduled deeper run, where a slow, stochastic result is investigated by a person rather than gating a deploy.

open as a page

Several teams run promptfoo red-team suites on a schedule, each with its own written description of the application under test. How do you keep those descriptions from quietly rotting, and who should own them?

level: principalimportance: should knowfreq 35%

basics

~20 s

Treat each description as a reviewed, version-controlled artifact owned with the application. When an app gains a tool, role or data source and the text does not, promptfoo's generator stops attacking the new surface and its graders stop failing it, so the pass rate rises while real risk grows. Diff it every release.

open as a page

You own promptfoo red-team suites for a dozen applications under one fixed monthly spend. How do you decide the per-plugin case counts across that portfolio, and why is a single uniform number the wrong answer?

level: principalimportance: should knowfreq 30%

basics

~20 s

Treat it as allocating a fixed recurring budget, not picking a setting. Give high counts to the harm classes where a miss is expensive and to apps with real exposure, and thin counts elsewhere. A uniform number overspends on harms an app cannot commit while under-sampling the ones it can, and hides both behind one green score.

open as a page

You own adversarial-testing coverage for a dozen LLM applications, all scanned with promptfoo's red-team mode. How do you govern which harm-class plugins each application runs, and what breaks when the available catalogue changes between quarters?

level: principalimportance: nice to knowfreq 29%

basics

~20 s

Set a baseline of harm classes every application must run, plus per-application additions derived from its capabilities, with every exclusion carrying a written reason and an owner. Record the enabled set with each report. When the catalogue grows, new classes change the denominator, so cross-quarter comparisons are invalid unless you re-baseline deliberately.

open as a page