skip to content

Attack Plugins

Each enabled plugin is a harm class the generator will try to elicit, and each one left off is a harm nobody tested. Interviewers ask because a clean report over three plugins says almost nothing.

on this pageshow

explore

questions

5

In promptfoo's red-team mode, what does adding a plugin to the red-team configuration actually change about the suite that gets generated?

level: juniorimportance: must knowfreq 68%

answer

  1. plugin = harm class
  2. generator plus its grader
  3. unselected means untested
  4. plugin list is the taxonomy
  5. record the enabled set

basics

~20 s

Each promptfoo red-team plugin stands for one harm class. Adding one makes the generator write adversarial test cases aimed at eliciting that harm, and it supplies the grading criteria used to judge the replies. A plugin you leave out generates nothing, so that harm is never attempted and never shows up in the report.

solid answer

~50 s

A plugin in promptfoo's red-team mode is a **harm class**, not an attack trick. Enabling one does two things: it makes the generator synthesise test cases whose goal is to make your application commit that specific harm, and it attaches the grading logic that decides whether a given response counts as a failure of that class. So the plugin list is really a **risk taxonomy** — a declaration of which harms you are testing for at all. The consequence that matters at review time is asymmetric: enabling a plugin can produce failures *or* passes, but leaving one off produces neither. There is no row, no denominator, no evidence. A report is only ever a statement about the classes you selected, and it says exactly nothing about the rest of the catalogue. That is also why the plugin list belongs in the artefact you hand to a reviewer, not just in a config file nobody reads.

go deeper

for a junior

Should say a plugin corresponds to a category of harm, that enabling it causes adversarial cases for that category to be generated, and that unselected categories are simply absent from the results.

for a middle

Adds that the plugin also determines how responses in that class are graded, and frames the plugin list as the risk taxonomy the report is scoped to.

for a senior

Stresses the asymmetry — enabling yields evidence either way, omitting yields none — and insists the enabled set be recorded with each result so pass rates stay comparable over time.

for a principal

Treats plugin selection as a governed decision mapped to the organisation's risk register, with per-application deviations recorded and reviewed rather than left to whoever wrote the config.

### Where the decision physically lives In a promptfoo red-team project the whole thing sits in one YAML block, `redteam:`, inside `promptfooconfig.yaml`. The key is `redteam.plugins`, and each entry is either a bare plugin id — `pii:direct`, `bola`, `excessive-agency`, `hallucination`, `policy` — or an object carrying its own settings, `{ id: pii:direct, numTests: 20, config: { ... } }`. Some entries are **collection aliases** rather than single plugins: `harmful` expands to the whole family of harm sub-classes, `pii` to the personal-data ones, and a framework alias such as `owasp:llm` expands to whatever plugin set your installed version maps that framework to. One line of YAML can therefore be twenty plugins, which matters for both cost and for knowing what you actually ran. ### What generation does with each entry Generation and evaluation are separate steps: `promptfoo redteam generate` writes a `redteam.yaml`, `promptfoo redteam eval` runs it against the target, and `promptfoo redteam run` does both in one go. For every enabled plugin the generator takes that plugin's own template, splices in `redteam.purpose` — your prose description of what the application is and is allowed to do — plus any per-plugin `config`, and asks the attack-generation provider for `numTests` cases (default 5 per plugin, overridable globally with `redteam.numTests` or per entry). That generation call goes to promptfoo's hosted generation service by default; `PROMPTFOO_DISABLE_REDTEAM_REMOTE_GENERATION=true` keeps it on your own provider at the price of weaker or unavailable generation for some plugins. What lands in `redteam.yaml` is ordinary promptfoo test cases, each tagged with `metadata.pluginId` naming the plugin that produced it, and each carrying an assertion of the form `promptfoo:redteam:<pluginId>`. ### The half people forget: the grader travels with the plugin That assertion **is** the plugin's grader, and it is specific to the harm. The `pii:direct` grader asks whether personal data about a third party was disclosed. The `excessive-agency` grader asks whether the application claimed to perform an action it has no tool for. The `bola` grader asks whether it returned an object belonging to a different user. These questions have nothing structural in common, which is why there is no single global "was that bad?" judge to fall back on. Enabling a plugin buys two things bound together — cases aimed at one harm class, and the rubric that decides whether a reply committed it. Delete the entry and both halves go; no generic safety check remains behind. ### What a plugin costs The arithmetic is worth carrying in your head. The `harmful` alias alone is north of twenty sub-plugins; at the default five cases each that is well over a hundred base cases before the strategy axis multiplies anything. Every case costs at least one call to the target application and one to the grading model, so a hundred cases is a couple of hundred model calls — typically single-digit dollars on a mid-tier hosted judge and single-digit minutes of wall clock at default concurrency. Money is almost never the constraint. The real budget is human: every failure a plugin produces has to be read by someone, and triage time scales with enabled classes far more painfully than the bill does. ### Where the report's number misleads Two specific misreadings. The first is the asymmetry: a plugin you never listed generates no cases, so the report has no card, no row and no zero for it. Silence renders identically to safety. The second is subtler and catches experienced people — `numTests` is a **request, not a guarantee**. Deduplication, a refusal from the generation provider, or a purpose too thin for that plugin to write against can all return fewer cases than you asked for. A plugin card showing "2 of 2 passed" is coloured exactly like one showing "50 of 50 passed", and the first is nearly worthless evidence. ### What I would check before believing a run - Count the generated cases per plugin, from `redteam.yaml` or the report's per-plugin counts, and compare against the `numTests` you configured. Any plugin far under quota needs re-generating, not reading. - Record the enabled plugin list — the resolved list, not the alias — beside the result, so the next reader does not have to reconstruct it from version control. - Diff that list against the previous run before comparing any two numbers.

  • If a promptfoo red-team run reports a pass rate, what is the denominator?
    The cases generated from the plugins you enabled, at the suite size you set. It is a set you chose, not the space of harms your application could commit.
  • Why does each promptfoo red-team plugin bring its own grading criteria instead of sharing one?
    Different harm classes fail in different ways. Disclosing another customer's record and answering outside the assistant's remit have nothing structural in common, so one universal rule would mislabel both.
  • Two promptfoo runs a month apart both show zero failures. What must you compare before calling that stable?
    The enabled plugin sets. If one was dropped in between, the second run is measuring less, and the identical headline hides a shrunken denominator.

saying these in an interview costs you the question

  • Describing a promptfoo plugin as an attack technique or an obfuscation trick, which confuses it with the other configuration axis.
  • Claiming a clean run means the application is safe, with no reference to which plugins were enabled.
  • Assuming plugins ship only generated prompts and that a single global grader judges every harm class.
  • Treating the enabled plugin set as a config detail not worth reporting alongside the result.

context

open as a page

A promptfoo red-team run over an internal support assistant comes back with zero failures, and the configuration enabled three plugins. What can and cannot you conclude from that report?

level: middleimportance: must knowfreq 61%

basics

~20 s

You can conclude that on this run, for those three harm classes, at that suite size, nothing failed. You cannot conclude the assistant is safe. Harm classes with no plugin enabled generated no cases, so the report is silent about them. Silence is a missing denominator, not a pass.

open as a page

How do you decide which promptfoo red-team plugins apply to a given application, and what goes wrong if you simply enable the whole catalogue?

level: middleimportance: should knowfreq 50%

basics

~20 s

Work from what the application can actually do: its tools, the data it can reach, who talks to it, and what it must never say. Enable the harm classes it could genuinely commit. Enabling everything costs generation and inference on irrelevant classes and buries real failures in noise nobody triages.

open as a page

In a promptfoo red-team run, one enabled harm-class plugin produces a large batch of failures that triage decides are not real problems for this application. Do you disable that plugin, and what does disabling it cost you?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Not before separating two causes: the application genuinely cannot commit that harm, or it can and the failures are being judged against the wrong expectation. Only the first justifies disabling. Disabling deletes the class from every future run too, so the next report is silently narrower while the headline number improves.

open as a page

You own adversarial-testing coverage for a dozen LLM applications, all scanned with promptfoo's red-team mode. How do you govern which harm-class plugins each application runs, and what breaks when the available catalogue changes between quarters?

level: principalimportance: nice to knowfreq 29%

basics

~20 s

Set a baseline of harm classes every application must run, plus per-application additions derived from its capabilities, with every exclusion carrying a written reason and an owner. Record the enabled set with each report. When the catalogue grows, new classes change the denominator, so cross-quarter comparisons are invalid unless you re-baseline deliberately.

open as a page