How do you decide which promptfoo red-team plugins apply to a given application, and what goes wrong if you simply enable the whole catalogue?
answer
- capabilities, data, callers, rules
- applicable / not applicable / deferred
- irrelevant failures kill triage
- noisy is not the same as inapplicable
- re-select when capabilities change
basics
~20 sWork from what the application can actually do: its tools, the data it can reach, who talks to it, and what it must never say. Enable the harm classes it could genuinely commit. Enabling everything costs generation and inference on irrelevant classes and buries real failures in noise nobody triages.
solid answer
~60 sI derive the plugin selection from the application's own capabilities rather than from the catalogue. Four questions drive it. **What can it do?** — a read-only FAQ bot and an assistant that can issue refunds have different harm surfaces. **What can it reach?** — customer records, internal documents, other tenants' data. **Who talks to it?** — anonymous public users, authenticated staff, minors. **What must it never say?** — the domain rules that are specific to this business, which usually need an application-specific policy class rather than a stock one. Enabling everything looks like the safe default and is not. You pay generation and inference on classes the application structurally cannot commit, and — worse — you get failures graded against harms that do not apply, which teach the team that the tool cries wolf. Once triage stops happening, the whole run is decoration. The opposite error is real too: dropping a class because it is inconvenient rather than because it is inapplicable. Those are different reasons and only one of them is defensible.
go deeper
Should say the selection depends on what the application does and what data it touches, and that enabling everything wastes time on irrelevant classes.
Derives the set from capabilities, data reach, caller population and business rules, and explains the noise cost and the false-economy of both extremes.
Produces a reviewable artefact with reasons for every disabled class, distinguishes noisy from inapplicable, and ties re-selection to capability changes rather than the calendar.
Frames selection as a risk decision with named owners, treats application-specific policy classes as first-class, and builds the review trigger into the change process instead of the scan schedule.
Plugin selection is a risk-modelling exercise wearing configuration clothes. It is worth doing deliberately once per application, and re-doing it when the application changes, rather than letting it accrete by copy-paste. ### The derivation Write down the capability surface before you open the catalogue. Four columns: | Column | What goes in it | Classes it implies | |---|---|---| | Actions | tools it can call, side effects it can cause | agency and authorization classes such as `excessive-agency`, `bfla`, `rbac` | | Data | corpora, records, other tenants' rows it can reach | disclosure classes such as the `pii` family, `bola`, `cross-session-leak`, `prompt-extraction` | | Callers | anonymous public, authenticated staff, minors, regulated users | the harm families whose consequence depends on audience | | Rules | what this business is contractually or legally bound never to say | a `policy` entry, plus `contracts`, `competitors`, `imitation` where relevant | Then walk the catalogue against those columns and mark every class one of three ways: **applicable**, **not applicable because the capability does not exist**, or **deferred, with an owner and a date**. That three-way marking is the artefact. It turns an opaque `redteam.plugins` list into a reviewable decision, and it makes the next person's question — "why isn't this class on?" — answerable without archaeology through git history. ### Why enabling the whole catalogue fails Three costs compound, and one of them is counter-intuitive. **Compute and time.** The catalogue is dozens of plugins, and collection aliases hide that: `plugins: [harmful]` is a single line and twenty-odd classes. At the default five cases each, the full catalogue is a few hundred base cases before `redteam.strategies` multiplies them; each case is a target call plus a grader call. That is tens of minutes and a real, if modest, bill — paid again on every scheduled re-run, forever. **A flattered number.** This is the part people get backwards. Enabling classes the application structurally cannot commit does not lower the pass rate — it *raises* it. A read-only FAQ bot passes every authorization and data-disclosure class trivially, because there is nothing to authorize and no data to leak. Twenty free passes dilute the two classes that actually matter, so the aggregate improves precisely because you tested things that were never at risk. Anyone comparing headline pass rates across configurations is measuring selection, not safety. **Triage collapse.** The one that kills programmes. Findings graded against inapplicable harms are not findings; they are distractions, and a report with a long tail of them gets skimmed, then ignored, and the two real failures go with it. Once the team believes the tool cries wolf, the run is decoration and the money keeps being spent. ### Why the minimal set fails too The classes people quietly drop are usually the ones that were noisy, and noisy correlates with *finding something*. Worse, applicability is a property of the current build. The day someone wires a tool that can move money or read another tenant's rows, a class that was genuinely inapplicable becomes the most important entry in the file, and nothing in the pipeline will announce that. Selection therefore needs a review trigger tied to capability changes, not to a calendar reminder. ### Application-specific harms No stock class knows that your assistant must never comment on an open legal matter, or must never quote a price outside the published tariff. That is what promptfoo's `policy` plugin is for: an entry whose `config.policy` states the rule in your own words, which the generator then attacks and the grader then enforces. In most real deployments these produce more actionable failures than the generic families, because they encode the rules the business actually gets sued over. ### What I check on an existing config - Does every class that is *not* in `redteam.plugins` have a written reason, and is that reason a capability argument rather than "it was noisy last quarter"? - Are there any `policy` entries at all? A config with none is testing generic harms and no business rules. - Are aliases used where a resolved list belongs? An alias is convenient and it hides both the case count and the fact that its membership can change under you. - When was selection last revisited, and against which change to the application's tools or data access?
- What is a defensible reason to disable a promptfoo red-team plugin, and what is not?Defensible: the application structurally cannot commit that harm — no such capability, no such data. Not defensible: it was noisy, slow, or produced findings the team did not want to own.
- What event should trigger a re-review of the enabled plugin set?Any change to the capability surface — a new tool, a new data source, a new caller population — not a calendar date.
- Why do application-specific policy classes often produce more useful failures than stock ones?They encode the rules that are actually enforceable against your business. Stock classes cover generic harms; only you can state that the assistant must never discuss an open dispute.
saying these in an interview costs you the question
- Enabling the entire catalogue and calling it thorough, with no account of the cost or the noise.
- Choosing classes by what looks impressive in a report rather than what the application could commit.
- Disabling a class because it produced awkward failures, without a capability argument.
- Never adding an application-specific policy class, so business rules go completely untested.
- Treating the selection as set-once when the application gains new tools or data access.