You own adversarial-testing coverage for a dozen LLM applications, all scanned with promptfoo's red-team mode. How do you govern which harm-class plugins each application runs, and what breaks when the available catalogue changes between quarters?
answer
- baseline plus per-app plus exclusions
- map to the risk register
- publish the enabled set
- catalogue growth moves the denominator
- re-baseline with an overlap period
basics
~20 sSet a baseline of harm classes every application must run, plus per-application additions derived from its capabilities, with every exclusion carrying a written reason and an owner. Record the enabled set with each report. When the catalogue grows, new classes change the denominator, so cross-quarter comparisons are invalid unless you re-baseline deliberately.
solid answer
~50 sI split the decision into three layers. **Baseline** — the harm classes every application in the portfolio runs, mapped to entries in the organisation's risk register so the selection can be justified to an auditor rather than to a tool. **Per-application** — classes added because this system has tools, data or a caller population the baseline does not assume, plus application-specific policy classes for its own business rules. **Exclusions** — every class not run carries a stated capability reason, a named owner and a review date. An exclusion with no reason is a defect, not a config. The artefact that makes this work is the enabled set published with each result. Without it, no number is comparable to any other number, across time or across applications. When the catalogue changes, the denominator changes. Adding classes usually depresses pass rates portfolio-wide, which looks like regression and is not. I re-baseline on a stated date, announce it, and treat the two sides as separate series.
go deeper
Not expected to answer at this scope; at most, that different applications need different harm classes and that the choice should be written down.
Describes a shared baseline plus per-application additions, and recognises that pass rates are not comparable when the enabled sets differ.
Adds exclusion hygiene with owners and review dates, insists the enabled set is published with every result, and handles catalogue drift as a measurement change.
Ties the baseline to the risk register, refuses cross-application ranking and single portfolio scores, designs the re-baseline with an overlap period, and instruments selection drift as its own signal.
At portfolio scale the hard problem stops being "can we run promptfoo" and becomes "can anyone believe the numbers". Selection governance is the answer to the second question. ### A three-layer selection **Baseline.** The harm classes every application in the portfolio runs, expressed as a shared `redteam.plugins` fragment that each application's `promptfooconfig.yaml` composes in. Map each baseline class to an entry in the organisation's risk register, because the risk register is the language a launch decision is actually argued in; a selection justified only by "it is in the tool" does not survive an auditor. **Per-application additions.** Classes added because this system has tools, data reach, a caller population or regulated content the baseline does not assume, plus `policy` entries carrying its own business rules. **Explicit exclusions.** Every catalogue class not run carries a capability argument, a named owner and a review date. An exclusion with no reason is a defect, not a configuration. Exclusions decay faster than anything else in this file, so any change to an application's tools or data access re-opens all of that application's exclusions at once. ### Comparability, and the ranking trap Two applications with different plugin lists have different denominators, so their pass rates are not on the same axis and must never be ranked side by side on a dashboard. Ranking is the failure mode that quietly corrupts the whole programme: it gives every team an incentive to narrow its own selection, and once selection is a performance lever the taxonomy stops describing risk. Publish per-application results with their resolved plugin list attached, and resist any single portfolio-wide safety score — a compressed number is exactly the artefact that loses the context that made it meaningful. ### Catalogue drift, and the version trap underneath it New harm classes appear between promptfoo releases, and existing collection aliases absorb them. This is the mechanism worth naming precisely: `plugins: [harmful]` or `plugins: [owasp:llm]` resolves to a **version-dependent** set. Upgrade promptfoo and that one unchanged line can expand to more plugins than it did last quarter, generating more cases, exposing more surface and lowering the pass rate — with no diff in any application's config to explain it. A team can spend a week hunting a regression that was a dependency bump. Three effects arrive together whenever the catalogue moves: applications that adopt new classes get worse-looking numbers for reasons unrelated to their behaviour; applications that do not adopt them look better than their peers for no merit; and any trend line spanning the change compares two different measurements. The disciplined handling: pin the promptfoo version and record it with every result; prefer resolved plugin lists over aliases in the stored artefact even if aliases are used in the config; re-baseline on an announced date; run one overlap period where both the old and new sets are scanned so the step change is quantified rather than argued about; then treat pre- and post-baseline results as separate series. The same discipline applies in reverse when a class is dropped. ### What this costs at twelve applications Twelve applications, twenty-five classes each, five cases per class is roughly fifteen hundred base cases per cycle before strategies multiply them, and each case is a target call plus a grader call. Nightly, that is a standing bill and a couple of hours of machine time — affordable. What is not automatically affordable is the triage: findings that nobody reads are the same as findings that were never generated, except more expensive. Budget the reviewer time before you budget the tokens, and let that budget, not the catalogue, set how broad each selection can honestly be. ### What I would instrument Selection drift per application over time; the count and age of unreviewed exclusions; the promptfoo version each result was produced under; and — the cheapest early warning in the whole programme — whether any application's resolved plugin list or per-class case count *shrank* between consecutive runs. A rising pass rate next to a shrinking selection is a denominator change wearing the costume of an improvement. ### What I would resist A mandate to enable everything everywhere, which converts the programme into a noise generator, flatters every aggregate with easy passes on inapplicable classes, and destroys triage. And a single number reported upward, which is how a narrow scan becomes a broad clearance in someone else's slide.
- Why should applications with different enabled harm-class sets not be ranked against each other?Their pass rates have different denominators. Ranking creates an incentive to narrow the selection, which turns the risk taxonomy into a scoring lever.
- How do you absorb newly available harm classes without inventing a false regression?Announce a re-baseline date, run one overlap period scanning both the old and new sets to quantify the step, and treat the periods as separate series rather than one trend.
- What single metric gives the earliest warning that a result improved for the wrong reason?Whether an application's enabled set shrank between consecutive runs. A rising pass rate alongside a shrinking selection is a denominator change, not an improvement.
Adding harm classes to the baseline is like adding two sections to an annual exam and then comparing this year's average with last year's. The average fell because the paper changed, not because the candidates got worse.
saying these in an interview costs you the question
- One portfolio-wide safety score with no enabled-set context attached.
- Ranking applications against each other when their harm-class selections differ.
- Exclusions with no reason, no owner and no review date.
- Reading a portfolio-wide drop after the catalogue grew as a real regression.
- Mandating the full catalogue everywhere and then wondering why nobody triages the reports.