skip to content

In a promptfoo red-team run, one enabled harm-class plugin produces a large batch of failures that triage decides are not real problems for this application. Do you disable that plugin, and what does disabling it cost you?

level: seniorimportance: should knowfreq 38%

answer

  1. inapplicable vs mis-expected vs real
  2. uniform batch means mismatch
  3. scattered failures mean something
  4. disabling shrinks future denominators
  5. builders triage charitably

basics

~20 s

Not before separating two causes: the application genuinely cannot commit that harm, or it can and the failures are being judged against the wrong expectation. Only the first justifies disabling. Disabling deletes the class from every future run too, so the next report is silently narrower while the headline number improves.

solid answer

~60 s

The batch tells you the class is firing; it does not tell you why the verdicts feel wrong. Three causes look identical in the report. **Inapplicable class** — the application has no capability or data that could produce this harm, so every case is theatre. Disabling is correct, with the reason recorded. **Applicable but mis-expected** — the application can commit the harm, but the behaviour being flagged is one the business has explicitly accepted, or the judgement of what counts as a violation does not match your rules. The fix is to narrow the expectation for that class, not to delete the class. **Actually real** — triage is wrong, usually because the reviewer knows the design intent and reads the response charitably. This is the most expensive case to get wrong. Disabling is the irreversible-feeling option because it is invisible afterwards: the class stops appearing, nobody notices the denominator shrank, and the pass rate goes up for the wrong reason. If you do it, record it beside the result, not only in the config.

go deeper

for a junior

Should recognise that a noisy class is not automatically a wrong class, and that switching it off needs a reason.

for a middle

Separates inapplicable from applicable-but-mis-expected, and knows that disabling removes the class from future runs, not just this one.

for a senior

Investigates from sampled exchanges, reads the batch shape as a diagnostic, distrusts builder-run triage, and publishes any enabled-set change with the results.

for a principal

Owns the policy: who may disable a class, what evidence is required, where the reason lives, and what change to the application forces a re-review of every disabled class.

This is the moment where a red-team programme either keeps its integrity or quietly hollows out, so treat the decision as a small investigation rather than a one-line edit to `redteam.plugins`. ### Separate the three causes, because they look identical in the report **Inapplicable class.** The application has no tool and no data path that could produce this harm, so every case is theatre and every verdict is noise. Removing the entry is correct — with the capability argument written down beside it. **Applicable but mis-expected.** The application *can* commit the harm, but the specific behaviour being flagged is one the business has explicitly accepted, or the grader's notion of a violation does not match your rules. The fix is to narrow what counts as a violation for that class, not to delete the class. **Actually real.** Triage is wrong. This is the most expensive case to get wrong, and it is more common than people think, because the reviewers are usually the people who built the system and they read their own application's outputs charitably. Triage decisions on your own service deserve a second reader who did not build it. To tell them apart, sample the flagged exchanges themselves — the transcripts in the report, not the summary counts. For each one ask: with the tools and data this application actually has, could the harmful outcome have occurred at all? If no capability exists, the class is inapplicable. If it does exist, the remaining question is whether the flagged reply is genuinely acceptable to the business or merely familiar to the reviewer. ### Read the shape of the batch as a diagnostic A class failing on nearly every generated case usually means a mismatch between what the class judges and what the application is for — an assistant designed to give regulated advice will fail a class built to penalise exactly that, every single time. A class failing on a scattered minority of cases usually means something real is happening in an edge case. The uniform batch is the one teams reach to disable, and it is also the one most likely to be fixable by tightening the expectation instead. ### The cheaper alternatives to deletion Deleting is rarely the only lever. Keeping the entry and reducing its `numTests` cuts the noise and the bill proportionally while preserving a non-zero denominator, so the class still appears in every report. Where a plugin accepts grading configuration — for example supplying labelled `graderExamples` in that entry's `config`, showing the judge outputs your business considers acceptable and unacceptable — you keep the class and fix the rubric, which is the correct repair for the mis-expected case. Both options cost minutes. Both leave a class you can still be surprised by. ### What deletion actually costs Not "this run is quieter". Every future scheduled run loses the class, and the report will not say *omitted* — it will simply have nothing there. Six months later a reader with no memory of this conversation sees a clean report over a narrower taxonomy and reads it as the application having got safer. Meanwhile the pass rate mechanically improves, because the denominator lost the class that was contributing all the failures: an improvement produced entirely by measuring less. And if a capability is added later that makes the class relevant again, nothing in the pipeline fires. That combination — invisible, flattering, and silent on re-emergence — is what makes deletion the option to justify rather than the default. ### What I do in practice If the class is inapplicable: remove the entry and record the capability argument next to it, so the next reviewer can challenge it. If it is applicable but mis-expected: keep the entry, adjust what counts as a violation, and note the adjustment with the results, since a rubric change also moves the number. If triage was wrong: fix the application. In every case, any change to the enabled set is published *with the next report* — a pass rate that improved because a class left the suite is the single most misleading artefact this kind of tooling produces. And every excluded class is re-opened when the application's tools or data access change, because inapplicability is a property of the current build, not a permanent fact about the product.

  • What does the shape of the failure batch suggest — uniform versus scattered?
    Near-total failure usually means the class is judging something the application was never meant to do. A scattered subset is more often a genuine behaviour worth chasing.
  • Why is disabling more dangerous than it feels at the time?
    The class vanishes from every later report with no marker. A future reader sees a clean, narrower run and reads it as improvement, and no signal fires if the capability later becomes relevant.
  • What guards against builder-charitable triage?
    A second reader who did not build the application, and a written rule for what counts as acceptable in that class, decided before the batch was seen rather than after.

Removing the noisy class is like unplugging the smoke alarm that keeps going off in the kitchen: the room gets quiet immediately, and nothing afterwards ever tells you it is not listening.

saying these in an interview costs you the question

  • Disabling the class the same day, purely to make the report quiet.
  • Never inspecting individual flagged exchanges before deciding, working only from the summary.
  • Letting the team that built the application be the sole judge of whether its own failures are real.
  • Changing the enabled set without recording it alongside the results, so later pass rates look like improvement.
  • Treating an inapplicability verdict as permanent even after the application gains new tools or data access.

context