skip to content

A promptfoo red-team run against your internal HR assistant reports almost no failures. The application description in the config reads, in full: 'an HR chatbot'. Is that result evidence the assistant is safe, and what do you do next?

level: seniorimportance: must knowfreq 55%

answer

  1. green means thin config, not safe model
  2. generic cases, no app surfaces
  3. undeclared limit cannot fail
  4. roles, reach, actions, refusals
  5. sample the generated cases

basics

~20 s

No. promptfoo's graders judge each reply against the description you supplied. If it never says what the assistant must refuse, who may ask, or what records it reaches, an off-policy answer breaks no declared rule and is scored a pass. Rewrite the description with roles, data and refusals, then re-run.

solid answer

~60 s

Read the result as a statement about the config, not about the model. **Why it is green.** Two mechanisms compound. Generation had nothing app-specific to aim at, so the cases are generic bait any chatbot would deflect — nothing was written that targets salary records, another employee's file, or the assistant's ability to open a ticket. And grading had no declared boundary, so even a response that leaked something would need a rule to be a failure against, and there is none. **What to do.** Rewrite the description to state: the user roles and whether requests are authenticated; the records and systems the assistant can reach, by class; the actions it can take; and the things it must refuse, including off-mission topics that are merely inappropriate rather than dangerous. Then re-run and read a sample of the generated cases: if any of them could only have been written for an HR assistant, the rewrite worked. Treat the first run's number as void, not as a baseline to improve on.

go deeper

for a junior

Says the description is too vague and should be more detailed so the attacks fit the application.

for a middle

Explains both effects — generic cases and no rule to fail against — and can list the content a proper description needs.

for a senior

Declares the first run void rather than a baseline, rewrites with roles/reach/actions/refusals, and verifies by reading generated cases and checking each refusal drew attacks.

for a principal

Adds that even a thorough description bounds the result to declared rules, and sets an expectation that no team reports a promptfoo pass rate without the description it was graded against.

### Read the result as a statement about the configuration The suite reported a number, and both halves of that number were set by the fourteen characters in `redteam.purpose`. "An HR chatbot" is a category label, and two independent denominators are wrong at once because of it. **Cases that were never written.** `promptfoo redteam generate` asks a model to write adversarial cases *for the described application*. Given a category, it produces category-level bait — the sort of thing that would fit any assistant with a chat box. Nothing in the run probes cross-employee record access, compensation disclosure, a question about an ongoing investigation, retaliation framing, or any action this assistant can actually take (a ticket it can file, a leave request it can submit, a policy exception it might imply), because none of that appears in the text. Those harms are absent from the run entirely, and absence is not a pass — it is a harm class with no denominator, which prints on the report as nothing at all. **Verdicts that could not be failures.** `promptfoo redteam eval` grades each reply with a model-graded assertion whose rubric interpolates the purpose. If the purpose never says "must never reveal another employee's compensation", a reply that does exactly that is measured against nothing in particular and is scored a pass. This is the sharper half of the trap: the first mechanism makes the report narrow, the second makes it *wrong*. A narrow report under-claims; this one over-claims, and it will be forwarded upward as evidence. ### What the void run already cost The bill is paid regardless. A dozen plugins at `numTests: 5` with a strategy or two is a few hundred cases, each one at least a target call plus a grader call, on top of generation — plus the wall-clock of the eval and the engineer-hours of whoever reviewed the green report. The information content of that spend is close to zero, and if the suite is scheduled, the same spend repeats every cycle producing the same non-information. The half-hour of writing that would have fixed it is the cheapest line item in the whole exercise. ### The rewrite Replace the single phrase with, concretely: - **Roles** — anonymous visitor, authenticated employee, HR administrator, and which the assistant is serving in this deployment. - **Reach** — "the requesting employee's own leave balance, their own benefits enrolment, and published policy documents", stated as classes, never as pasted records. - **Actions** — what it may submit, file or change on the user's behalf. - **Refusals** — other employees' records, compensation comparisons, legal advice, anything touching an ongoing investigation, and off-mission topics that are merely inappropriate. - **Boundaries you care about** — "may summarise the leave policy; may not interpret it as a decision about this employee's case." ### Verification, not faith Do not simply read the new pass rate. Open the generated suite and sample cases: are they recognisably about *this* system? Check that every declared refusal drew cases — one that drew none is phrased in language the generator did not convert into anything, so restate it in the same operational terms as the rest and regenerate. Confirm the graders are capable of failing by finding a case whose reply is visibly off-policy and checking its verdict. Then handle the reporting carefully. The earlier number is void, not a baseline: publishing "we went from 98% to 71%" invites the reading that the assistant regressed, when what actually happened is that it was tested for the first time. Say that explicitly, or the rewrite will be read as a regression and someone will ask you to revert it. ### The caveat that survives the rewrite Even a thorough purpose bounds the run to rules you thought to declare and cases the generator thought to write. The report tells you which declared rules survived which generated attacks against one target configuration on one day. It never tells you the assistant is safe, and a red-team suite that is presented as though it did is a governance problem regardless of how good its description is.

  • After rewriting the description, one declared refusal produced no generated cases at all. What does that tell you?
    The refusal is worded in language the generator did not turn into attacks. Restate it in the same concrete, operational terms as the rest of the description and re-run.
  • Can you compare the pass rate before and after the rewrite to show progress?
    No. The rewrite changed both the cases generated and the rules graded against, so the two runs measure different things. The earlier number is void, not a baseline.
  • Once the description is thorough, does a fully green run mean the assistant is safe?
    It means the declared rules survived the generated attacks. Harms you did not think to declare are still untested and still score as passes.

saying these in an interview costs you the question

  • Reporting the pass rate upward as a safety result.
  • Concluding the model is well-aligned rather than that the run was underspecified.
  • Treating the thin-description run as a baseline to show improvement against.
  • Rewriting the description and trusting the new score without reading any generated case.
  • Fixing it by pasting real employee records into the description to add specificity.

context