skip to content

In a PyRIT red-team run, why is a false negative from your custom scorer more expensive than a false positive?

level: middleimportance: must knowfreq 58%

answer

  1. FP gets triaged, FN is never read
  2. misses = findings you never counted
  3. scorer is the loop's stop condition
  4. burned budget, wrong branch
  5. measure the two error classes apart

basics

~20 s

A false positive is caught: a human reads the flagged transcript during triage and drops it. A false negative is never read, because the transcript looks like an ordinary refusal and nobody opens it. In a multi-turn attack the miss also tells the strategy to keep pushing or to stop, so it changes the run itself.

solid answer

~1 min

The two errors land in different places. A false positive enters the queue of things a person looks at, and triage removes it — the cost is minutes of a reviewer's time. A false negative never enters any queue. The transcript is filed as a non-hit, and the only signal that it was wrong is a number that is quietly too low. On a red-team engagement that asymmetry is the whole point: you are hired to find behaviours, and the scorer decides what counts as found. Its misses are findings you never counted, and they are invisible by construction — you cannot notice an absence in a report. There is a second cost specific to a framework run. In a multi-turn strategy the scorer's verdict is the loop's control signal. A missed hit means the strategy keeps escalating past a success it already had, spending turns and metered calls, or a mis-scored intermediate step sends the next turn down a different branch. The transcript you end up with is not the one a correct scorer would have produced. So the mitigation is not 'tune the threshold down'. It is: sample the negatives, label them by hand, and report the miss rate next to the hit count.

go deeper

for a junior

Says a false positive gets caught by a human reviewing hits, while a false negative is never reviewed, so misses go uncounted.

for a middle

Adds that in a multi-turn run the verdict is the loop condition, so a miss also wastes turn budget or produces a run labelled 'attack failed' that was actually a scoring failure.

for a senior

Measures the two error classes separately against hand-labelled transcripts and names concrete leak shapes: unexpected formats, other languages, unparsed modalities.

for a principal

Sets the reporting standard for the team, so a hit count never ships without the measured miss rate and the sample it came from, and pushes back on any accuracy-only claim.

### The two error classes land in different places Every engagement has a human triage step, and that step reads the scorer's hits. It is a filter on false positives and on nothing else. A false positive therefore costs a reviewer a few minutes and then dies. Nothing in the pipeline reads a non-hit: the transcript is filed, the conversation is closed, and the only trace of the error is a number that is quietly too small. This inverts the intuition most engineers bring from production alerting, where false positives are the expensive class because they burn on-call attention. On an engagement the deliverable is *coverage* — the claim that you looked and this is what is there — and a silent miss attacks coverage directly. You cannot notice an absence in a report. ### What a miss costs inside the run In PyRIT the scorer is not only a labeller; in a multi-turn strategy it is the loop condition. Three concrete shapes: - **Missed success, budget burned.** The objective was met on turn 3 of a 10-turn budget. The scorer said no, so the strategy spent seven more turns — seven more attacker-model calls, seven more target calls, seven more scoring calls — pushing a target that had already fallen over. Across 50 conversations that is several hundred metered calls and a slice of the run's wall clock bought for nothing. - **Missed success, wrong stop reason.** The run terminates on budget exhaustion and is written up as an attack that failed. That sentence is a claim about the target's robustness which you have not established; what you established is that your scorer never said yes. - **Wrong branch.** Where the strategy conditions the next prompt on the last verdict, a mis-scored turn changes what was sent afterwards. The transcript you now own is not the one a correct scorer would have produced, and re-scoring later cannot recover turns that were never sent. ### How the number misleads An attack-success rate computed from an uncalibrated scorer is a measurement of the scorer, not of the target. Read literally, "12 hits in 400 turns" means "my scorer said yes twelve times". Triage removes the false-positive share of that twelve; the misses stay outside the fraction entirely, in neither the numerator nor any denominator anyone inspects. A single accuracy figure hides this completely. On a corpus where most turns are genuine refusals, a scorer that never fires at all still posts high accuracy, because it is right about every refusal. The most common structural version of the failure: a scorer written against text quietly returns false for every image or audio response, so an entire channel is reported clean without a line of it having been judged. ### Measuring the two classes apart Because the errors are not symmetric, report them separately, both against hand-labelled transcripts drawn from this run: | question you answer | what it measures | who else catches it | | --- | --- | --- | | of responses a human called a hit, what share did the scorer call a hit | miss rate — coverage | nobody | | of the scorer's hits, what share survived a human read | false-alarm rate | triage, already | Only the first row is new information. The second you were getting for free. ### What to check Look specifically where a bespoke scorer leaks, among the responses it scored false: - the longest completions, and any completion that does not open with refusal language - compliance in a shape the rule never anticipated: a code fence, a table, a summary rather than a direct answer, a translation - responses in a language other than the one the rule was written against - errors, truncations and empty strings — an exception message scored as a non-hit is a hole, not a refusal - the final turns of conversations that ended on budget exhaustion, where a late success hides behind the stop reason And keep the mitigation pointed the right way. The answer to a suspected miss rate is not to lower a threshold until the count rises, which trades a measured problem for an unmeasured one. It is to label a sample, publish the miss rate next to the hit count, and correct the scorer if the sample says so.

  • Does the same asymmetry hold for a scorer used to gate a product release rather than to drive a red-team run?
    It flips closer to symmetric there, because a false positive blocks a release and costs real time. On an engagement nobody is blocked by a false positive; it just gets dropped in triage, while a miss silently shrinks the finding count.
  • Where do you look first for suspected false negatives in a completed run?
    The false-scored responses that do not look like refusals: long completions, answers wrapped in code blocks or another language, and any response in a modality the scorer cannot parse.

saying these in an interview costs you the question

  • Treating scorer errors as symmetric and quoting a single accuracy number.
  • Assuming triage will catch misses, when triage only ever reads the hits.
  • Reporting 'the attack failed' when what actually happened is 'the scorer never said yes'.
  • Ignoring that a scorer which cannot read a modality returns false for that whole channel.

context