A batch of PyRIT multi-turn runs reports a high objective-achieved rate, but most transcripts stop on an early turn at a hedged, non-compliant answer. What is happening, and how do the two directions of scorer error differ in what they cost you?
answer
- short transcripts + high success = scorer firing early
- false hit truncates, false miss bills
- the miss is recoverable by re-scoring
- stopped run looks like a real one
- early-turn achievement = review trigger
basics
~20 sThe scorer is firing on responses that are not real compliance, and because a positive verdict ends the run, every false hit is also a truncated attack. False positives inflate the success rate and delete the turns you never sent. False negatives cost the whole turn budget and undercount. Read the stop-turn responses.
solid answer
~60 sThe symptom is diagnostic: a high achieved rate whose transcripts are short and end on non-compliance is a scorer firing early, not a weak target. The two error directions are not symmetric, because the verdict is control flow. A **false positive** does two kinds of damage at once. It adds a success that is not one, and it terminates the conversation, so the turns that would have shown the target actually holding the line - or actually breaking - were never generated. You cannot fix that in the report; the evidence does not exist. A **false negative** is loud but cheaper. The run continues to its turn budget, so you pay the full call cost and the real success is filed as a miss. The transcript, however, still contains the compliance, so a later re-score can recover it. So the recoverable error is the miss; the unrecoverable one is the false hit. That asymmetry is the argument for tuning the stop rule conservatively when transcripts are persisted and re-scoring is available.
go deeper
Should notice the mismatch between the reported success and the transcripts and say the scorer needs checking before the numbers are used.
Should explain that a positive verdict ends the run, so a false hit both inflates the rate and shortens the transcript.
Should state the asymmetry explicitly - the miss is recoverable by replay, the false hit is not - and give a concrete recovery plan of sample, re-score, re-run.
Should set the program's default error direction deliberately and require early-turn successes to be reviewed before any rate leaves the team.
**Reading the symptom.** The batch in front of you has a high objective-achieved rate and transcripts that stop early on hedged, non-compliant answers. That pairing is diagnostic, and it is not a weak target: it is a scorer firing before compliance happened. The evidence lives in two run-level fields PyRIT records for every attack result - the outcome, and how many turns actually executed. Sort by executed turns and read the last response of the short runs, together with the stored rationale the scorer wrote. If those responses read as deflections, partial refusals or topic-adjacent hedges, the verdict is the fault. The opposite tail needs the same treatment for the opposite reason: runs that exhaust the budget look identical whether the target held or the scorer is blind, so exhaustion alone proves nothing either. **Why the two error directions are not symmetric.** Because the objective scorer's verdict is the loop condition, its errors do not merely mislabel data - they decide which data gets produced. - **False positive.** The loop breaks. The rate is inflated by a success that is not one, *and* the turns that would have shown real compliance or real resistance were never sent. The report cannot repair this, because the evidence does not exist. Worse, the error is self-concealing: in any summary a prematurely stopped run is indistinguishable from a genuine one, and its short transcript is the easiest thing in the corpus to skim past. - **False negative.** The loop keeps going and ends on the turn budget. You pay attacker, target and scorer calls for every remaining turn, and a real finding is filed as a miss. But the compliance is sitting in persisted memory, so a corrected scorer replayed offline recovers it without touching the target again. One direction is a bill; the other is a claim you cannot audit. **What it costs, in both directions.** A false negative's price is arithmetic: the remaining turns times three metered calls each. Across a campaign of a few hundred objectives that is a bounded number of dollars and some wall-clock - annoying, recoverable. A false positive's price is a re-run of the affected objectives, which costs the *entire* conversation again (attacker, target and scorer from turn one), plus the human time to find out you needed to. And there is an incentive gradient worth naming out loud: an over-eager scorer produces a higher success rate on shorter conversations, so the misconfiguration that damages your findings most is also the one that makes the campaign look cheaper and more productive. Nothing in the tooling will flag that for you. **Where the numbers mislead.** Both headline metrics are corrupted in the same direction by the same fault. The achieved rate goes up; the mean turns-to-success goes down, which reads as "the target is easy to break quickly" when it actually means "the scorer stops early". A "success on turn one" bucket is the single most informative slice in the corpus and almost nobody looks at it. And any comparison across time or across endpoints inherits the fault: if the scorer changed between two campaigns, the delta measures the scorer. **Recovering the batch you already have.** Sample the stop-turn responses and label them yourself; that gives the false-positive rate directly, on the exact population that produced the headline. Re-score the persisted conversations with a tightened scorer to get a corrected rate for the runs you already paid for - this recovers the misses for free. Then re-run the objectives whose runs were truncated, because those need turns that were never generated; re-scoring cannot conjure them. Do not republish the old number with a footnote: the two halves of the fix are different operations and only one of them is cheap. **Preventing the repeat.** Tune the stop rule against a frozen set of your own hard cases rather than by intuition. Persist the raw score next to the verdict so the corpus can be re-thresholded without new model calls. Treat "achieved on turn one or two" as a review trigger rather than a headline. And choose the default error direction on purpose: when transcripts are persisted and re-scoring is available, bias the loop toward continuing, because a miss is a bill you can pay later and a false hit is a lie you cannot audit.
- Which of the two error directions can be repaired without touching the target again?The false negative. The run went to full budget, so the compliance is in the stored transcript and a corrected scorer replayed offline recovers it. A false positive destroyed turns that were never generated, so only a re-run fixes it.
- The same batch instead shows every run exhausting its budget with zero successes. What do you check first?Whether the scorer can recognise success at all - hand it a handful of responses you know are compliant, from your own stored transcripts, and see what it says. Exhaustion is the shared exit of a resistant target and a blind scorer.
- How do you keep this from recurring across a campaign?Tune the stop rule against labelled hard cases, persist the raw score alongside the verdict so the corpus can be re-thresholded, and flag runs that report success on the first turn or two for human reading.
A false hit is a referee blowing the whistle early: the play stops, so nobody ever learns what would have happened next. A false miss just lets the clock run - expensive, but the tape still shows everything.
saying these in an interview costs you the question
- Reports the success rate as-is and blames the target
- Never opens the response at the turn where the run stopped
- Treats false positives and false negatives as equally costly because both are 'one wrong label'
- Believes re-scoring the stored transcripts fully repairs a batch of premature stops
- Raises the turn budget in response to a false-positive problem