In PyRIT, what is a scorer's job during an attack run, and what does its verdict do to the run itself?
answer
- verdict is the loop condition
- stops on a hit, never on a miss
- no scorer, no result
- read the stop-turn response
- budget exhausted is the other exit
basics
~20 sIn PyRIT, the scorer reads the target's response and decides whether the attack objective was met. In a multi-turn run that verdict is control flow: a positive verdict ends the run early and marks it a success. So the scorer, not you, decides when testing stops.
solid answer
~50 sA PyRIT scorer is the run component that reads a response and returns a verdict on whether the objective was achieved. Two things follow from that. First, it is the only thing that turns raw transcripts into a result. Nothing else in the run labels an attempt, so every success number you report traces back to one scorer configuration. Second, in a multi-turn attack the verdict is the loop condition. The strategy keeps sending turns until the scorer says the objective was met or the turn budget is exhausted, and a positive verdict stops the conversation at that turn. That is why the scorer is not a reporting step you bolt on afterwards. It changes which conversations get explored and how deep they go: an over-eager one stops on a hedged non-answer and you never see the turns that would have followed, while a blind one keeps pushing past a real success and spends the budget for nothing.
go deeper
Should say the scorer decides whether the objective was achieved, and that a positive verdict ends a multi-turn run.
Should add that the verdict is the loop condition alongside the turn budget, and that scoring is per-turn work with a cost.
Should frame scorer error as experiment-design error at runtime - a false hit destroys the turns that would have followed - and describe checking the stop-turn responses.
Should talk about who owns the scorer configuration behind a published success rate, and about re-scoring stored transcripts as the standing check on it.
## What the object actually is In PyRIT a **scorer** is an object deriving from the `Scorer` base class with one job: it is handed a piece of a target's response together with the objective string for the run, and it returns a list of `Score` objects. A `Score` is richer than a boolean. It carries: - `score_value` (a string - `"True"`/`"False"` for a true/false scorer, a normalised number for a float-scale one); - `score_type`; - `score_category`; - `score_rationale` (the judge's own written justification for the verdict); - and the id of the response piece it judged. Every score PyRIT produces is written to memory next to the conversation it came from, which is what makes after-the-fact re-reading and re-scoring possible at all. ## Where it sits in the run PyRIT's multi-turn attacks - the red-teaming attack, Crescendo, the tree-search attack - are constructed with a scoring configuration that names one **objective scorer** and, optionally, a list of **auxiliary scorers**. Only the objective scorer steers, and it must produce a true/false verdict, because the attack loop consumes exactly one bit per turn. One turn is: 1. the adversarial chat model composes the next attacker prompt; 2. any converters transform it; 3. the prompt target answers; 4. the objective scorer judges that answer. A true verdict breaks the loop and the attack result is recorded as achieved. A false verdict sends another turn, until the configured maximum number of turns is reached and the run ends unachieved. Auxiliary scorers annotate every response for the record and change nothing about control flow. So exactly two things end a run: **a positive verdict, or the turn budget**. Nothing else in the loop labels an attempt, which means every success number you will ever report traces back to one scorer configuration. ## What it costs With a model-backed scorer, a turn is three metered calls: attacker model, target, scorer. A campaign of 100 objectives with a budget of 8 turns is up to 800 turns and therefore up to ~2,400 model calls. The scorer's call is usually the cheapest of the three - a short rubric prompt in, a couple of hundred tokens of verdict and rationale out - so the money is small, often single-digit dollars for a campaign that size. The real bill is **wall-clock and attention**. The scorer sits on the critical path of every turn, so its latency (typically a second or a few) is added to every turn of every conversation, serialised; a cheap local scorer such as a substring check costs nothing and no time, and is coarse in exactly the way described next. Engineer time is the largest line of all: someone writes the rubric, labels transcripts and reads stop turns. ## Where the number misleads The headline of a PyRIT campaign is an **objective-achieved rate**, and that rate is a *scorer output*, not a measurement of the target. Because the verdict is control flow, a scorer error is not a reporting error - it is an experiment-design error committed at runtime, because it changes which data ever exists. - A **false positive** stops the conversation, so the turns that would have shown genuine compliance, or genuine resistance, were never generated; the run is filed as a success and its conspicuously short transcript is easy to skim past. - A **false negative** lets the run burn its full budget and files a real success as a miss - expensive and undercounted, but the compliance is still in memory and a corrected scorer can recover it. The rate also moves when nothing about the target moved: swap the judge model behind a self-ask scorer, or edit its rubric, and this month's number is not comparable with last month's. And "achieved" means one scorer believed one response met one objective; it is not a claim that the output was harmful, usable or reproducible. ## What you would check - Sort the runs by executed turns and read the response at the **stop turn** of the short ones, together with the stored rationale - that is the exact text the scorer acted on, and no other sample answers the question. - Then look at the other tail: a batch that always exhausts the budget looks identical whether the target resisted or the scorer is blind. - **Smoke-test the scorer** by handing it a handful of stored responses you already know are compliant and confirming it says so. - Finally, **re-score** the persisted conversations offline with a second scorer; that costs scorer calls only, since no target turns are repeated, and every disagreement is a transcript worth reading.
- If the scorer never returns a positive verdict, what ends the run?The turn budget. The run terminates as exhausted, and from that fact alone you cannot tell a resistant target from a scorer that fails to recognise compliance - a batch that always burns the full budget is as often a scorer symptom as a target one.
- Can you score a conversation after the run has finished?Yes, if the run's conversations were persisted. Re-scoring stored transcripts is the cheapest way to compare scorers, because no target turns are repeated. What it cannot recover are the turns a premature stop never generated.
- Does the scorer decide what gets sent next?Not directly - the attack strategy composes the next turn. But the scorer decides whether there is a next turn at all, which is why its errors change the transcript, not just the label on it.
The scorer is the smoke alarm wired to the fire drill's stop button: when it sounds, the drill ends right there. A false alarm does not just log a wrong event - it means nobody ever finds out what the rest of the building would have done.
saying these in an interview costs you the question
- Describes the scorer as a reporting step that runs after the attack finishes
- Cannot say what terminates a multi-turn run
- Treats a positive verdict as proof the objective was met without opening the transcript
- Never asks what text the scorer was actually shown