skip to content

Scorers

The scorer is a run component whose verdict stops a multi-turn attack, so an uncalibrated one decides your results before you read them. Interviewers ask how you picked it and how you checked it.

on this pageshow

explore

questions

5

In PyRIT, what is a scorer's job during an attack run, and what does its verdict do to the run itself?

level: juniorimportance: must knowfreq 72%

answer

  1. verdict is the loop condition
  2. stops on a hit, never on a miss
  3. no scorer, no result
  4. read the stop-turn response
  5. budget exhausted is the other exit

basics

~20 s

In PyRIT, the scorer reads the target's response and decides whether the attack objective was met. In a multi-turn run that verdict is control flow: a positive verdict ends the run early and marks it a success. So the scorer, not you, decides when testing stops.

solid answer

~50 s

A PyRIT scorer is the run component that reads a response and returns a verdict on whether the objective was achieved. Two things follow from that. First, it is the only thing that turns raw transcripts into a result. Nothing else in the run labels an attempt, so every success number you report traces back to one scorer configuration. Second, in a multi-turn attack the verdict is the loop condition. The strategy keeps sending turns until the scorer says the objective was met or the turn budget is exhausted, and a positive verdict stops the conversation at that turn. That is why the scorer is not a reporting step you bolt on afterwards. It changes which conversations get explored and how deep they go: an over-eager one stops on a hedged non-answer and you never see the turns that would have followed, while a blind one keeps pushing past a real success and spends the budget for nothing.

go deeper

for a junior

Should say the scorer decides whether the objective was achieved, and that a positive verdict ends a multi-turn run.

for a middle

Should add that the verdict is the loop condition alongside the turn budget, and that scoring is per-turn work with a cost.

for a senior

Should frame scorer error as experiment-design error at runtime - a false hit destroys the turns that would have followed - and describe checking the stop-turn responses.

for a principal

Should talk about who owns the scorer configuration behind a published success rate, and about re-scoring stored transcripts as the standing check on it.

## What the object actually is In PyRIT a **scorer** is an object deriving from the `Scorer` base class with one job: it is handed a piece of a target's response together with the objective string for the run, and it returns a list of `Score` objects. A `Score` is richer than a boolean. It carries: - `score_value` (a string - `"True"`/`"False"` for a true/false scorer, a normalised number for a float-scale one); - `score_type`; - `score_category`; - `score_rationale` (the judge's own written justification for the verdict); - and the id of the response piece it judged. Every score PyRIT produces is written to memory next to the conversation it came from, which is what makes after-the-fact re-reading and re-scoring possible at all. ## Where it sits in the run PyRIT's multi-turn attacks - the red-teaming attack, Crescendo, the tree-search attack - are constructed with a scoring configuration that names one **objective scorer** and, optionally, a list of **auxiliary scorers**. Only the objective scorer steers, and it must produce a true/false verdict, because the attack loop consumes exactly one bit per turn. One turn is: 1. the adversarial chat model composes the next attacker prompt; 2. any converters transform it; 3. the prompt target answers; 4. the objective scorer judges that answer. A true verdict breaks the loop and the attack result is recorded as achieved. A false verdict sends another turn, until the configured maximum number of turns is reached and the run ends unachieved. Auxiliary scorers annotate every response for the record and change nothing about control flow. So exactly two things end a run: **a positive verdict, or the turn budget**. Nothing else in the loop labels an attempt, which means every success number you will ever report traces back to one scorer configuration. ## What it costs With a model-backed scorer, a turn is three metered calls: attacker model, target, scorer. A campaign of 100 objectives with a budget of 8 turns is up to 800 turns and therefore up to ~2,400 model calls. The scorer's call is usually the cheapest of the three - a short rubric prompt in, a couple of hundred tokens of verdict and rationale out - so the money is small, often single-digit dollars for a campaign that size. The real bill is **wall-clock and attention**. The scorer sits on the critical path of every turn, so its latency (typically a second or a few) is added to every turn of every conversation, serialised; a cheap local scorer such as a substring check costs nothing and no time, and is coarse in exactly the way described next. Engineer time is the largest line of all: someone writes the rubric, labels transcripts and reads stop turns. ## Where the number misleads The headline of a PyRIT campaign is an **objective-achieved rate**, and that rate is a *scorer output*, not a measurement of the target. Because the verdict is control flow, a scorer error is not a reporting error - it is an experiment-design error committed at runtime, because it changes which data ever exists. - A **false positive** stops the conversation, so the turns that would have shown genuine compliance, or genuine resistance, were never generated; the run is filed as a success and its conspicuously short transcript is easy to skim past. - A **false negative** lets the run burn its full budget and files a real success as a miss - expensive and undercounted, but the compliance is still in memory and a corrected scorer can recover it. The rate also moves when nothing about the target moved: swap the judge model behind a self-ask scorer, or edit its rubric, and this month's number is not comparable with last month's. And "achieved" means one scorer believed one response met one objective; it is not a claim that the output was harmful, usable or reproducible. ## What you would check - Sort the runs by executed turns and read the response at the **stop turn** of the short ones, together with the stored rationale - that is the exact text the scorer acted on, and no other sample answers the question. - Then look at the other tail: a batch that always exhausts the budget looks identical whether the target resisted or the scorer is blind. - **Smoke-test the scorer** by handing it a handful of stored responses you already know are compliant and confirming it says so. - Finally, **re-score** the persisted conversations offline with a second scorer; that costs scorer calls only, since no target turns are repeated, and every disagreement is a transcript worth reading.

  • If the scorer never returns a positive verdict, what ends the run?
    The turn budget. The run terminates as exhausted, and from that fact alone you cannot tell a resistant target from a scorer that fails to recognise compliance - a batch that always burns the full budget is as often a scorer symptom as a target one.
  • Can you score a conversation after the run has finished?
    Yes, if the run's conversations were persisted. Re-scoring stored transcripts is the cheapest way to compare scorers, because no target turns are repeated. What it cannot recover are the turns a premature stop never generated.
  • Does the scorer decide what gets sent next?
    Not directly - the attack strategy composes the next turn. But the scorer decides whether there is a next turn at all, which is why its errors change the transcript, not just the label on it.

The scorer is the smoke alarm wired to the fire drill's stop button: when it sounds, the drill ends right there. A false alarm does not just log a wrong event - it means nobody ever finds out what the rest of the building would have done.

saying these in an interview costs you the question

  • Describes the scorer as a reporting step that runs after the attack finishes
  • Cannot say what terminates a multi-turn run
  • Treats a positive verdict as proof the objective was met without opening the transcript
  • Never asks what text the scorer was actually shown

context

open as a page

A PyRIT scorer can return a true/false verdict, a scaled numeric score, or a category label. How do you choose between them for a multi-turn attack run, and what extra decision does a scaled scorer force on you?

level: middleimportance: must knowfreq 62%

basics

~20 s

A true/false scorer gives the loop the yes-or-no it needs to stop, so it is the default for objective-driven runs. A scaled scorer returns a degree of compliance and forces you to pick the threshold that counts as success. A category scorer says what kind of harm, not whether the attack worked.

open as a page

Before you let a PyRIT scorer end multi-turn runs on its own, how do you check it against transcripts you labelled yourself, and which responses have to be in that labelled set?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Pull stored conversations from earlier runs, label the responses yourself, then re-score them offline with the candidate scorer and compare. Load the set with the hard cases: refusal-then-comply, partial compliance, in-fiction compliance, and confident nonsense. Only then let it stop runs. Replaying stored text costs no target turns.

open as a page

A batch of PyRIT multi-turn runs reports a high objective-achieved rate, but most transcripts stop on an early turn at a hedged, non-compliant answer. What is happening, and how do the two directions of scorer error differ in what they cost you?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The scorer is firing on responses that are not real compliance, and because a positive verdict ends the run, every false hit is also a truncated attack. False positives inflate the success rate and delete the turns you never sent. False negatives cost the whole turn budget and undercount. Read the stop-turn responses.

open as a page

Your team runs PyRIT campaigns against several internal endpoints on a recurring schedule. What should the scorer be allowed to decide on its own, and what has to reach a human before it counts?

level: principalimportance: should knowfreq 40%

basics

~20 s

Let the scorer decide only inside the run: when to stop a conversation and how to rank what gets read first. Anything leaving the team as a finding needs a human who read the transcript. Keep a fixed labelled transcript set as the scorer's own regression test, re-run whenever the scorer changes.

open as a page