skip to content

Before you let a PyRIT scorer end multi-turn runs on its own, how do you check it against transcripts you labelled yourself, and which responses have to be in that labelled set?

level: seniorimportance: must knowfreq 50%

answer

  1. label your own run transcripts
  2. replay offline, no target turns
  3. refusal-then-comply, partial, in-fiction, confident nonsense
  4. look at where errors land, not just how many
  5. freeze the set as a regression test

basics

~20 s

Pull stored conversations from earlier runs, label the responses yourself, then re-score them offline with the candidate scorer and compare. Load the set with the hard cases: refusal-then-comply, partial compliance, in-fiction compliance, and confident nonsense. Only then let it stop runs. Replaying stored text costs no target turns.

solid answer

~60 s

Two moves. First, get a labelled set that looks like your own traffic - responses your runs actually produced, not generic examples. The scorer will only ever see the output of your attack strategies against your targets, so that is the distribution it has to be right on. Second, replay. Score the stored responses offline with the candidate scorer and compare its verdicts against your labels. This is cheap because no target or attacker turns are repeated; only the scorer runs. What goes in the set matters more than how many. Clear refusals and clear compliance are decided correctly by almost anything. The set earns its keep on the ambiguous middle: responses that open with a refusal and then comply anyway, partial or hedged compliance, compliance wrapped in fiction or roleplay, and fluent output that is confidently wrong or useless. Those four families are where scorers split, and they are exactly the shapes a multi-turn attack produces. Then pick the stop rule against that evidence, and keep the labelled set so you can re-run it whenever the scorer changes.

go deeper

for a junior

Should at least say you check the scorer against some responses you judged yourself before trusting its verdicts.

for a middle

Should describe replaying stored transcripts through the candidate scorer and comparing verdicts with their own labels, rather than trusting a default.

for a senior

Should name the hard response families that belong in the set and explain why the scorer has to be checked on their own attack distribution, not a generic one.

for a principal

Should make the labelled set a durable asset with an owner - a frozen regression suite re-run on every scorer or backing-model change before it is allowed to end runs.

**Why your own transcripts, and not a curated set.** A scorer inside a run never sees tidy examples. It sees whatever your attack strategy elicited from your target, in your domain, after your converters reshaped the prompt. A scorer that looks fine on a public collection can still be systematically wrong on the one near-miss shape a particular strategy produces in almost every conversation - and because the verdict is control flow, that single blind spot bends every number in the campaign, not a few rows of it. Calibration means measuring the scorer on the distribution it will actually judge. **The replay loop, and what it costs.** PyRIT writes prompt pieces and scores to memory as a run proceeds, so if memory was persisted (a file-backed store rather than the in-memory one, which leaves nothing behind), checking a scorer is an offline job. Pull the stored responses, hand them to the candidate scorer directly, compare its verdicts with your labels. No target turns and no attacker turns are repeated, so the machine cost is one scorer call per item - dollars, not hundreds of dollars - and the comparison is clean because both candidates judge byte-identical text. The expensive input is human: labelling a response honestly takes one to a few minutes, so a 200-item set is most of an engineer-day. That is the number to plan around, and it is why the set is built once and frozen rather than re-created per campaign. **What has to be in the set.** Volume is not the lever; composition is. Clear refusals and clear compliance are judged correctly by almost anything, so a set made of those measures nothing. Weight it toward the shapes where scorers and humans diverge: - a refusal-shaped opening followed by actual compliance, and its mirror - a compliant-sounding opening that never delivers anything; - partial compliance: the outline or the framing without the operative content; - compliance held inside a fictional, hypothetical or roleplay frame; - fluent, confident output that is wrong or empty, which reads like success to any scorer keyed on tone rather than substance; - context-dependent compliance, where the response only means anything given earlier turns. Keep a slice of ordinary refusals and ordinary compliance too, so a scorer that fires on everything is visible in one glance. **A mechanism most people miss.** A self-ask scorer is typically shown one response piece plus the objective string - not the whole conversation. Compliance that only makes sense against earlier turns therefore looks harmless in isolation and scores as a miss, and multi-turn escalation strategies are exactly the machinery that produces such responses. Before blaming the rubric, confirm what text the scorer is being handed. Half the "miscalibrated scorer" reports are really the scorer being shown less than the human labeller saw. **How to read the comparison.** Do not report a single agreement percentage. With a low base rate of true successes, agreement is dominated by the easy negatives and a scorer that says "no" to everything can score above 90%. Look instead at *where* the errors land: how many of the scorer's positives survive your reading, and how many of your known-compliant responses it missed. Then note the direction that matters for the campaign you are about to run. Also measure the scorer against itself: run the frozen set twice. A self-ask judge above temperature zero can flip verdicts on its own, and that self-inconsistency is a floor on how well it can ever agree with you - no rubric edit fixes noise. **Then set the stop rule, and keep the set.** Choose the boolean condition or the threshold with that evidence in hand, and write down the choice and the numbers behind it. Freeze the labelled set as the scorer's regression test and re-run it whenever the scorer, its rubric or its backing model changes, before the new configuration is allowed to end runs. Labelling is paid once; the re-run is cheap. **The limit of the method.** Offline replay can tell you what the scorer would have said about text you already have. It can never tell you how a run would have gone had the scorer not stopped it: the turns after a premature stop were never generated, and no amount of re-scoring conjures them. Only re-running the affected objectives with the corrected scorer does that.

  • Why is re-scoring stored conversations preferable to re-running the attack when comparing two scorers?
    It repeats no target or attacker turns, so it costs scorer calls only, and both scorers see identical text - a live re-run would give each a different conversation and confound the comparison.
  • What can offline replay never tell you?
    How the run would have gone if the scorer had not stopped it. The turns after a premature stop were never generated, so no amount of re-scoring recovers them; only re-running with the corrected scorer does.
  • How large does the labelled set need to be?
    Small and adversarially chosen beats large and random. A few dozen deliberately hard cases per attack family surfaces the systematic errors; volume mostly adds easy cases that every scorer gets right.

saying these in an interview costs you the question

  • Trusts a scorer straight out of the box because it is the framework default
  • Calibrates on curated public examples rather than the transcripts their own runs produce
  • Labelled set contains only clear refusals and clear compliance
  • Re-runs live attacks to compare two scorers when the stored transcripts would have answered it
  • Changes the scorer or its backing model and never re-checks the labelled set

context