A PyRIT scorer can return a true/false verdict, a scaled numeric score, or a category label. How do you choose between them for a multi-turn attack run, and what extra decision does a scaled scorer force on you?
answer
- the loop needs a boolean
- threshold is the hidden parameter
- keep the raw score, re-threshold later
- category = triage axis, not stop condition
- cheap in the loop, rich afterwards
basics
~20 sA true/false scorer gives the loop the yes-or-no it needs to stop, so it is the default for objective-driven runs. A scaled scorer returns a degree of compliance and forces you to pick the threshold that counts as success. A category scorer says what kind of harm, not whether the attack worked.
solid answer
~60 sThe run loop needs a boolean. Whatever kind of scorer you attach, something has to reduce its output to "stop" or "keep going", so the choice is really about where that reduction happens and what you keep alongside it. A true/false scorer does the reduction inside the scorer. It is the simplest thing to wire into a multi-turn attack, and it throws away the evidence of how close the near-misses were. A scaled scorer returns a number on a fixed range and pushes the decision onto you: the threshold at which a response counts as objective-achieved. That threshold is a real experimental parameter - move it down and the success rate rises while runs stop earlier and shallower; move it up and runs go deeper and you undercount. Its benefit is that stored scores can be re-thresholded later without re-running anything. A category scorer answers "what kind of content is this", which is a triage and reporting axis, not a stop condition. Use it to label conversations you already stopped, not to decide when to stop.
go deeper
Should know the three shapes exist and that the true/false one is what a run loop consumes directly.
Should explain that a scaled score has to be thresholded to drive the loop, and that the threshold moves both the success rate and how deep runs go.
Should push the expensive judgment off the hot path - cheap stop condition in the loop, richer re-scoring of stored transcripts afterwards - and warn that a scaled score is not a probability.
Should treat the threshold as a program-level parameter that has to be written down and justified, because it silently defines what the organisation calls a success.
**Start from what the loop can consume.** A PyRIT objective-driven multi-turn attack continues until something answers one binary question: was the objective achieved? Every scorer shape therefore gets collapsed to a boolean somewhere - inside a true/false scorer, at a threshold applied over a float score, or in a mapping you write over category labels. Choosing a shape is really choosing *where that collapse happens* and *what evidence you keep beside it*. **True/false.** A true/false scorer (in PyRIT, the self-ask true/false family, or a cheap deterministic check such as a substring scorer) answers a question you phrase yourself and returns `"True"`/`"False"` in the `Score.score_value`, with a rationale. It plugs straight into the attack's objective-scorer slot with no adapter. Cost: one model call per scored response for a self-ask scorer, effectively zero for a substring or local-classifier one. What you lose is gradation: a run that stopped and a run that came within a hair of stopping are indistinguishable afterwards if the boolean is all you kept. **Float scale.** A scaled scorer (the Likert-style self-ask scorers, or a hosted content-filter scorer returning per-category severities) returns a normalised number. PyRIT will not let that drive the loop directly; you wrap it in a threshold scorer, which converts the float to true/false at a cut-off you supply. That cut-off is a real experimental parameter and it is where most of the trouble lives. Three specific traps: | trap | what actually happens | |---|---| | granularity | A 1-5 Likert rubric normalises onto five reachable values only. Thresholds that fall between two adjacent steps behave identically, so "tuning" from 0.55 to 0.65 can change nothing at all. | | meaning | The float is not a probability that the response is harmful. It is whatever the rubric or the vendor's severity scale emits, and its middle is precisely where it disagrees with humans most. | | direction | Lowering the threshold raises the success rate *and shortens the runs*, because conversations stop earlier. The rate goes up while the evidence behind it gets thinner. | The compensating benefit is real: if the float is persisted in memory alongside the verdict, you can re-threshold the whole corpus offline for nothing. If you kept only the boolean, re-thresholding means re-scoring, which is a fresh model call per item - and a self-ask judge above temperature zero is not guaranteed to reproduce its own earlier score. **Category / classifier.** A category scorer answers a different question: which harm family this content belongs to. That is a triage and routing axis. Using it as the stop condition means encoding "any label except the benign one counts as achieved", which is usually wrong, because a refusal that explains *why* it is refusing is still about the topic and will be labelled accordingly. Content-filter scorers also carry a second mismatch: their categories are the vendor's taxonomy, not your objective, so a policy-relevant success can score zero on every category the service knows about. **The arrangement that usually wins.** Drive the loop with something cheap and deliberately tuned - a true/false scorer or a thresholded float - and enrich afterwards. Attach the richer scorers as auxiliary scorers, or better, run them offline over the persisted conversations for the record and for grouping. Scoring stored text repeats no target and no attacker turns, so the expensive judgment never has to sit on the hot path of every turn of every conversation. Composite scorers let you require agreement between two cheap checks before stopping, which is often a better use of a second call than one expensive judge. **What you would check.** Plot the distribution of the raw float over a few hundred stored responses before you trust any threshold. If the mass sits at the two ends, your "scale" is a boolean that cost you a graded scorer's price. If a large mass sits within one step of your cut-off, the reported rate is fragile and small rubric or model changes will move it. Then take a sample straddling the threshold, label it yourself, and confirm the cut-off separates your labels rather than the scorer's habits. Finally, confirm the raw score is actually being persisted - the day you need to re-threshold is not the day to discover only the boolean was kept.
- You lowered the threshold on a scaled scorer and the success rate doubled. What else changed?The depth of the runs. Runs now stop earlier, so the transcripts are shorter and the later turns that would have tested real compliance were never sent. The rate rose and the evidence behind it got thinner.
- When is a boolean scorer clearly the right call?When the objective is a crisp, checkable event - a specific string, format or action appears in the response - and cheapness per turn matters because the run is long. Grading adds nothing when the target condition is discrete.
- Where should a category scorer run?Over persisted transcripts after the run, to group and route what you already decided was a hit. Running it in the loop pays for a label on every turn you were going to discard anyway.
A threshold is a dial on a metal detector, not a fact about the ground. Turn the sensitivity up and you find more, faster, and stop digging sooner - which is exactly why the pile of finds says as much about the dial as about the field.
saying these in an interview costs you the question
- Picks a scaled scorer and never states the threshold that counts as success
- Reads a scaled score as a probability that the response is harmful
- Uses a harm-category label as the objective-achieved condition
- Persists only the boolean, so nothing can be re-thresholded later