In PyRIT, what does a scorer do during an attack run, and what must a custom scorer you write return for the run to be able to use it?
answer
- scorer = the run's stop signal
- score object, not bare True
- verdict + rationale + piece id
- boolean family vs graded family
- persisted to memory, replayable
basics
~20 sA PyRIT scorer reads the target's response and decides whether it satisfies the run's objective. A custom scorer must return what the built-in ones return: a verdict, boolean or a normalized number, plus a rationale and the identifier of the piece it scored, so the attack strategy can branch and memory can store it.
solid answer
~60 sIn PyRIT the scorer is the object that decides whether a response counts as objective-achieved. It is not decoration on the report; it is the run's control signal. A multi-turn attack strategy asks the scorer after each target reply and uses the answer to decide whether to send another turn or to stop and declare success. So a custom scorer has to satisfy the same contract as the shipped ones: subclass the scorer base class in the version you have installed, implement its scoring method, and return score objects — not a bare `True`. Each score carries the verdict (a true/false value, or a value on a normalized scale, depending on which scorer family you are extending), the rationale text explaining the verdict, a category, and a reference to the prompt piece it scored so memory can join the score back to the transcript. The rationale is the part people skip and then regret: without it, a triager reading the run later sees a hit with no reason, and cannot tell a real finding from a scorer bug.
go deeper
Says a scorer decides whether the target's response met the objective, and that a custom one must return the framework's score object with a verdict and a rationale.
Adds that the verdict is the multi-turn stop condition, and distinguishes the boolean family from the graded family and when each is appropriate.
Talks about memory persistence, replaying the scorer over stored transcripts, and handling responses the scorer cannot parse instead of returning a silent false.
Frames the scorer as the run's control plane and the report's evidence chain at once, and sets a team rule that no custom scorer ships without a rationale field and a defined behaviour for unscoreable input.
### Where the scorer sits in a run PyRIT drives an attack as a loop around three objects you configure: an *attack strategy* (the algorithm deciding what to send next), a *prompt target* (the adapter wrapping the system under test), and a *scorer*. Each iteration builds a prompt, passes it through any attached *prompt converters* (objects that transform the outgoing text), sends it to the target, writes the response into *memory* (PyRIT's persistence layer — a local SQLite file by default, a shared database on a real engagement), and then hands that response to the scorer with exactly one question: does this satisfy the objective? Two things consume the answer, which is why a scorer is not report furniture. **The attack strategy.** A multi-turn strategy loops until the scorer reports objective-achieved or the turn budget is exhausted. Your scorer is the loop's exit condition — it decides when the run stops, and in strategies that condition the next prompt on the last verdict, it also shapes what gets sent next. **Memory and the deliverable.** Scores are persisted alongside the prompt pieces and keyed to the piece that was scored. That key is what lets you re-open a conversation months later, join a verdict to the exact response that earned it, and re-score stored transcripts offline without touching the live target again. ### The contract a custom scorer must satisfy Subclass the scorer base class present in your installed version, implement its scoring entry point, and return a *list* of score objects — never a bare `True`. A list, because one response can attract several verdicts in different categories. Each score object carries four things that matter downstream: the value (a true/false verdict, or a number on a normalised scale, depending on which scorer family you extend), the rationale text, the category, and the identifier of the prompt piece scored. Pick the family deliberately. A boolean verdict is what a stop condition wants; a graded value is what a trend across runs wants. Returning a float and letting the caller decide that 0.7 means yes hides the threshold outside the scorer, so nobody can reconstruct the hit count from the stored scores alone. ### What it costs A deterministic scorer costs effectively nothing per turn — it is a function over a string. A model-backed one is a third metered call on every turn, alongside the attacker model and the target: a 50-conversation run with an 8-turn budget is up to 400 scored turns, so 400 extra calls and 400 extra round-trips in a loop that runs serially per conversation. The engineering cost is the line teams under-budget: writing the scorer is an afternoon, while establishing that it agrees with a human is a day of hand-labelling — and that day is the one that makes the number reportable. ### Where the number misleads The dangerous failure is silent. A scorer that returns an empty list, or returns false because it was handed something it cannot parse — an image, an audio response, an error string, an empty completion — produces a run with zero objective-achieved. Zero reads in a report as *the target held*. That is a different claim from *we looked and found nothing*, and nothing in the pipeline separates them, because no human ever opens a non-hit. The second misread is a hit with no evidence. A bare boolean, with no rationale and no piece reference, hands a triager a count with nothing to check, so a scorer bug and a real finding look identical in the deliverable. ### What to check before you trust it - Run three turns and read the scores table in memory directly. Confirm rows exist, that the piece reference resolves to the response you expect, and that the rationale is populated. - Feed the scorer one hand-labelled known hit and one hand-labelled known refusal. A scorer that has never been shown a true positive is untested, however clean the code reads. - Confirm it is actually attached to the strategy. An unattached scorer and a broken one look identical from outside. - Check class and field names against the installed package rather than a write-up. These have been renamed across PyRIT releases. - Define the behaviour for input it cannot score, and make it visible: an unscored turn you can count, not a false verdict you cannot.
- Why does a PyRIT score reference the prompt piece it scored rather than just the conversation?One conversation holds many pieces and one response can get several scores. The piece reference is what lets memory join a verdict back to the exact response, so a later re-score or a triage read lands on the right turn.
- Your custom scorer returns an empty list for every response. What does the multi-turn strategy do?It never sees objective-achieved, so it keeps sending turns until the turn budget is exhausted and reports the run as a failure. The symptom is a run that always burns its full budget and always finds nothing.
saying these in an interview costs you the question
- Describing the scorer as report metadata rather than the thing that ends a multi-turn run.
- Returning a bare boolean with no rationale and no reference to the scored piece.
- Assuming a specific class or field name from a tutorial without checking the installed package, since these names have been renamed across releases.
- Silently returning false when the scorer was handed content it cannot parse.