Mid-engagement you fix a bug in your custom PyRIT scorer. You re-score last week's stored transcripts instead of re-running the attacks. What does that recover, and what does it not?
answer
- pieces and scores stored separately
- single-turn repairs fully
- multi-turn: paths not taken
- re-run where control flow changes
- stamp verdicts with scorer identity
basics
~20 sRe-scoring recovers verdicts on responses you already collected: single-turn runs are fully repaired and missed hits reappear. It cannot recover the run itself. In multi-turn attacks the old scorer chose when to escalate and when to stop, so the transcripts hold only the paths it steered into; turns never taken stay unexplored.
solid answer
~1 minRe-scoring is cheap and it is the right first move, because the target calls are already paid for and persisted. What it genuinely fixes: - **Single-turn runs.** Each response is independent of the verdict, so a corrected scorer over the stored responses gives exactly the number a correct scorer would have produced originally. - **The evidence trail.** You keep both verdicts against the same pieces, which lets you say precisely how many findings the bug cost. What it cannot fix is the counterfactual. In a multi-turn strategy the scorer is the loop condition and often shapes the next prompt, so the corpus you are re-scoring was produced under the buggy scorer. Runs that should have stopped at turn 3 kept going; runs that should have kept going stopped. Turns that a correct scorer would have caused simply do not exist in memory, and no amount of re-scoring conjures them. The practical protocol: re-score everything, then re-run the subset where the corrected verdict would have changed the loop's behaviour — the runs with a newly-found hit before the last turn, and the runs that ended early on a verdict the fix reverses. Report which numbers came from re-scoring and which from a re-run.
go deeper
Knows the responses are stored and can be scored again without hitting the target.
Distinguishes single-turn runs, which re-score cleanly, from multi-turn runs, where the verdict drove the conversation.
Prioritises re-runs by whether the corrected verdict changes control flow, and keeps both verdicts stamped with the scorer identity.
Treats it as an incident with a blast radius set by how much control the scorer holds, sets the reporting rule for mixed-version figures, and pushes calibration onto a pilot run so an engagement is never spent behind an unvalidated scorer.
### Why re-scoring is possible at all PyRIT persists prompt pieces and scores as separate records joined by the piece identifier. That separation is the whole reason a scorer bug is a recoverable incident rather than a lost week: verdicts are a layer over stored evidence, not something baked into it. A corrected scorer can be run over a week of stored conversations offline, without a single new call to the target. ### What re-scoring genuinely recovers - **Single-turn and dataset-style runs, completely.** Each response was produced without reference to any verdict, so relabelling the stored responses reconstructs exactly what a correct scorer would have produced on the day. - **The per-turn truth of the conversations you happen to hold.** In multi-turn runs you learn the correct verdict for every stored response, and missed hits reappear as findings. - **The size of the bug.** Keep both verdict sets and you can state precisely how many findings the defect cost — a far better answer to a client than a quiet correction. ### What it cannot recover The counterfactual. In a multi-turn strategy the scorer is the loop condition and often shapes the next prompt, so the corpus you are re-scoring was *produced under the buggy scorer*: - **Paths not taken.** Where the next prompt depended on the last verdict, the entire downstream conversation is a function of the bug. Re-scoring relabels a branch a correct scorer would never have walked. - **Early stops.** A false hit ended a conversation at turn 2. The remaining eight turns were never sent; the corrected verdict only tells you the run should have continued. - **Late stops.** A missed hit burned the budget. Here the finding does come back — the happy case — but the extra turns are money already spent. ### Deciding what to re-run Rank by whether the corrected verdict changes control flow, not by whether it changes a label: 1. Conversations where the fix reverses the stop decision. Counterfactual from that turn onward; only a fresh run produces the turns that should have existed. 2. Conversations where the fix flips an intermediate verdict the strategy branched on. 3. Conversations where only the final label changes. Re-scoring settles these completely and a re-run buys nothing. ### What each option costs Re-scoring 400 stored responses pays at most one scoring call each and zero target calls — minutes of wall clock, and nothing at all if the corrected scorer is deterministic. A re-run pays attacker, target and scorer on every turn again: fifteen conversations at an 8-turn budget is roughly 360 metered calls, plus target-side rate limits, plus fresh hours inside the engagement's authorised testing window. On most engagements that window, not the bill, is the binding constraint — which is exactly why ranking by control flow matters, since it keeps the re-run set to the conversations that need it. ### Where the number misleads A hit count assembled from two scorer versions and presented as one figure. If seven of nineteen findings came from re-scored history under the fixed scorer and twelve from live runs under the buggy one, nobody reproducing the engagement arrives at nineteen with either version. Stamp every stored score with the scorer identity that produced it, never overwrite the old verdicts, and label each reported figure as re-scored history or fresh run. The second misread is subtler. After a fix the re-scored count usually rises, and it is tempting to present the rise as new attack progress. It is not: it is the same evidence read correctly. Say so explicitly, or the client reads a scoring repair as a deteriorating security posture. ### What to check - Confirm the corrected scorer is deterministic over the stored corpus before trusting any diff between versions. If it is model-backed, part of the delta is variance, not the fix, and you need the flip rate first. - Diff the two verdict sets per conversation and classify each change by whether it would have altered control flow. That classification *is* the re-run list. - Check that the fix did not simply move the error. A scorer loosened until it catches the misses needs its false-alarm rate re-measured on the hits. - Draw the standing lesson: the blast radius of a scorer bug is proportional to how much control the scorer holds. That is the argument for calibrating a bespoke scorer on a small pilot run before an engagement's budget is spent behind it.
- Which runs from the affected week do you re-run first?The ones where the corrected verdict reverses the stop decision, since everything after that turn is counterfactual. Runs where only the final label changes need no re-run at all.
- How do you keep the report honest when figures come from two scorer versions?Stamp each stored score with the scorer identity, keep both verdicts rather than overwriting, and label each reported figure as re-scored history or fresh run with its scorer version.
Re-scoring a multi-turn run is like re-marking an exam that adapted its questions to the marker's earlier decisions. You can correct every mark on the questions the candidate actually saw, but the questions a correct marking would have led the exam to ask were never asked, and no amount of re-marking creates them.
saying these in an interview costs you the question
- Claiming a re-score fully repairs a multi-turn engagement.
- Overwriting the old verdicts so the cost of the bug becomes unmeasurable.
- Mixing re-scored and freshly-run figures in one number without saying so.
- Re-running everything at full cost when only the control-flow-changing subset needs it.
- Not drawing the lesson that a bespoke scorer should be calibrated on a pilot before an engagement is spent behind it.