A finished PyRIT engagement is stored, and you now want to relabel it with a stricter judgment of what counts as a hit without spending another call against the target. What makes that possible, and what can relabelling stored transcripts not tell you?
answer
- generation costs, judgment does not
- relabel stored turns, no target calls
- verdict was a control signal
- unexplored branches do not exist
- old transcripts are dated evidence
basics
~20 sBecause the run stored every prompt and response, you can point a different scorer at the saved transcripts and relabel them without spending another target call. What that cannot tell you is what the attack would have done under the new judgment: the strategy branched on the old verdicts, so the turns it never explored simply do not exist.
solid answer
~60 sThe store decouples generation from judgment. Sending prompts costs money, time and exposure against a live system; judging text does not touch the target at all. So a durable run is re-labellable: you read the turns back, apply whichever scorer you now trust, and produce a second labelling of the same evidence. That is how you recover from a scorer that was too loose, compare two judgments on identical material, and build a labelled set from real transcripts. The limit is that only the labels are counterfactual, not the run. In a multi-turn strategy the verdict is a control signal: it decides whether to continue, to branch, or to stop. A stricter judgment applied afterwards would have kept a run going that actually stopped early, so the transcripts you are relabelling are a truncated sample chosen by the old scorer. Relabelling can turn old hits into non-hits honestly; it cannot discover the hits the run never reached. When the change is material, you re-run — and the store is also what tells you what the old run cost, so you can size that.
go deeper
Should grasp that the saved prompts and responses can be read back and judged again without contacting the target.
Explains the decoupling and names the obvious use — fixing a bad labelling cheaply — and that only labels change, not the conversations.
Articulates the selection bias from verdict-driven branching and early stopping, and reports relabelled numbers with the caveat attached.
Decides when a relabelling is sufficient and when the engagement must be re-run, and sets how such numbers are allowed to be stated to stakeholders.
### Why offline relabelling is possible at all A stored request piece is self-contained text: what was sent, what came back. A PyRIT scorer is a function over that text plus its own judge — it takes response pieces and emits score records, and it never touches the target to do so. Relabelling therefore means reading pieces back out of memory, running the new scorer over them, and writing the results with the memory instance's score-adding entry point. That is the whole mechanism, and the important detail is what it does **not** do: the new records are *appended alongside* the old ones, referencing the same pieces. Nothing is overwritten. A piece can end up carrying two verdicts from two scorers, and every later count has to filter by which scorer produced the score. ### What it costs, against what a re-run costs Take the same forty-objective, ten-turn engagement. Driven live, it bills three metered calls per turn — attacker model, target, scorer — roughly 1,200 calls, plus wall-clock, plus a fresh window of attack traffic against a production system and usually a fresh authorisation to send it. Relabelling the stored result pays only the judge's third: about 400 calls, minutes of wall-clock, no traffic to the target at all, no rate-limit risk on the customer's endpoint, and no scheduling conversation. That ratio is why the durable store is worth keeping even when you think the run is finished. ### What relabelling is genuinely good for Fixing a labelling mistake without buying the engagement twice. Running two candidate judges over identical transcripts so the comparison is not confounded by different conversations. Producing a human-labelled sample to calibrate a judge against, from real traffic against the real system rather than synthetic examples that lack the awkward partial-compliance cases. Answering an application team's challenge by showing the exact exchange behind a finding. ### Where the number misleads — the important part - **Selection.** In a multi-turn strategy the verdict is a *control signal*, not just an annotation: it decides whether to continue, branch, or stop. Early stopping on a false hit truncated conversations that would have gone further; a missed hit abandoned an objective. The surviving transcripts are therefore a sample chosen by the old scorer. Relabelling can honestly turn old hits into non-hits; it can never surface the hits the run never reached. - **Double counting.** Because re-scoring appends, a naive count over score records mixes two labellings of the same exchanges and can nearly double the apparent numerator. Filter by scorer identity, or by the timestamp window of the relabelling pass, and say which you used. - **Rate arithmetic.** Recomputing a success rate over relabelled transcripts and presenting it as though the run had been *driven* by the new judgment overstates it. The conversations were generated under the old one. - **Drift.** A transcript is evidence of what a system did on a date. Systems get patched and get guards put in front of them. A relabelled old run is not a current measurement, and presenting it as one is a common report defect. - **Missing context.** Anything a response depended on that was never stored — retrieved documents, tool results, server-side session state, a system prompt you could not see — is absent from the piece, so a later judge is scoring text whose provenance it cannot inspect. - **Judge calibration on a curated set.** Measuring a new scorer's agreement with humans on transcripts the old scorer selected flatters it. Its behaviour on a run it actually drives is not what you measured. ### What to check Filter the score records by scorer before you count anything, and report the count with the filter stated. Hand-label a random sample — fifty exchanges is usually enough to see a problem — and report the new scorer's agreement on it rather than asserting that it is stricter. Look at why conversations ended: a high proportion terminating at the strategy's stop condition rather than exhausting the turn budget is the signature of selection bias, and is the number that tells you whether relabelling is defensible or whether you owe a re-run. Then state, in the report itself, the run date, the target deployment, which labelling produced the reported figures, and whether the conversations were generated under that same labelling. When those two differ, that is a caveat in the body, not a footnote — and if the conclusion turns on the number, the honest next step is a fresh run driven by the new judgment, sized from what the store says the old one consumed.
- Your stricter relabelling drops the hit count by half. What do you report?The stricter count as the finding, with the run date, a note that the conversations were generated under the looser judgment, and a recommendation to re-run if the conclusion turns on the number.
- Why can relabelling not raise coverage?Coverage is a property of what was sent. Judgment only relabels what came back, so no relabelling reaches a prompt or a turn the run never issued.
- What makes stored transcripts better than synthetic examples for checking a judge?They are real responses from the real target under real attack pressure, including the awkward partial-compliance cases that synthetic examples rarely produce.
Relabelling a stored run is like re-marking an exam with a stricter key: you can lower the score on every answer that was written, but you cannot mark the questions the student was never asked. The old marking scheme is what decided which questions got asked at all.
saying these in an interview costs you the question
- Reporting a relabelled success rate as if the run had been driven by the new judgment.
- Believing relabelling can surface objectives the original run abandoned.
- Treating an old stored run as a current measurement of a system that has since changed.
- Ignoring that responses may have depended on context that was never stored.