A multi-turn PyRIT attack run finishes with the objective not achieved. What are the ways that loop can terminate, and why is that outcome not evidence that the target is safe?
answer
- achieved / cap exhausted / error / cancelled
- false negative leaves no trace
- attacker-side refusal invalidates the run
- turns used vs cap tells you which
- hand re-score a sample of misses
basics
~20 sIt ends three ways: the scorer returns an achieved verdict and the loop stops early, the turn cap is exhausted, or something raised — target error, rate limiting, the attacker side declining to continue. Not-achieved conflates all of the non-success endings, so it says the run stopped, not that the target held.
solid answer
~60 s**Terminations.** An achieved verdict from the configured scorer ends the loop early. Otherwise the turn cap runs out. Otherwise something breaks: the target errors or rate-limits, a converter or the attacker endpoint fails, the operator cancels. **Why not-achieved is weak evidence.** Several very different situations collapse into the same word: - The scorer missed a real success — a miss never stops the loop, so it looks like turns simply ran out. - The turn cap was set below the number of turns this objective needs. - The attacker-side model applied its own safety behaviour and stopped producing useful next prompts, so the transcript is polite and empty. - The target surface does not carry history, so the loop was really repeated one-shots. - The objective and the scorer's criteria never described the same thing. **How I confirm.** Read the stored transcript, not the verdict. Check turns used against the cap: a run that stopped at the cap and one that stopped at turn two mean different things. Look for refusals originating on the attacker side. Re-score the transcript by hand on a sample.
go deeper
Knows the run ends on success or when it runs out of turns, and that a failure to succeed does not prove much.
Lists the endings including errors, and explains that a scorer miss keeps the loop going so it is indistinguishable from an honest exhaustion.
Runs the triage: turns used against cap, who refused in the transcript, whether history was carried, a hand-scored sample, and repeats for non-determinism. Refuses to write a robustness claim from verdicts alone.
Sets the standard that a report states endings and denominators rather than a pass count, and requires a measured scorer miss rate before any negative result is published.
### The endings A multi-turn PyRIT run has exactly one success ending and a family of non-success endings, and the summary result flattens the family into a single word. 1. **Achieved.** The configured scorer returned a positive verdict on some turn's response and the loop stopped there — possibly at turn two of a cap of fifteen. 2. **Turn cap exhausted.** The loop spent its full budget without a positive verdict. 3. **Error.** The target timed out, rate-limited or rejected at the transport level; the attacker-side model endpoint failed; a converter raised; credentials expired mid-run; the memory store was unwritable. 4. **Cancelled.** An operator or an outer harness stopped the run. Only the first is informative from the summary alone. The other three need the transcript. ### Why the negative is soft evidence Genuinely different situations collapse into “not achieved”: - **The scorer missed a real success.** The response satisfied the objective; the rubric did not recognise it. A miss never stops the loop, so the run simply continues and exhausts. - **The cap was below what the objective needs.** Nothing distinguishes “ran out of turns while progressing” from “ran out of turns while getting nowhere” without reading the exchange. - **The attacker side declined.** The model composing the next prompts applied its own safety behaviour, the transcript is polite and empty, and the target was never pressed. - **The target adapter carried no history.** The loop ran; each send arrived context-free. Repeated one-shots at multi-turn prices. - **Objective and scorer criteria were never the same thing.** The run chased one goal and judged another. The asymmetry is the point. A scorer **false positive** ends the run early and leaves a conspicuously short transcript in a batch — you notice it, you open it, you catch it. A scorer **false negative** leaves no trace at all: the loop keeps going and then exhausts, which is exactly what an honest failure looks like. The error that inflates a safety claim is precisely the error that is invisible in the summary. Any process that reads verdict fields and never opens transcripts is therefore *systematically* biased toward reporting the target as robust. ### What it costs to find out Triage is engineer time before it is call budget. Pulling turns-used and error counts per run is a query over the memory store — minutes. Hand-labelling a sample of not-achieved transcripts is the expensive part: an experienced reader gets through a few dozen transcripts in an hour or two, and that is the only way to put a number on the scorer's miss rate. Re-running costs the same three calls a turn as the original, so a ten-percent repeat reserve on a deep campaign is a real line item. All of it is cheaper than publishing a robustness claim that the customer's own team disproves in a week. ### The triage I would actually run - **Turns used against the cap, per run.** Runs bunched at the cap say the cap is binding. Runs bunched well short of the cap say something errored or somebody refused — honest exhaustion pushes runs to the cap. - **Who produced the refusals**, the target or the attacker-side model. Attacker-side refusals invalidate the run as a test of the target entirely. - **Whether the target adapter carried history at all**, which decides whether “multi-turn” describes what happened. - **A hand-scored sample of not-achieved transcripts**, to estimate the miss rate before any negative reaches a report. - **Repeats.** The endpoint and its guardrail are both stochastic. One not-achieved is one draw, not a property of the system. ### What the report should say Not “the model resisted N objectives”. Rather: N objectives were run under this strategy shape at this turn cap; M ended at the cap, K ended in errors, J were cancelled; a sample of S cap-exhausted transcripts was re-scored by hand and the measured miss rate was X. The endings and the denominator are the finding. A verdict count on its own is a number whose meaning nobody downstream can reconstruct — and it will be reconstructed, wrongly, as safety.
- Every run in a batch stopped well before the turn cap with no success. What is your first hypothesis?An error path or an attacker-side model that declined to continue. Real exhaustion should push runs to the cap, so uniformly short runs point at something breaking or refusing rather than at a resistant target.
- Why is a scorer false positive easier to catch than a false negative in this loop?A false positive stops the run early and leaves a short transcript that stands out and invites inspection. A miss just lets the loop continue to exhaustion, so it looks exactly like an honest failure.
- What would you put in the report instead of a count of objectives the model resisted?The strategy shape and turn cap used, how many runs ended at the cap versus in errors, and the hand-checked miss rate on a sample of the not-achieved transcripts.
A scorer's false positive is a smoke alarm going off with no fire: loud, and somebody comes to look. A false negative is the alarm that stays silent, and silence is exactly what a safe building sounds like.
saying these in an interview costs you the question
- Treats not-achieved as a pass and writes it up as model robustness.
- Cannot name the turn cap as a termination cause.
- Never opens the transcript, working only from the verdict field.
- Misses that refusals in the transcript may have come from the attacker-side model rather than the target.
- Reports a single run per objective on a non-deterministic endpoint as a result.