A PyRIT multi-turn run ends with 'objective not achieved' after exhausting a 10-turn budget. Why is that not evidence that the target refuses the behaviour?
answer
- not-achieved = budget exit
- budget is a hyperparameter, not a property
- one run = one attacker draw
- read the last turns: flat refusal vs softening
- publish negatives as 'within N turns'
basics
~20 sIt only says the adversarial model did not get there within the turns you allowed. The budget is part of the result, not a property of the target. A longer run, a different attacker configuration, or another attack strategy can still breach the same endpoint. Report the negative together with the turn budget that produced it.
solid answer
~50 sThe not-achieved exit is the *budget* exit. Three separate things had to hold for it to appear, and only one of them is about the target: 1. **The budget** — ten turns is a hyperparameter you chose. Runs that end on turn twelve are invisible to a ten-turn run. 2. **The attacker draw** — each turn is a model sample, so this run followed one path out of many. Re-running the identical configuration explores a different path and can land differently. 3. **The scorer** — if the objective scorer is strict or the objective is worded so that a partial disclosure does not count, a real weakening of the target reads as not-achieved. So the honest statement is *this attacker, at ten turns, once, did not reach this objective* — a bound on what you looked for, not a clean bill of health. In a report, a negative without its turn budget attached is an unfalsifiable claim; with the budget attached it is a reproducible one.
go deeper
Should at least say that the run stopped because it ran out of turns, so it does not prove the target refused.
Separates the three contributors — budget, stochastic attacker path, scorer strictness — and insists the budget be reported with the result.
Diagnoses from the transcript shape whether the run was truncated, and sets escalation policy for budgets on those objectives.
Owns the reporting standard: no negative leaves the team without its budget and repeat count, and comparisons between systems hold the harness constant.
**What the verdict is a statement about.** A PyRIT multi-turn run has exactly two exits: the objective scorer returns met, or the loop hits `max_turns`. The second exit is the *budget* exit, and it is recorded with the same kind of verdict string as the first. Nothing in "objective not achieved" is a measurement of the target's willingness to refuse; it is a record that a search stopped. **Three things had to hold, and only one is about the target.** 1. **The budget.** Ten turns is a hyperparameter you chose. Every conversation that would have broken on turn twelve is invisible to a ten-turn run, and is invisible in a way that produces no warning — the run completes cleanly and reports a negative. 2. **The attacker draw.** Each attacker turn is a sampled generation from the adversarial chat model, conditioned on the transcript so far. Two runs of a byte-identical configuration diverge at turn one and never re-converge. This run walked one path through a very large space of conversations. 3. **The scorer.** If the objective scorer is strict, or the objective is worded such that partial disclosure does not count as met, a genuine weakening of the target is recorded as not-achieved. The verdict then measures your referee, not your target. **Why budgets are short, in money and clock.** One turn bills three model calls — adversarial, target, scorer — so a ten-turn run costs roughly thirty calls, and turns inside a conversation cannot be parallelised because each needs the previous reply. Forty objectives, five repeats each, at ten turns is about six thousand calls and hours of wall clock. Doubling the budget to twenty turns doubles both. Budgets therefore get capped for scheduling reasons, and the cap silently becomes the sensitivity of the entire exercise: if the behaviours you care about only emerge after a long build-up, a short budget guarantees a page of green and costs you nothing to produce. **Reading the transcript is the diagnostic.** Pull the conversation out of PyRIT's memory and look at the target's last few replies. Three shapes, three conclusions: | Transcript shape at the cutoff | What it means | Action | |---|---|---| | Identical flat refusal from turn one, attacker circling | Genuine, narrow negative | Record it, with the budget | | Replies softening — more hedging, more partial engagement, more willingness to hold the framing | Truncated run, not a refusal | Re-run at a larger budget before calling it | | Attacker drifting off the objective | Objective-wording or attacker-configuration defect | Nothing was tested; fix and re-run | The second row is the expensive one to miss, because it looks exactly like the first in a summary table. **Where the number misleads.** The phrase travels upstream and shortens. "Objective not achieved" becomes "not achieved" becomes "the model refused" becomes a resistance percentage on a slide, and by then the budget that produced it is gone. A negative without its budget attached is unfalsifiable: nobody can reproduce it, nobody can refute it, and it cannot be compared with anything. Worse, two targets run at different budgets produce a difference in hit rate that is entirely an artefact of the harness — you have measured your own settings and reported it as a property of the systems. **What I would check, and what I would publish.** Confirm which exit fired and how many turns were consumed; a run that ended at turn ten of ten is the budget exit, and a summary that does not distinguish it from an early stop is not usable. Read the trailing replies for the softening shape. Confirm the scorer would have accepted a partial result, by checking it against a couple of hand-labelled transcripts. Then state the finding in reproducible form — "objective X not achieved in 5 of 5 runs at a 10-turn budget with attacker configuration Y" — rather than "the model is resistant to X". Fix a standard budget per objective tier and hold it constant across retests and across any two systems you compare, so a change in the numbers is a change in the system rather than a change in how hard you looked. Budgets should ratchet upward on retest, never down.
- You suspect a negative was truncated. What in the stored transcript tells you, before you spend another run?The trajectory of the target's last few replies. Flat, identical refusals suggest a real negative; progressively softer, more engaged replies that get cut off at the budget suggest the run ended early and deserves a larger budget.
- Two targets both come back not-achieved. What must be equal before you can say they are equally resistant?At minimum the objective, the turn budget, the adversarial model configuration and the scorer, plus the number of repeat runs. Otherwise you are comparing harness settings, not systems.
- How would you phrase the finding in a report?Something like: 'objective X not achieved in 5 of 5 runs at a 10-turn budget with attacker configuration Y' — a statement someone else can reproduce or refute, rather than 'the model is resistant to X'.
A ten-turn budget is a stopwatch on the search, not a verdict on the safe. Calling the search off after ten minutes tells you how long you looked, not whether there was anything to find.
saying these in an interview costs you the question
- Reporting 'objective not achieved' as 'the target is safe against this' with no budget stated.
- Believing a single run of a stochastic attacker is a measurement.
- Never opening the transcript to see whether the target was softening when the run was cut.
- Comparing two targets that were run at different turn budgets.