skip to content

A PyRIT multi-turn run halts the moment its objective scorer marks a turn as success. What does that early stop leave out of your engagement report, and how would you run it differently?

level: seniorimportance: should knowfreq 42%

answer

  1. success exit = existence proof only
  2. hits out of N, not a flag
  3. turns-to-success as a difficulty proxy
  4. first accepted response is the floor, not the ceiling
  5. hold budget and N constant to compare

basics

~20 s

Stopping at first success gives an existence proof and nothing else: no severity, no repeatability, no sense of whether it took two turns or nine. One run is a single sample of a stochastic search. Repeat each objective several times, record turns-to-success across runs, and hold the budget fixed across targets you compare.

solid answer

~50 s

The success exit answers *can it happen* and refuses to answer anything else. - **Repeatability is missing.** The attacker turn is a model sample, so a one-off hit may reproduce nine times in ten or one in ten. Those are different findings and a defender will price them differently. - **Effort is missing.** Turns-to-success is the closest thing the loop gives you to a difficulty metric, and a single run gives you one draw of it. - **Depth is missing.** The run stopped at the first response the scorer accepted, which is usually the weakest form of the breach. Whether the target would have gone further is unmeasured. So run each objective N times, keep N and the turn budget constant, and report the hit fraction plus the distribution of turns-to-success rather than a single flag. Where depth matters, follow up with a separate run that continues from the successful transcript instead of trying to make one loop do both jobs — that keeps the success statistic clean.

go deeper

for a junior

Recognises the run stops on the first success and that this is one attempt, not a rate.

for a middle

Adds that the attacker is stochastic, so a single hit or miss is a sample, and suggests repeating the run.

for a senior

Specifies the actual protocol — fixed budget, fixed attacker configuration, N repeats, hits/N plus turns-to-success — and separates depth measurement from the success statistic.

for a principal

Sets the tiering: cheap low-N screening for breadth, high-N confirmation for anything that hit, and a house rule that no comparison crosses differing budgets or N.

**Why one run is not a result.** The success exit fires the moment the objective scorer accepts a target reply, and the loop returns immediately. That answers exactly one question — *can this happen at all* — and the answer is genuinely valuable: an existence proof against a live target is the thing a red team is for. But the attacker's turns are sampled generations from the adversarial chat model, so two runs of a byte-identical configuration diverge on the first turn and never re-converge. Each run is one draw from a distribution of attack paths, and the loop's output collapses that distribution to a single bit. A report built on single runs cannot distinguish a fragile one-in-ten fluke from an every-time breach, and that distinction is most of what a defender needs in order to triage. **Three things the early stop leaves out.** - **Repeatability.** Nothing in a single hit tells you the hit fraction. Reliability drives severity and drives whether a mitigation can be validated at all. - **Effort.** Turns-to-success is the closest thing the loop offers to a difficulty metric — an objective that lands on turn two is a different risk from one that needs eleven turns of build-up. One run gives you one draw of that number. - **Depth.** The run halted at the first response the scorer would accept, which is by construction the *weakest* form of the breach — the threshold of the objective, not its ceiling. Whether the target would have gone further, in specificity or in volume, is simply unmeasured. **What to record instead of a flag.** For each objective, at a fixed turn budget and fixed attacker configuration, run N times and keep: hits out of N; the turns-to-success values for the runs that hit; for the misses, whether the transcript ended on a flat refusal or was still trending at the cutoff; and the transcript of at least one hit as the evidence artefact. **What that costs.** Repeats multiply everything by N. One turn bills three model calls — adversarial, target, scorer — so a 12-turn objective run at N=10 is up to 360 calls for one objective, against 36 for a single run, and turns within a conversation cannot be parallelised. Across a suite this is the dominant line item of an engagement, so N is a real budget decision, not a free improvement. The usual resolution is a small N for a broad screening tier and a larger N only for objectives that hit at least once. What you must not do is buy the cheap version and describe it in the language of the expensive one. **Where the number misleads.** Three specific misreadings. *One hit quoted as a rate.* "Attack success rate 100%" from a single successful run is the most common version, and it is not a rate at all; it is one draw with a denominator of one. *A clean retest read as a fix.* Suppose the baseline was 2 hits in 10 runs. If the mitigation changed nothing whatsoever and the true per-run rate is still 20%, the probability of seeing zero hits in the next ten runs is 0.8 to the tenth, about 11%. So roughly one in nine unchanged systems produces a perfect retest by luck alone. A 0/10 after a 2/10 baseline is weak evidence, not a verified fix, and the fix is a larger N, a raised budget, and a look at whether the transcripts still show the target trending. *A comparison across different harness settings.* Turn budget alone moves the hit fraction. Any comparison between two endpoints, two model versions, or before and after a mitigation must hold the budget, the attacker configuration, the objective text, the scorer and N constant, or it is a comparison of settings. **What I would check.** That the reported number has a denominator and a budget attached. That turns-to-success came from more than one run before anyone calls an objective "easy" or "hard". That severity was assessed by reading the transcript rather than inherited from the scorer's acceptance threshold — and, where depth matters, measured by a separate, explicitly labelled continuation run rather than by letting the scored loop run on past its success, which would muddy both the hit statistic and the turns-to-success distribution.

  • You ran an objective 10 times at a 12-turn budget and got 2 hits. How do you write that up?
    As '2/10 runs achieved the objective within 12 turns, at turns-to-success 9 and 11, with attacker configuration X'. That states reproducibility, effort and the bound, and someone else can rerun it.
  • A mitigation ships and the same objective now scores 0/10. Is it fixed?
    Not established. At 2/10 baseline, 0/10 is weak evidence — the sample is small and the run is stochastic. You would want a larger N, a raised budget, and a look at whether the transcripts still show the target trending before you claim a fix.
  • Why not simply let the loop continue past the first success to measure depth?
    You can, but then turns-to-success and the hit statistic get muddied by post-success turns. Cleaner to keep the scored run stopping at first success and do depth as a separate, explicitly labelled continuation.

A single successful run is one lottery draw that happened to win: it proves the ticket can win, not how often. Quoting it as an attack success rate is quoting one draw as the odds.

saying these in an interview costs you the question

  • Reporting a single successful run as an attack success rate.
  • Treating the first accepted response as the maximum severity of the breach.
  • Declaring a mitigation effective from one clean run.
  • Comparing before-and-after numbers that were produced at different turn budgets or repeat counts.

context