In PyRIT, what distinguishes a single-send attack strategy from a multi-turn one, and what actually happens when you point a single-send strategy at an objective that only succeeds after several turns?
answer
- single-send = fan-out, no feedback
- multi-turn = loop + stopping condition
- attacker + target + scorer per turn
- mismatch manufactures a false negative
- not-achieved describes the run, not the model
basics
~20 sA single-send strategy delivers each prompt once, scores the reply and stops; nothing feeds forward. A multi-turn strategy loops: it reads the last reply, composes the next prompt with an attacker model, and scores each turn until the scorer says achieved or the turn cap runs out. Single-send on a multi-turn goal returns not-achieved, having proved nothing.
solid answer
~60 sThe difference is **feedback**. A single-send strategy is a fan-out: prompts in, one exchange each, a score per exchange. There is no state carried from a response into the next prompt, so it can only find things reachable in one shot — a seed prompt that lands, a converter transform that slips past a filter. A multi-turn strategy is a **loop with a stopping condition**. Each iteration takes the target's last response, has an attacker-side model produce the next prompt from it, sends it through the same target, and asks the scorer whether the objective is now met. The loop ends on an achieved verdict or when the turn cap is spent. Run single-send against an objective that needs accumulation and you spend the whole prompt list, get uniform not-achieved results, and learn nothing about the target — the run failed because the strategy could not express the attack, not because the target held. That is a **false negative manufactured by the harness**, and it is worse than no test, because it lands in a report as a clean result.
go deeper
Knows single-send sends each prompt once and multi-turn keeps going, and that the wrong choice can miss things.
Explains the feedback loop, the per-turn cost of attacker plus target plus scorer calls, and names the mismatch as a false negative produced by the harness.
Adds how they would detect the mismatch after the fact — turns used against cap, whether the target surface carries history, whether the attacker side refused — and refuses to report not-achieved as evidence of safety.
Treats strategy shape as a coverage claim in the report, and insists the report say which objectives were only ever tested single-send.
### The two shapes An **attack strategy** in PyRIT is the object that owns the send-and-score loop. It is built with a **prompt target** (the adapter that reaches one endpoint), a **scorer** (the object that reads a reply and returns a verdict), optionally a **converter** chain (transforms applied to an outgoing prompt), and — for the looping variants — a **turn cap**, the maximum number of exchanges one execution may spend. A **single-send** strategy is a fan-out. It takes each prompt in turn, sends it once, hands the reply to the scorer, records the result and moves on. Nothing a reply contains reaches the next prompt. Its state is empty by design. A **multi-turn** strategy is a loop with a stopping condition. Each iteration reads the transcript so far, has an **attacker-side model** — a second model, configured separately from the target — compose the next prompt from what the target just said, sends that through the same target, and asks the scorer whether the objective is now satisfied. The loop exits on an achieved verdict or when the turn cap is spent. The whole difference is one word: **feedback**. ### What each costs | shape | unit | parallel? | metered calls per unit | |---|---|---|---| | single-send | one prompt | yes, up to the endpoint's rate limit | 1 target + 1 scorer (2 when the scorer is model-backed) | | multi-turn | one turn | no — turns are serial inside a run | 1 attacker + 1 target + 1 scorer (3) | The multiplier is what surprises people. A turn cap of ten is not a ten-call run; it is up to thirty metered calls, across as many as three different models on three price tiers, every one of them serial — so wall-clock scales with the cap as well as money. A hundred objectives at a cap of ten is a four-figure call count before a single repeat. The same hundred objectives single-send is two hundred calls that finish as fast as your concurrency allows. ### The mismatch, and why it is the worst outcome Point a single-send strategy at an objective that only succeeds once the target has already produced something to build on, and the run completes cleanly. Every prompt scores not-achieved. The results file is uniform, complete, and wrong in spirit: nothing in it separates *the target refused* from *this strategy could not express the attack in the first place*. That is a **false negative manufactured by the harness**, and it is worse than not running the test at all. A missing test is visibly missing. A full sheet of not-achieved results is read as coverage by everyone downstream who never sees which shape ran — and the number that travels is the objective count, which is identical either way. ### The reverse mismatch Multi-turn on an objective a single prompt would already satisfy is mostly a budget problem: three calls a turn to reach a conclusion the first send already carried. It has one correctness hazard too. Over several turns the attacker-side model can drift onto an adjacent, easier goal that the scorer still accepts, producing an achieved verdict for something nobody asked about. The transcript is the only place that shows it; the verdict field looks like a win. ### What I would check - **Which shape actually ran, per objective**, and how many turns each run used against its cap. A looping strategy whose runs all stopped at turn one did not perform multi-turn runs. - **Whether the target adapter carries conversation history at all.** A stateless surface silently reduces the loop to repeated one-shots: the strategy loops correctly, the endpoint just never sees the prior exchange — at multi-turn prices. - **Who refused in the transcript.** An attacker-side model with its own guardrails will decline to compose the next prompt, and the run ends polite, short and empty, having tested nothing about the target. - **A sample of transcripts, read by a human**, not just the verdict column. A not-achieved verdict is a statement about the run — its shape, its cap, its scorer, its target adapter — and only derivatively about the model.
- Your multi-turn run against a chat endpoint behaves exactly like repeated one-shots. What would you suspect first?That the target surface is not carrying conversation history — a stateless adapter or an endpoint that ignores prior turns reduces the loop to independent sends even though the strategy is looping correctly.
- When is single-send the better choice even though multi-turn is strictly more capable?For breadth and regression: it parallelises, costs one target call plus one score per prompt, and is enough for anything reachable in one shot, such as sweeping a seed list or re-checking a fix.
- How would you avoid burning a whole budget on a strategy that cannot express the objective?Pilot a handful of runs first and read the transcripts, not just the verdicts. If no run ever gets a response worth building on, the shape is wrong and no amount of extra prompts fixes it.
Every turn of a multi-turn run is metered three times over: once to think of the next thing to say, once to say it, once to judge the answer. A turn cap of ten is a thirty-call bill, not a ten-call one.
saying these in an interview costs you the question
- Reports a not-achieved sweep as evidence the model is robust, without saying which strategy shape ran.
- Thinks multi-turn just means sending more prompts, with no notion of a response feeding the next prompt.
- Cannot say what ends a multi-turn loop.
- Ignores that multi-turn adds an attacker-model call per turn and treats it as free breadth.
- Believes a stateless endpoint still gives a real multi-turn run.