In PyRIT, what do you have to configure before a multi-turn adversarial run can start, and what are the two ways the loop can stop?
answer
- objective, attacker model, target, scorer
- two exits: scorer met, budget exhausted
- attacker turn is a model call, not a script
- transcript and scores land in memory
- not-achieved is budget-bounded
basics
~20 sYou set four things: an objective describing what the system under test should be made to do, an adversarial chat model that writes each attacker turn, the target endpoint being tested, and a scorer that judges the target's reply. The loop ends when that scorer says the objective was met, or when the turn budget runs out.
solid answer
~60 sA PyRIT multi-turn run wires four pieces together and then loops. - **Objective** — the natural-language goal of the run, the thing you want the target to end up doing. - **Adversarial chat model** — given the objective plus the conversation so far, it writes the next attacker turn. It is a model call, not a fixed script. - **Target** — whatever endpoint you registered as the system under test; the attacker turn is sent to it and its reply comes back. - **Objective scorer** — reads the target's reply each turn and returns a met / not-met judgement. That judgement is one stop rule; the maximum turn count is the other. The part candidates miss is that the two exits mean opposite things. Stopping because the scorer said *met* is a claim about the target. Stopping because you ran out of turns is a claim about your budget — but both land in the report as *objective not achieved* versus *achieved*, with the budget only visible if you kept it.
go deeper
Name the four pieces and both stop conditions, and know the attacker turn is generated by a model rather than replayed from a file.
Add that the objective is handed to both the attacker and the scorer, and that the not-achieved exit is bounded by the turn budget.
Talk about what you inspect in memory after a run — which exit fired, turns consumed, the score that stopped it — and about running each objective more than once.
Frame it as evidence quality: define what an organisation's standard multi-turn run configuration is, so negatives from different engagements are comparable at all.
**The four pieces, defined.** A PyRIT multi-turn run is an automated conversation between two models, refereed by a third. Four things must exist before it can start. - **The objective** — a single sentence of free text saying what state you want the system under test to end up in. It is not a label for the report; the framework passes it into the run as a live input, twice per turn. - **The adversarial chat model** (PyRIT's *adversarial chat* endpoint) — a second, separately configured model whose job is to write the attacker's next message given the objective and the transcript so far. It is a generation, not a replay of a file. - **The target** — the endpoint under test, registered as a PyRIT *prompt target*. In this framework "target" means the thing being attacked; the same word means a garak generator or a promptfoo provider elsewhere in the tooling landscape, so name the tool when you use it. - **The objective scorer** — a PyRIT *scorer* that reads the target's reply each turn and returns a met / not-met judgement against the objective. It is a referee attached to your harness, not a filter running inside the target's own stack; that distinction is the one interviewers probe. Optionally you register **converters**, which transform the attacker's message before it is sent, and everything the run touches is written to PyRIT's **memory** — attacker turns, target responses, score records — which is what you query afterwards. **One iteration, in order.** Pseudocode for the control flow, deliberately not an API listing, because these class names get renamed between releases: ```text turn = 0 while turn < max_turns: attacker_msg = adversarial_model(objective, transcript) response = target(convert(attacker_msg)) transcript += [attacker_msg, response] if objective_scorer(response, objective).met: return SUCCESS(turn) turn += 1 return NOT_ACHIEVED(max_turns) ``` **What a turn costs.** In the ordinary configuration one turn bills three model calls: the adversarial model writes, the target answers, the scorer judges. A ten-turn run that goes the distance is therefore about thirty metered calls, not ten, and an LLM-backed converter adds a fourth per turn. Turns within one conversation are strictly sequential — turn *n+1* needs turn *n*'s reply — so a ten-turn run's wall clock is roughly ten times the sum of three round trips, commonly one to three minutes; the only parallelism available is across objectives or repeats, never inside a conversation. Sweeping forty objectives five times each at ten turns is on the order of six thousand billed calls before a human reads a single transcript. That arithmetic, not curiosity, is what makes people set the turn budget low. **The two exits, and why they are not symmetric.** `SUCCESS(turn)` is a claim about the target: an attacker got it to a state your scorer accepted. `NOT_ACHIEVED(max_turns)` is a claim about your budget: this attacker, at this cap, on this draw, did not get there. Both land in the report as a single verdict string, and the budget only survives if you deliberately carry it. That asymmetry is the central defect of the loop as a reporting instrument. **Where the number misleads.** A run summary reading "objective not achieved" gets read upstream as "the model refused". It is not: the target may have been visibly softening on the last turn when the cap fired. In the other direction, the attacker turn is a sampled generation, so two runs of a byte-identical configuration diverge on the first turn and never re-converge — a single verdict, positive or negative, is one draw from a distribution, not a measurement. And a met verdict is only as good as the scorer: a hedged reply that names the topic without doing the thing can trip a loose scorer and end the run early, giving you a "success" nobody can confirm from the transcript. **What I would check before believing a run.** Which of the two exits fired. What the turn budget was and how many turns were actually consumed — a run that stopped at turn ten of ten is a different artefact from one that stopped at turn three of ten. The specific score record that ended the loop, read against the response it scored. And the attacker turns themselves, to confirm the adversarial model was pursuing the stated objective rather than wandering across framings, because a wandering run tested nothing at all.
- Where do the attacker turns and the score records live after a PyRIT run ends?In the framework's memory store, alongside the conversation. That is what you query to rebuild the transcript, count turns consumed and see which score triggered the stop — the console summary alone will not tell you.
- If you leave the objective scorer out, what happens to the loop?You lose the success stop rule, so the run simply plays out the full turn budget and you triage the transcript by hand. That is a legitimate exploratory mode, but it produces no machine-readable verdict.
- Does the adversarial model have to be the same model as the target?No, and usually it should not be. The attacker is a separate configured endpoint chosen for willingness to generate adversarial turns; the target is the system under test.
saying these in an interview costs you the question
- Describing the attacker turns as a fixed, pre-written script rather than generated per turn.
- Treating 'objective not achieved' as equivalent to 'the target refused'.
- Not knowing there is a maximum turn count at all.
- Confusing the scorer that ends the loop with an output filter on the target.