skip to content

In a PyRIT multi-turn run, how many model calls does a single turn bill, and which components make them?

level: juniorimportance: must knowfreq 72%

answer

  1. turn bills three: adversary, target, scorer
  2. unit of cost is a turn, not a prompt
  3. adversarial context grows each turn
  4. refused turns still bill
  5. rule-based scorer bills nothing

basics

~20 s

Three. A PyRIT turn bills the adversarial model that writes the next prompt, the target under test that answers it, and the scorer that judges the reply. Every turn pays all three, so a fifty-turn objective is roughly one hundred fifty billed calls, not fifty.

solid answer

~50 s

PyRIT drives a loop with three distinct model-shaped dependencies, and a naive estimate counts only one of them. - **The adversarial model** generates the next attack prompt from the objective and the conversation so far. One call per turn. - **The prompt target** is the system under test. One call per turn (more if converters fan the prompt out into variants). - **The scorer** decides whether the target's response counts as a hit. One call per response, per scorer configured — and a scorer backed by a model is itself a metered endpoint. So the unit of cost is a turn, not a prompt, and each turn is at least three billed calls across up to three different accounts. The simple answer breaks when a scorer is rule-based rather than model-backed, or when the adversarial side is a fixed seed-prompt list rather than a model — then the turn bills two, or one. Say which shape you are running before quoting a number.

go deeper

for a junior

Should say a turn is not one call — the attacking model, the target and the scorer are each billed — and that cost scales with turns.

for a middle

Adds that converters fan out target calls and that each configured scorer adds a call per response, and knows that the adversarial leg's input grows across the transcript.

for a senior

Estimates a real objective's spend by leg, notes that failed and refused turns bill identically, and checks the stored transcript for actual call counts instead of trusting arithmetic.

for a principal

Frames the three legs as three separate quotas and vendors, and decides which legs are worth a frontier model at engagement scale versus a cheaper or self-hosted substitute.

The mental model most people arrive with is "a red-team run sends prompts to the model I am testing". That counts one of three legs, and it is why cost estimates for PyRIT runs are routinely wrong by a multiple rather than by a margin. ### The three legs, by the name that configures them PyRIT's multi-turn attack objects — the red-teaming attack, and the automated jailbreak strategies the package ships such as PAIR and tree-of-attacks-with-pruning — are constructed with three separate model-shaped dependencies. The constructor arguments name them: - **The adversarial chat** (PyRIT's `adversarial_chat` target, supplied through the attack's adversarial configuration). This is a model whose job is to read the objective and the transcript so far and compose the next attack prompt. One call per turn. - **The objective target** (PyRIT's `objective_target`, an implementation of PyRIT's `PromptTarget` interface — `OpenAIChatTarget`, an HTTP target, a self-hosted endpoint). This is the system under test. One call per prompt actually sent: one per turn before converters, more once a `PromptConverter` fans the prompt into variants. - **The objective scorer** (PyRIT's `objective_scorer`, an implementation of its `Scorer` interface). A model-backed scorer such as `SelfAskTrueFalseScorer` sends the target's response to a judge model and asks whether the objective was met — that is a metered call per response, per scorer attached. A rule-based scorer such as `SubStringScorer` matches text in-process and calls nothing. The loop is: adversarial chat writes a prompt, converters may rewrite it, each resulting prompt goes to the objective target, each response goes to every attached scorer, and the score conditions the next adversarial turn. Only the middle leg touches the system you were hired to assess. The other two are your own spend, usually on your own accounts, and usually invisible in whichever dashboard you happened to be watching. ### Why it is shaped that way The adversarial side has to be a model if you want turns that adapt to what the target just refused; a fixed list of seed prompts cannot escalate. The scorer has to be a model if you want a verdict on open-ended prose rather than a substring match. Both capabilities are bought with tokens, on every turn, whether or not the turn produces anything. ### What it costs Call count is the easy part: roughly three per turn, so a `max_turns` of 50 is about 150 billed calls for a single objective, not 50. Token cost is the part that surprises people, because the three legs do not grow alike. The objective target sees one prompt and answers it; its input is roughly flat across the run. The scorer sees one response at a time; also flat, and cheap per call, but numerous. The adversarial chat is conditioned on the conversation so far, so its input grows with turn number — by turn 50 it may be sending tens of thousands of tokens of transcript, and its cumulative input across an objective grows roughly with the square of the turn count. On a long objective the adversarial leg, not the target, is often the largest line on the bill. Wall clock follows the same structure. The three calls in a turn are strictly sequential — you cannot score a response you have not received — so a turn costs the sum of three round trips, commonly several seconds each. A 50-turn objective takes minutes of real time before any other objective is considered. ### Where the number misleads The classic bad estimate is `turns × the target's per-call price`. It is wrong three ways at once: it omits two legs, it treats a flat per-call price as if the adversarial leg's input were constant, and it silently assumes the turn budget is the expected number of turns rather than a ceiling. Two more readings go wrong in the field. First, **a refused or failed turn bills exactly like a successful one** — the adversarial model wrote a prompt, the target answered with a refusal, the scorer judged it a miss, and all three were metered. A run that reported nothing was not a free run; in a low-success sweep it is close to the most expensive kind. Second, **the target's usage dashboard understates the engagement by roughly two thirds**, because the other two legs bill to different keys, often different vendors, and frequently a different team's budget. People sign off on a red-team programme having looked at one of the three invoices. ### What to check before believing a figure Establish, for the specific configuration in front of you, which legs are model-backed at all — a rule-based scorer removes a whole leg, a seed-prompt-driven attack removes another. Check whether the adversarial transcript is truncated or summarised or allowed to grow unbounded, since that decides whether token cost is linear or superlinear in turns. Then stop estimating and measure: run one objective end to end and count what actually happened out of PyRIT's memory store, which persists every request and response piece with the target that produced it. Compare that count against the three providers' usage for the same window. The ratio between your arithmetic and the store is the correction factor to apply to the whole engagement, and it is the number to put in front of whoever approves the spend.

  • Which of the three legs grows in cost as an objective runs longer, and why?
    The adversarial model. Its input is the conversation so far, so per-turn input tokens climb with turn number unless the transcript is truncated or summarised. The target and scorer see roughly constant-size inputs.
  • When does a PyRIT turn bill fewer than three calls?
    When the scorer is rule-based rather than model-backed (a pattern or classifier you run locally), or when the attacking side is a fixed list of seed prompts rather than an adversarial model. Then a turn bills two legs, or one.
  • You self-host the target. Does that make the run free?
    No. It moves the target's cost from tokens to your own compute and removes one metered quota, but the adversarial model and a model-backed scorer are still billed, and hosted inference still consumes GPU time you are paying for.

saying these in an interview costs you the question

  • Quoting a run's cost as turns times the target's price only.
  • Assuming the scorer is free because it 'just labels' the output.
  • Not knowing whether the scorer in their own run was model-backed or rule-based.
  • Believing a run that found nothing cost nothing.

context