skip to content

Driving a Run

An objective, a turn budget, a stop rule and the surfaces you actually wired a target for decide what a campaign can report. Interviewers ask what the run cost and what it never touched.

on this pageshow

explore

questions

15

In a PyRIT multi-turn run, how many model calls does a single turn bill, and which components make them?

level: juniorimportance: must knowfreq 72%

answer

  1. turn bills three: adversary, target, scorer
  2. unit of cost is a turn, not a prompt
  3. adversarial context grows each turn
  4. refused turns still bill
  5. rule-based scorer bills nothing

basics

~20 s

Three. A PyRIT turn bills the adversarial model that writes the next prompt, the target under test that answers it, and the scorer that judges the reply. Every turn pays all three, so a fifty-turn objective is roughly one hundred fifty billed calls, not fifty.

solid answer

~50 s

PyRIT drives a loop with three distinct model-shaped dependencies, and a naive estimate counts only one of them. - **The adversarial model** generates the next attack prompt from the objective and the conversation so far. One call per turn. - **The prompt target** is the system under test. One call per turn (more if converters fan the prompt out into variants). - **The scorer** decides whether the target's response counts as a hit. One call per response, per scorer configured — and a scorer backed by a model is itself a metered endpoint. So the unit of cost is a turn, not a prompt, and each turn is at least three billed calls across up to three different accounts. The simple answer breaks when a scorer is rule-based rather than model-backed, or when the adversarial side is a fixed seed-prompt list rather than a model — then the turn bills two, or one. Say which shape you are running before quoting a number.

go deeper

for a junior

Should say a turn is not one call — the attacking model, the target and the scorer are each billed — and that cost scales with turns.

for a middle

Adds that converters fan out target calls and that each configured scorer adds a call per response, and knows that the adversarial leg's input grows across the transcript.

for a senior

Estimates a real objective's spend by leg, notes that failed and refused turns bill identically, and checks the stored transcript for actual call counts instead of trusting arithmetic.

for a principal

Frames the three legs as three separate quotas and vendors, and decides which legs are worth a frontier model at engagement scale versus a cheaper or self-hosted substitute.

The mental model most people arrive with is "a red-team run sends prompts to the model I am testing". That counts one of three legs, and it is why cost estimates for PyRIT runs are routinely wrong by a multiple rather than by a margin. ### The three legs, by the name that configures them PyRIT's multi-turn attack objects — the red-teaming attack, and the automated jailbreak strategies the package ships such as PAIR and tree-of-attacks-with-pruning — are constructed with three separate model-shaped dependencies. The constructor arguments name them: - **The adversarial chat** (PyRIT's `adversarial_chat` target, supplied through the attack's adversarial configuration). This is a model whose job is to read the objective and the transcript so far and compose the next attack prompt. One call per turn. - **The objective target** (PyRIT's `objective_target`, an implementation of PyRIT's `PromptTarget` interface — `OpenAIChatTarget`, an HTTP target, a self-hosted endpoint). This is the system under test. One call per prompt actually sent: one per turn before converters, more once a `PromptConverter` fans the prompt into variants. - **The objective scorer** (PyRIT's `objective_scorer`, an implementation of its `Scorer` interface). A model-backed scorer such as `SelfAskTrueFalseScorer` sends the target's response to a judge model and asks whether the objective was met — that is a metered call per response, per scorer attached. A rule-based scorer such as `SubStringScorer` matches text in-process and calls nothing. The loop is: adversarial chat writes a prompt, converters may rewrite it, each resulting prompt goes to the objective target, each response goes to every attached scorer, and the score conditions the next adversarial turn. Only the middle leg touches the system you were hired to assess. The other two are your own spend, usually on your own accounts, and usually invisible in whichever dashboard you happened to be watching. ### Why it is shaped that way The adversarial side has to be a model if you want turns that adapt to what the target just refused; a fixed list of seed prompts cannot escalate. The scorer has to be a model if you want a verdict on open-ended prose rather than a substring match. Both capabilities are bought with tokens, on every turn, whether or not the turn produces anything. ### What it costs Call count is the easy part: roughly three per turn, so a `max_turns` of 50 is about 150 billed calls for a single objective, not 50. Token cost is the part that surprises people, because the three legs do not grow alike. The objective target sees one prompt and answers it; its input is roughly flat across the run. The scorer sees one response at a time; also flat, and cheap per call, but numerous. The adversarial chat is conditioned on the conversation so far, so its input grows with turn number — by turn 50 it may be sending tens of thousands of tokens of transcript, and its cumulative input across an objective grows roughly with the square of the turn count. On a long objective the adversarial leg, not the target, is often the largest line on the bill. Wall clock follows the same structure. The three calls in a turn are strictly sequential — you cannot score a response you have not received — so a turn costs the sum of three round trips, commonly several seconds each. A 50-turn objective takes minutes of real time before any other objective is considered. ### Where the number misleads The classic bad estimate is `turns × the target's per-call price`. It is wrong three ways at once: it omits two legs, it treats a flat per-call price as if the adversarial leg's input were constant, and it silently assumes the turn budget is the expected number of turns rather than a ceiling. Two more readings go wrong in the field. First, **a refused or failed turn bills exactly like a successful one** — the adversarial model wrote a prompt, the target answered with a refusal, the scorer judged it a miss, and all three were metered. A run that reported nothing was not a free run; in a low-success sweep it is close to the most expensive kind. Second, **the target's usage dashboard understates the engagement by roughly two thirds**, because the other two legs bill to different keys, often different vendors, and frequently a different team's budget. People sign off on a red-team programme having looked at one of the three invoices. ### What to check before believing a figure Establish, for the specific configuration in front of you, which legs are model-backed at all — a rule-based scorer removes a whole leg, a seed-prompt-driven attack removes another. Check whether the adversarial transcript is truncated or summarised or allowed to grow unbounded, since that decides whether token cost is linear or superlinear in turns. Then stop estimating and measure: run one objective end to end and count what actually happened out of PyRIT's memory store, which persists every request and response piece with the target that produced it. Compare that count against the three providers' usage for the same window. The ratio between your arithmetic and the store is the correction factor to apply to the whole engagement, and it is the number to put in front of whoever approves the spend.

  • Which of the three legs grows in cost as an objective runs longer, and why?
    The adversarial model. Its input is the conversation so far, so per-turn input tokens climb with turn number unless the transcript is truncated or summarised. The target and scorer see roughly constant-size inputs.
  • When does a PyRIT turn bill fewer than three calls?
    When the scorer is rule-based rather than model-backed (a pattern or classifier you run locally), or when the attacking side is a fixed list of seed prompts rather than an adversarial model. Then a turn bills two legs, or one.
  • You self-host the target. Does that make the run free?
    No. It moves the target's cost from tokens to your own compute and removes one metered quota, but the adversarial model and a model-backed scorer are still billed, and hosted inference still consumes GPU time you are paying for.

saying these in an interview costs you the question

  • Quoting a run's cost as turns times the target's price only.
  • Assuming the scorer is free because it 'just labels' the output.
  • Not knowing whether the scorer in their own run was model-backed or rule-based.
  • Believing a run that found nothing cost nothing.

context

open as a page

In PyRIT, what do you have to configure before a multi-turn adversarial run can start, and what are the two ways the loop can stop?

level: juniorimportance: must knowfreq 68%

basics

~20 s

You set four things: an objective describing what the system under test should be made to do, an adversarial chat model that writes each attacker turn, the target endpoint being tested, and a scorer that judges the target's reply. The loop ends when that scorer says the objective was met, or when the turn budget runs out.

open as a page

A PyRIT run against your deployed chat assistant finishes with no successful attacks recorded. What has that run actually tested, and what has it not?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Only the targets you configured. PyRIT sends prompts to the prompt targets you wired and scores the replies; anything it was never pointed at, such as another endpoint, an upload path or an internal service, is untested rather than proven safe. A clean result describes your target list, not the application.

open as a page

Before launching a PyRIT run, how do you estimate its total call count when the configuration includes several seed prompts, converters and scorers?

level: middleimportance: must knowfreq 60%

basics

~20 s

Multiply, do not add. Target calls are objectives times turns times converter variants. Scoring calls are target responses times the number of scorers attached. Adding one converter and one scorer roughly doubles two legs at once. Work the product out on paper first, then dry-run one objective to check it.

open as a page

A PyRIT multi-turn run ends with 'objective not achieved' after exhausting a 10-turn budget. Why is that not evidence that the target refuses the behaviour?

level: middleimportance: must knowfreq 62%

basics

~20 s

It only says the adversarial model did not get there within the turns you allowed. The budget is part of the result, not a property of the target. A longer run, a different attacker configuration, or another attack strategy can still breach the same endpoint. Report the negative together with the turn budget that produced it.

open as a page

When you configure a PyRIT prompt target, what is the difference in what a clean run proves if you wire it at the model provider's raw API versus at your application's own endpoint?

level: middleimportance: must knowfreq 58%

basics

~20 s

The raw API tests the model alone, with no production system prompt, retrieval, tools or guard. The application endpoint tests the whole deployed stack. A clean run at one says nothing about the other: a guard can hide model weakness, and a bare-model result ignores every layer users actually pass through.

open as a page

In a PyRIT multi-turn run the objective text is given both to the adversarial model and to the objective scorer. Why does a vaguely worded objective make the run untrustworthy in both directions?

level: middleimportance: should knowfreq 45%

basics

~20 s

The same sentence steers the attacker and defines success. If it names no observable artefact in the target's reply, the adversarial model has nothing concrete to drive toward and drifts, while the scorer has nothing concrete to check and can mark a hedged, harmless answer as met. Write objectives a reviewer can verify from the transcript.

open as a page

Your PyRIT run is too slow, so you raise its concurrency. What does that change about the run's wall-clock time and about its cost?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Concurrency changes arrival rate, not work. The same calls are made, so spend is unchanged; you only reach the quota sooner. Wall clock improves until the tightest-quota of the three endpoints saturates, then stops improving, and past that point the other two legs idle behind the bottleneck.

open as a page

A PyRIT run bills an adversarial model, the prompt target and a scorer. If you must cut spend, which of those three can you move to a cheaper or locally-hosted model, and what does each substitution cost you?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The target cannot change — it is the system under test, and swapping it invalidates the run. The adversarial model and the scorer can. A weaker adversarial model finds fewer breaches, so results understate risk. A weaker scorer misjudges responses, which corrupts the number the whole run reports.

open as a page

A PyRIT multi-turn run halts the moment its objective scorer marks a turn as success. What does that early stop leave out of your engagement report, and how would you run it differently?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Stopping at first success gives an existence proof and nothing else: no severity, no repeatability, no sense of whether it took two turns or nine. One run is a single sample of a stochastic search. Repeat each objective several times, record turns-to-success across runs, and hold the budget fixed across targets you compare.

open as a page

A PyRIT run against a support assistant returned no hits, but an incident later shows the same behaviour was reachable in production. How do you work out whether the target wiring, rather than the attack content, produced the clean result?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Check reachability before blaming the prompts. Confirm the incident path was in the target list at all, replay a positive control through each target, read the stored attempts for empty bodies, errors, auth rejections and timeouts, and compare the target's endpoint, tenant and deployment identifier against what production served.

open as a page

PyRIT is wired to a support assistant's chat endpoint and the run is clean. Which other intake paths of that application would you wire as separate prompt targets before calling the surface covered, and what does wiring each one require?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Any path whose content reaches the same model: document or ticket ingestion, file and attachment parsing, transcribed voice or extracted image text, scheduled and batch jobs, webhook consumers, admin or internal endpoints, other tenants and older routed API versions. Each needs its own send path, matching credentials, and somewhere the effect can be observed.

open as a page

You have a fixed spend for a PyRIT engagement. How do you divide it between many objectives at a short turn budget and few objectives explored deeply, and how do you report what the money bought?

level: principalimportance: should knowfreq 38%

basics

~20 s

Split the budget: a broad shallow pass to find which objectives show movement, then deep turn budgets spent only on those. Cap spend per objective so one runaway conversation cannot eat the pass. Report cost per objective and per confirmed finding, and state which objectives were cut short rather than cleared.

open as a page

You have a fixed pool of adversarial turns to spend across roughly 40 objectives in a PyRIT engagement. How do you allocate the turn budget between deep runs on a few objectives and shallow runs on many, and how do you report the choice?

level: principalimportance: should knowfreq 33%

basics

~20 s

Do not spread it evenly. Run a short screening pass over all objectives, then spend deep budgets only on the ones whose transcripts showed the target trending rather than refusing flat. Report every negative with the budget that produced it, treat shallow negatives as untested rather than safe, and keep budgets fixed between retests.

open as a page

Your team re-runs the same PyRIT target set against a fast-moving product every release. How do you keep the wired target set from silently drifting behind the surfaces the product actually ships?

level: principalimportance: should knowfreq 34%

basics

~20 s

Derive the target set instead of remembering it. Reconcile wired targets against the artefacts the platform deploys from — routes, service registry, data-flow records — and flag any model-reaching path with no target and no expiring waiver. Make registering a target part of shipping a new intake path, and track wired-over-deployed as a first-class number.

open as a page