skip to content

You have a fixed call allowance against a metered chat endpoint for one PyRIT engagement. How do you split it between many single-send runs across a broad objective list and a few deep multi-turn runs, and how do you choose the per-run turn cap?

level: principalimportance: should knowfreq 34%

answer

  1. breadth to rank, depth to prove
  2. three metered calls per multi-turn turn
  3. cap from observed success turn, not cost
  4. reserve budget for repeats
  5. publish the split, not just the counts

basics

~20 s

Spend the first slice on broad single-send runs to find where the target is soft, then concentrate deep multi-turn runs on those objectives. Set the turn cap from observed successes, not from a guess, and reserve a slice for repeats, because one run against a non-deterministic endpoint is a sample rather than a result.

solid answer

~60 s

I treat it as a **search under budget**, in three slices. **Breadth first.** Single-send runs are cheap and parallel — one target call and one score each — so a broad sweep is the cheapest map of where the target moves at all. It cannot reach anything needing accumulation, and I say so in the report rather than letting it read as coverage. **Depth where breadth showed movement.** Multi-turn costs an attacker call, a target call and a scorer call per turn, and turns are serial. I spend it on objectives whose one-shot transcripts showed partial compliance, hedged refusals or topic drift, not on the ones that were flatly refused. **Repeats reserved.** These endpoints are non-deterministic and guardrails are stochastic. A slice held back for re-running both the hits and a sample of the misses is what turns verdicts into findings. For the cap: pilot a few runs uncapped-ish, look at the turn at which successes actually landed, and set the cap above that distribution's tail. A cap chosen purely for cost manufactures not-achieved results.

go deeper

for a junior

Says do the cheap broad runs first and the expensive deep ones on whatever looked promising.

for a middle

Quantifies the cost difference — one call per single send versus attacker plus target plus scorer per turn, serial — and picks depth targets from one-shot transcripts.

for a senior

Calibrates the turn cap from the observed distribution of success turns, reserves budget for repeats against non-determinism, and recalibrates when the target configuration changes.

for a principal

Treats the allocation as part of the report: states which objectives were only tested at one-shot depth, how the cap was derived, and what share went to repeats, so nobody reads a pass count as coverage.

### What you are actually buying The allowance is a fixed number of metered calls against one endpoint, and the two strategy shapes buy different things per call. A **single-send** run spends one target call plus one scoring decision on one prompt, and executions are independent, so they parallelise up to whatever concurrency the endpoint tolerates. A **multi-turn** run spends three metered calls per turn — the attacker-side model composes the next prompt, the target answers, the scorer judges — and the turns are serial inside a run, so a cap of twelve is up to thirty-six calls and twelve round-trips of wall-clock for one objective. Roughly, one deep run at that cap costs what eighteen one-shot probes cost. That exchange rate is the whole decision. ### Three slices **Breadth first, to rank.** A broad single-send sweep is the cheapest map of where the target moves at all. Its output is not a pass/fail list; it is a sorting — flatly refused, refused with hedging, partially complied, drifted off-topic. It cannot reach anything that requires accumulation, and the report has to say so rather than let the sweep read as coverage. **Depth where breadth showed movement.** Spend the expensive runs on objectives whose one-shot transcript already wobbled. Depth spent on an objective the target refused flatly in one shot is usually the worst return in the engagement; depth spent where it hedged is usually the best. **Repeats reserved.** Ten to twenty percent held back. Both the hits and a sample of the misses get re-run, because a single run against a stochastic model behind a stochastic guardrail is one draw. ### Choosing the turn cap The cap is the knob most often set wrong, and it goes wrong in the direction that flatters the target. Method: - Pilot a handful of runs with a deliberately generous cap. - Record the turn index at which the successes actually landed. - Set the production cap **above the tail** of that distribution, not at its median — a median cap silently drops the slower half of the real successes and files them as non-successes. - Recalibrate whenever the target, its system prompt, its guardrail or the attacker-side model changes. The cap is calibrated against one configuration, not against the objective in the abstract. ### Where the numbers mislead - **“We tested 200 objectives.”** A count of executions, not coverage. If 180 of them were single-send only, the sentence overstates the depth of the engagement by an order of magnitude. - **“The model resisted 96%.”** The denominator silently includes objectives the chosen shape could not express and runs cut off by a cost-derived cap. Both are non-successes that say nothing about the target. - **Cap-bunching.** If every deep run ended exactly at the cap, the cap is binding, and some of those runs were on track when they were cut off. - **A one-off hit reported as a finding.** Without repeats you cannot separate reproducible from lucky, and an irreproducible finding costs the customer's team far more time than the repeat would have cost you. - **Errors counted as resistance.** Rate-limit aborts and expired credentials land in the same non-success column as genuine refusals unless somebody splits them out. ### What I would check, and what I publish Before writing anything: the distribution of turns used against the cap; error counts per objective; who refused inside the deep transcripts, target or attacker side; and a hand-scored sample of the misses to estimate the scorer's miss rate. Then the allocation itself becomes part of the deliverable — how many objectives were only ever tested at one-shot depth, what cap the deep runs used and how it was derived, what share of the allowance went to repeats, and how much was lost to errors and rate limits. A report that gives a pass count without the allocation is asserting coverage it did not buy, and the first competent reader will ask exactly that question.

  • All your deep runs ended exactly at the turn cap. What does that tell you about the cap?
    That it is probably binding rather than generous. Some of those runs may have been on track and were cut off, so I would re-run a sample with a higher cap before reporting any of them as failures to reach the objective.
  • How do you justify spending part of a scarce budget on re-running objectives you already have a verdict for?
    Because a single run against a stochastic endpoint and a stochastic guardrail is one sample. Without repeats you cannot distinguish a reproducible failure from a one-off, and a finding that does not reproduce wastes far more of the customer's time than the repeat cost.

saying these in an interview costs you the question

  • Spreads the budget uniformly across objectives with no ranking pass.
  • Sets the turn cap from the money available and never checks it against where successes actually landed.
  • Keeps no budget for repeats on a non-deterministic endpoint.
  • Reports objective counts without saying which were only tested single-send.
  • Spends the deep budget on objectives that were flatly refused one-shot, purely because they sound severe.

context