Your PyRIT run is too slow, so you raise its concurrency. What does that change about the run's wall-clock time and about its cost?
answer
- arrival rate, not work
- min of three service rates
- one objective is sequential; parallel across objectives
- past saturation: queue depth, not completions
- cut work to cut both time and cost
basics
~20 sConcurrency changes arrival rate, not work. The same calls are made, so spend is unchanged; you only reach the quota sooner. Wall clock improves until the tightest-quota of the three endpoints saturates, then stops improving, and past that point the other two legs idle behind the bottleneck.
solid answer
~1 minTwo things are worth separating. **Cost.** The run performs a fixed amount of billed work — objectives × turns × the converter and scorer multipliers. Running them in parallel does not remove any of those calls, so total spend is identical. The only cost concurrency adds is failed work that has to be repeated. **Throughput.** A PyRIT run pulls on three separate endpoints — adversarial model, prompt target, scorer — usually with three different quotas. Parallelism helps only until the *lowest*-capacity of them saturates. After that, adding workers just deepens a queue in front of the bottleneck; the other two legs sit idle waiting for it, and per-turn latency rises even though throughput is flat. The non-obvious part is that the bottleneck is frequently *not* the system under test. A generous production target paired with a modest quota on the adversarial model or the scorer means you are throttled by your own tooling, and the natural fix — a cheaper or self-hosted scorer — changes what the run measures, not just what it costs. When the simple answer breaks: a multi-turn objective is sequential within itself. Turn N+1 depends on turn N's scored result, so concurrency can only run *different objectives* in parallel, never the turns of one.
go deeper
Should say that running calls in parallel does not reduce how many calls are made, so the bill is the same.
Adds that throughput is capped by the tightest of the three endpoints and that a multi-turn objective is internally sequential.
Diagnoses which leg saturates from per-leg latency and in-flight counts, and knows that partially-failed turns bill the legs that already succeeded and that heavy load changes the target's behaviour.
Decides whether to buy quota, cut scope or accept a longer run, and insists the report states which knob was turned because reducing work changes coverage while raising concurrency does not.
Two quantities are being confused whenever someone says raising concurrency will "help with cost". A run has a fixed amount of **work** — a set of billed calls determined entirely by configuration: objectives, `max_turns`, converter fan-out, scorers attached. It also has an **arrival rate** — how fast those calls are issued. Concurrency is a knob on the second only. Spend tracks the first. ### Why the bill does not move Billing is per call and per token, not per unit of elapsed time. Running the same 6,500 calls over one hour instead of six hours issues the same 6,500 calls. Nothing is skipped, and there is no volume behaviour that rewards density. The only cost concurrency can add is work that fails and is repeated, and past saturation that term stops being negligible: a turn that generated an adversarial prompt and got a target response and then failed at the scoring step has already burned two of its three legs and stored nothing you can report. ### Where the parallelism actually is, and is not PyRIT exposes the knob in more than one place — a concurrency bound on the executor that fans multiple objectives out at once, and a per-target throttle such as `max_requests_per_minute` on the target object itself, which paces that endpoint regardless of how many workers sit above it. What neither can do is parallelise inside a multi-turn objective. That loop is a feedback loop: the adversarial chat composes turn N+1 from the target's response to turn N and the scorer's verdict on it. Turn N+1 does not exist until turn N has completed all three of its legs. The parallel units are therefore *across* objectives and seed prompts, and across the converter variants of a single turn — never the turns themselves. A run with three objectives and a 50-turn budget cannot be made fast by raising a worker count; its floor is roughly 50 sequential turns of latency no matter what you set. ### Where it stops helping Model the three legs as three queues with three service rates. Pipeline throughput is the minimum of the three, and every worker added past that minimum converts into queue depth rather than completions: throughput flat, per-turn latency rising. The non-obvious part is that the saturating leg is frequently **not** the system under test. A generously provisioned production target paired with a modest quota on your own adversarial model or judge model means you are throttled by your own tooling, and the natural-looking fix — move the scorer to something cheaper or self-hosted — changes what the run measures, not merely what it costs. ### Where it makes the result wrong, not just slow Two failure modes matter more than the wasted wall clock. First, **partially-completed turns bill the legs that succeeded** and produce no scored row, so a saturated run's cost-per-finding quietly worsens while its output does not improve. Second, **heavy concurrency changes the thing you are measuring.** A target under abnormal load may shed requests, degrade to a smaller served model, trigger platform-side throttling, or route through different capacity than it does under production traffic. A run that hammered the endpoint has tested a condition you did not intend to test, and any conclusion drawn about the deployed system's behaviour is now confounded with the load you generated. On someone else's production system that is also a courtesy and often a contractual problem, not only a measurement one. ### What to check, and what to change instead Diagnose per leg, never in aggregate. Instrument in-flight counts and latency separately for the adversarial chat, the objective target and each scorer; raise concurrency one step and watch which of the three has latency climbing while the others' utilisation falls. That leg is the bottleneck. Total run duration cannot distinguish the three and will send you to tune the wrong one. Then size concurrency to the tightest of the three quotas rather than to the machine's core count, keep the parallel unit the objective rather than the turn, and if the wall clock is genuinely the binding constraint, **cut work rather than raise arrival rate**: fewer converter variants, one scorer instead of two, a shorter turn budget, fewer objectives in this pass. Those reduce cost and time together, which raising concurrency never does. The catch is that they also change what the run covers, so the report must say which knob was turned — a shorter turn budget changes the findings and a higher worker count does not, and a reader who cannot tell the two apart cannot interpret a clean result.
- Why can't you parallelise the turns inside a single multi-turn objective?Because the loop is a feedback loop: the next adversarial prompt is conditioned on the previous response and its score. Turn N+1 does not exist until turn N completes. Parallelism lives across independent objectives, seeds and converter variants.
- How do you identify which of the three legs is the bottleneck?Instrument per-leg latency and in-flight counts separately. Raise concurrency a step: the bottleneck is the leg whose latency rises while the other two show falling utilisation. Aggregate run duration alone cannot distinguish them.
- The wall clock is genuinely unacceptable. What do you change?Reduce work rather than arrival rate — fewer converter variants, a single scorer, a shorter turn budget, or fewer objectives in this pass. Those cut cost and time together, but each changes what the run covers, so state the tradeoff in the report.
Opening more checkout lanes clears the queue faster but does not change the total on the receipts, and it helps only until everyone is waiting on the single card reader they all have to use.
saying these in an interview costs you the question
- Claiming higher concurrency makes the run cheaper.
- Assuming the system under test is the bottleneck without measuring the adversarial and scoring legs.
- Trying to parallelise the turns within one multi-turn objective.
- Ignoring that partially-completed turns still bill the legs that succeeded.