Before launching a PyRIT run, how do you estimate its total call count when the configuration includes several seed prompts, converters and scorers?
answer
- product, not sum
- turns x (1 + V + V*S)
- chained converters multiply
- turn budget is a ceiling, early stop undercuts it
- dry-run one objective, read the transcript store
basics
~20 sMultiply, do not add. Target calls are objectives times turns times converter variants. Scoring calls are target responses times the number of scorers attached. Adding one converter and one scorer roughly doubles two legs at once. Work the product out on paper first, then dry-run one objective to check it.
solid answer
~60 sTreat the configuration as a product of independent multipliers rather than a list of features. - **objectives** — each is its own conversation. - **× turns** — the per-objective turn budget, which is an upper bound; a loop that succeeds early spends less. - **× converter variants** — a converter that rewrites a prompt into N encodings sends N requests to the target for one adversarial turn, and chained converters multiply, they do not add. - **× scorers** — every response is judged by every attached scorer, so two scorers double the scoring leg. The adversarial leg is roughly objectives × turns; the target leg is that times converter fan-out; the scoring leg is the target leg times scorer count. The main tradeoff is that converters and extra scorers are the cheapest-looking config lines in the file and the most expensive at runtime. Where the simple product breaks: early-stopping objectives make it an upper bound, retried turns push it above the bound, and the adversarial leg's growing transcript means token cost is superlinear in turns even when call count is linear.
go deeper
Should at least recognise that converters and scorers multiply the number of calls rather than adding a fixed overhead.
Writes the product out — turns × (1 + variants + variants × scorers) — and knows chained converters multiply.
Separates calls from tokens, notes the transcript growth on the adversarial leg, and validates the estimate against one dry-run objective before scaling.
Turns the estimate into an approval artefact with a stated correction factor, and requires re-estimation whenever the converter chain or scorer set changes.
The estimate that matters has two units, and conflating them is what produces the approval figure that turns out to be an order of magnitude low. **Calls** form a clean product over the configuration. **Tokens** do not, and tokens are what you are billed for. ### The call product Per objective, with `V` = the number of prompt variants a converter chain emits per adversarial turn and `S` = the number of scorers attached to each response: ``` adversarial calls = turns target calls = turns * V # PromptConverter fan-out; chained converters multiply scoring calls = target calls * S # every scorer judges every response total = turns * (1 + V + V*S) ``` Three details in that formula do the damage. **Converters chain multiplicatively, not additively.** A PyRIT converter takes a prompt and returns rewritten prompts; put a 3-variant converter ahead of a 4-variant one and each of the first converter's three outputs is fed through the second, so the objective target receives 12 requests for one adversarial turn, not 7. **Scorers multiply the target leg, not the turn count.** Attaching a second scorer does not add one call per objective; it adds one call per response, and the response count has already been multiplied by `V`. **The adversarial leg is not multiplied at all** — one prompt is composed per turn no matter how many variants it is expanded into afterwards — which is why the `1` in the parenthesis stands alone. Worked: 5 turns, 4 variants, 2 scorers is 5 × (1 + 4 + 8) = 65 calls for **one** objective, and 6,500 for a hundred. Against the naive `turns` figure of 5 per objective that is a 13× multiplier, bought with two config lines that each looked like a feature toggle. ### What it costs beyond the call count Token cost departs from call count in two directions, both upward. The adversarial chat is conditioned on the growing transcript, so its per-turn input climbs with turn number and its cumulative input over an objective grows roughly quadratically in `turns` — call count is linear in turns, spend is not. And converter variants are not the same size as the base prompt: encoding, translation and paraphrase converters can inflate a prompt substantially, so the target leg's tokens are not simply `V` times the original. Wall clock has its own product: turns are sequential within an objective, so `turns` multiplies latency directly while `V` and `S` multiply only work that can be issued in parallel. ### Where the number misleads The product is an **over-estimate** in one specific and common way: `max_turns` is a ceiling, not an expectation. A multi-turn loop stops when the objective scorer says the objective was met, so a high-success pass comes in well under the product. It is an **under-estimate** in several ways that are easier to miss. Transient failures re-bill: a turn that got a prompt generated and a target response and then failed at the scoring step has paid for two thirds of a turn and recorded nothing usable. Retries repeat legs that already succeeded. Any leg you assumed was free but is not — a scorer you believed was a regex and that is in fact model-backed, a converter that is itself LLM-driven rather than deterministic, which several paraphrase-style converters are — inserts a whole multiplier you never wrote into the formula. The worst reading of all is to quote the product as **the** cost. It is a ceiling on calls and a floor on nothing. Presented without the words "assuming every objective runs to its cap", it will be treated as a forecast, and the variance between a high-success and a low-success pass on the same configuration is large enough to make that forecast useless in both directions. ### What to check Compute the product by hand first — it is the only way to notice that a converter chain you thought was additive is a multiplier. Then run **one** objective end to end and count the actual calls out of PyRIT's memory store, which records every request and response piece with the target that issued it, grouped by conversation. Compare observed to predicted; the ratio is your correction factor, and it is what belongs in front of whoever signs off on the full run, alongside an explicit statement of whether the figure assumes every objective exhausts its turn budget. Finally, re-derive the number whenever the converter chain or the scorer set changes, and only then. Those are precisely the edits that read as one line in a diff and move the total by a multiple, and they are the edits most likely to be made by someone who was not in the room when the budget was approved.
- Two converters are chained. Does the target see the sum or the product of their variants?The product. Each variant produced by the first converter is fed through the second, so a 3-variant and a 4-variant converter chained give 12 requests to the target per adversarial turn.
- Why is the computed product usually an over-estimate for a multi-turn run?Because the turn budget is a ceiling, not a target. An objective that reaches its goal on turn two stops there, so runs with a high success rate spend far less than the worst case. Runs that mostly fail spend the full ceiling.
Adding a converter and a second scorer reads like adding two lines to a config file, but it bills like adding two digits to a price tag: each one multiplies everything underneath it instead of adding a fixed amount on top.
saying these in an interview costs you the question
- Adding multipliers instead of multiplying them.
- Treating an extra scorer as a rounding error rather than a whole extra leg.
- Quoting a turn-budget figure as the actual cost with no note that early stopping makes it a ceiling.
- Never dry-running a single objective before authorising a large run.