skip to content

In PyRIT, some prompt converters transform text locally while others call a model to rewrite the prompt. If a converter chain with two model-backed links runs over a 200-prompt seed set, what does that add to the run's billed calls, and what is already metered besides the converters?

level: middleimportance: should knowfreq 48%

answer

  1. k model-backed links = k calls per prompt
  2. converter, target, scorer, attacker
  3. variants x turns x surfaces
  4. convert once and cache
  5. throttled on a surface not under test

basics

~20 s

Each model-backed link bills one call per prompt, so two links over 200 prompts add roughly 400 calls before anything is sent anywhere. On top of that the run still pays the target call and the scorer call per attempt, plus the attacker model's calls if the strategy uses one.

solid answer

~50 s

Local converters are free string work. A model-backed converter is a full inference request per prompt, and because the chain is a series, two such links means two calls for every single prompt you send: 200 seeds becomes about 400 converter calls, and more if a link is configured to emit several variants or if retries fire. The part candidates miss is that converters are not the only meter. A red-team run pays on several surfaces per attempt: the converter chain, the target itself, the scorer that judges the reply, and — in a multi-turn strategy — an attacker model generating the next turn. Those multiply, they do not add: variants x turns x metered surfaces. So before a full sweep, price one prompt end to end, then multiply. If the converter rewrite is deterministic enough to be reusable, generate the converted corpus once and re-send it rather than re-converting on every run.

go deeper

for a junior

Know that some converters call a model and therefore cost money and latency, while others are plain local string work.

for a middle

Do the multiplication: one call per model-backed link per prompt, plus the target and scorer calls, and note that variants and turns multiply it further.

for a senior

Add the operational consequences — caching converted corpora, rate limits on the converter's own endpoint, and never letting a failed rewrite count in the success denominator.

for a principal

Own the budget question: which of the metered surfaces the team is allowed to spend on, and whether a non-deterministic rewrite makes a standing suite's trend line meaningless.

### Do the arithmetic out loud For **one** prompt travelling through a chain with *k* model-backed links, the metered surfaces are: - *k* converter calls — one per link configured with a `converter_target` - 1 target call — the system actually under test - 1 scorer call, if the scorer is model-backed (PyRIT's self-ask family) rather than a local string check - plus 1 attacker-model call per turn, if the attack strategy generates the next turn with a model For the question as posed — 200 seeds, two model-backed links, single-send strategy, model-backed scorer: | surface | calls per prompt | over 200 prompts | |---|---|---| | model-backed converters (k = 2) | 2 | 400 | | prompt target | 1 | 200 | | model-backed scorer | 1 | 200 | | **total** | **4** | **800** | Three of those four surfaces are yours, not the customer's. Only the 200 target calls touch the system you were hired to assess; the other 600 are the cost of asking the question. ### Then multiply, because that is what actually bites The seed count is the least interesting number in a real suite. Fan the same corpus across 4 converter variants and the whole table quadruples. Run a 5-turn multi-turn strategy and the converters, target, scorer and attacker model all fire per turn. 200 seeds x 4 variants x 5 turns x 4 surfaces is on the order of 16,000 billed calls for one sweep — from a config that a status update will describe as "we ran 200 prompts". ### Wall clock is a different constraint from the bill PyRIT sends with bounded concurrency (a batch size or max-concurrency argument on the send call). Raising it shortens wall clock and changes the bill not at all; what it does change is which endpoint you saturate first. The converter's `converter_target` and the scorer are frequently the *same provider account*, often with a smaller quota than the target you are testing, so a run can throttle itself on a surface that is not under test. The symptom is a slow run with a completely idle target rate limit, and the wrong diagnosis is "the target is rate-limiting us". ### Where the numbers mislead - **Sizing by seeds.** "200 prompts" understates spend roughly fourfold here, and by an order of magnitude once variants and turns are in. - **Denominator rot.** A converter call that errors or comes back as a refusal to rewrite is a *run* failure. Fold it into the attempt count and it reads as an attempt the target survived, which quietly deflates the success rate and makes the system look safer than the evidence supports. - **A trend line made of noise.** If a model-backed rewrite produces a materially different string on every run, the suite is not a regression suite; it is a fresh experiment each time, and quarter-over-quarter comparison of its numbers is meaningless. - **Latency attributed to the wrong system.** As above: the harness's own endpoints, not the target's. ### The cheap fix, and what it costs you If a rewrite does not need to be fresh, convert once and store the converted corpus, then re-send the stored strings. Per-run cost drops from four surfaces to two (target plus scorer), and — more valuable — the input set becomes stable, which is the precondition for comparing this month's number to last month's. What you give up is per-run novelty, so keep a separate, smaller exploratory lane where the rewrite is regenerated on purpose. ### What I would check before the sweep Dry-run the chain over five seeds with the target stubbed, then count the requests your provider account actually logged and compare with the predicted `k x n`. It is common to find retries, or a link fanning out into several variants, that the config did not obviously imply. Run those same five seeds twice and diff the stored `converted_value`s: identical means you can cache and compare across runs, divergent means you have to say so next to every trend you publish. Finally, price one prompt end to end and multiply by the real matrix before anyone approves the budget — not after the invoice arrives.

  • How would you cut the converter spend without losing the test?
    Convert the corpus once and store it, then re-send the stored converted prompts. Cost drops to the target plus scorer calls, and the input set becomes stable across runs.
  • The run is slow but the target's rate limit is untouched. Where do you look?
    The converter and scorer endpoints. They are often the same provider account, so the harness throttles itself on a surface that is not the system under test.

Sizing a converter run by its seed count is like pricing a road trip by the number of destinations while ignoring that every leg passes four toll booths. Three of those booths are your own harness — the rewrite model, the scorer, the attacker model — and only one is the system you were paid to test.

saying these in an interview costs you the question

  • Assuming every converter is a free local transform.
  • Sizing the run by seed count alone.
  • Ignoring that the scorer and attacker model are separately metered.
  • Counting errored or refused conversions as attempts that the target survived.

context