A PyRIT run bills an adversarial model, the prompt target and a scorer. If you must cut spend, which of those three can you move to a cheaper or locally-hosted model, and what does each substitution cost you?
answer
- target is not a knob — it is the system under test
- adversarial cut = false-negative bias
- scorer cut = corrupts the number itself
- regression sweep tolerates a weak attacker
- cut work before cutting model quality
basics
~20 sThe target cannot change — it is the system under test, and swapping it invalidates the run. The adversarial model and the scorer can. A weaker adversarial model finds fewer breaches, so results understate risk. A weaker scorer misjudges responses, which corrupts the number the whole run reports.
solid answer
~60 sRank the three legs by what a substitution destroys. **The prompt target is not a knob.** It must be the deployed configuration — same model, same system prompt, same tools, same guardrails. Testing a cheaper stand-in produces a number about a system nobody ships. **The adversarial model is the safest cut, with a known bias.** A cheaper generator produces less inventive prompts, so it fails to find breaches a stronger one would. The error is one-directional: results understate risk, and a clean run becomes ambiguous between "the target held" and "the attacker was weak". Acceptable for regression sweeps over known objectives, poor for discovery. **The scorer is the most dangerous cut**, because it is the leg that produces the number. A weaker judge introduces error in both directions — false hits inflate the report and force triage, missed hits erase real findings — and unlike the adversarial leg you cannot reason about the direction of the bias without labelled transcripts. So the ordering is: cut the adversarial leg first and say so in the report, cut scoring only against a held-out labelled set that shows the cheaper judge agrees, and never cut the target.
go deeper
Should know the target is fixed because it is the system being tested, and that the other two legs are where the savings are.
Explains that a weaker adversarial model biases toward false negatives and that a weaker scorer errs in both directions.
Prefers cutting work over cutting fidelity, requires labelled transcripts before downgrading the judge, and records per-leg model class alongside results so runs stay comparable.
Sets the policy: which engagement types may run a cheap attacker, what evidence licenses a cheaper judge, and how a downgraded pipeline is disclosed to whoever consumes the number.
This is a question about which errors are recoverable, and the three legs sit in a strict order. ### The objective target is not a cost knob PyRIT's `objective_target` is the system under test: a specific deployed configuration — the model version, the system prompt, the tool bindings, the retrieval sources, the input and output guardrails in front of it. The run's claim is about that configuration and nothing else. Swapping in a cheaper sibling model, or the raw model without its guardrail stack, is not a saving; it is a change of subject. The number produced answers a question about a system nobody ships and cannot be carried into a ship decision. When budget genuinely forces the issue, the honest move is **fewer objectives against the real target**, never all objectives against a proxy, because a reduced objective set is a stated coverage gap while a substituted target is a silently invalid result. ### The adversarial chat is the safest cut, with a known direction of error PyRIT's adversarial chat composes each next prompt from the objective and the transcript. A weaker generator plateaus earlier: it repeats phrasings, fails to adapt to a refusal, and burns its turn budget without escalating. The consequence is a **false-negative-biased** run — it under-finds, and it does not mislabel what it does find. That bias is legible, which is what makes the substitution defensible in the right context. It is fine for a regression sweep, where the objectives were already proved reachable by a previous engagement and the question is "is this still broken". It is not fine for discovery over a new surface, because a clean run is then ambiguous between "the target held" and "the attacker was weak" — and every stakeholder will read it as the first. ### The objective scorer is the most dangerous cut PyRIT's `objective_scorer` is what converts transcripts into the run's headline number: it decides whether each response counted as a hit, and therefore fixes the per-objective success rate, the ranking of which surfaces are worst, and the contents of the triage queue. Scoring error propagates into all of it, and unlike the adversarial leg it errs in **both** directions: false hits inflate the report and consume triage hours on nothing, missed hits erase real findings and are never noticed. Worse, a cheaper judge degrades hardest exactly where judgement is hard — partial compliance, hedged or in-character refusals, responses that supply the shape of an answer without the substance — which is precisely where the interesting findings live. On obvious hits and obvious refusals every judge agrees, which is why a naive agreement check looks reassuring and proves nothing. ### What it costs, and what it costs you The saving is real: on a long objective the adversarial leg's input grows with the transcript, so it is often the largest single line on the bill, and moving it down a model tier can halve an engagement's spend. The scorer is cheap per call but numerous — one call per response per scorer, on a response count already multiplied by converter fan-out — so downgrading it saves meaningfully too. What it costs you is comparability. A number produced by a downgraded pipeline is not comparable to last quarter's, and a trend line drawn across a silent substitution is worse than no trend line. ### The prerequisite before downgrading the judge A held-out set of labelled transcripts from a prior run. Score them with both the incumbent and the candidate judge and compare agreement **on the disputed middle** — the partial and hedged cases — reporting that separately from overall agreement, which the easy cases will dominate. Without that evidence you have not saved money; you have made the output unfalsifiable, which is the one outcome a red-team run must never produce. ### The cheaper structural option, and what to record Before downgrading any leg, reduce **work**: fewer converter variants, one scorer instead of two, a shorter turn budget, fewer objectives per pass. These cut spend without changing the fidelity of any leg. They change coverage instead — and a coverage reduction is a sentence you can write honestly in a report, whereas a quietly weakened judge is a defect in the number itself. Whatever you choose, record per-leg model class alongside the results: which model generated, which answered, which judged, and at what version. That line is what lets the next reader tell a resilient system from an under-funded run, and it is the line most often missing.
- When is a cheap adversarial model actually the right choice?For regression sweeps over objectives a previous engagement already proved reachable. The question is 'is this still broken', which a weaker generator can answer. For discovery on a new surface it is the wrong choice, because a clean result would be misread as safety.
- What do you need before downgrading the scorer?A held-out set of labelled transcripts from a prior run, and evidence that the cheaper judge agrees with the labels on the ambiguous cases — partial compliance and hedged answers — not just on the obvious hits and obvious refusals.
- Budget is halved and none of the three can be downgraded. What now?Cut work instead of quality: fewer objectives, a shorter turn budget, fewer converter variants, a single scorer. That reduces coverage, which is a statement you can put in the report, rather than reducing fidelity, which quietly changes what the number means.
The scorer is the ruler, not the thing being measured. A cheaper target measures the wrong object; a cheaper ruler leaves you unable to say whether anything you already wrote down was the right length.
saying these in an interview costs you the question
- Testing a cheaper stand-in and reporting the result as being about the deployed system.
- Downgrading the scorer with no labelled transcripts to check agreement.
- Reporting a clean run from a weakened adversarial model without noting the false-negative bias.
- Not recording which legs ran on which class of model, then comparing the number to a prior run.