In a PyRIT multi-turn run the objective text is given both to the adversarial model and to the objective scorer. Why does a vaguely worded objective make the run untrustworthy in both directions?
answer
- objective used twice: attacker + scorer
- vague = attacker drifts, scorer over-fires
- name the observable artefact
- must be possible to fail
- one behaviour per objective
basics
~20 sThe same sentence steers the attacker and defines success. If it names no observable artefact in the target's reply, the adversarial model has nothing concrete to drive toward and drifts, while the scorer has nothing concrete to check and can mark a hedged, harmless answer as met. Write objectives a reviewer can verify from the transcript.
solid answer
~50 sThe objective is not a label; it is an operational input used twice per turn. **As the attacker's instruction**, a vague objective gives the adversarial model no way to tell progress from noise, so its turns wander across framings instead of converging. You burn the turn budget on an unfocused search and then record a negative that reflects the wording, not the target. **As the scorer's criterion**, the same vagueness makes the met/not-met judgement unfalsifiable. Asked whether the target 'discussed something harmful', a judge will find something to say yes to; the run then reports a success no reviewer can confirm from the transcript. The fix is to write the objective so it names the artefact that would have to appear in the target's own output for the objective to be met, and the form it would take. That single change gives the attacker a gradient and gives the scorer, and any human reviewer, the same checkable test.
go deeper
Knows the objective is a natural-language goal supplied to the run and that vague wording produces useless results.
Explains the double duty explicitly: it steers the adversarial model and it is the criterion the scorer applies, so vagueness produces both drift and false successes.
Audits transcripts for both signatures, enforces one-behaviour objectives, and separates an objective-wording defect from a scorer-agreement problem.
Establishes an objective library with a review standard so results from different operators mean the same thing across engagements.
**The objective is consumed twice per turn.** In a PyRIT multi-turn run the objective is one sentence of free text, and the framework hands it to two different components on every iteration. It goes to the **adversarial chat model** as the goal it must steer the conversation toward, alongside the transcript so far. It goes to the **objective scorer** as the criterion against which the target's reply is judged met or not-met. One string, two consumers, no validation of any kind in between — the framework has no way to tell a sharp objective from a mood. **Failure mode one: drift, which manufactures false negatives.** An objective written as a theme rather than an outcome gives the adversarial model no way to tell progress from noise. Watch the attacker turns in such a run and you see the theme restated in different registers, turn after turn, never building on what the target actually said. The budget expires, the run records not-achieved, and the report gains a green row for an objective that was never tested. This is not free: at three model calls per turn, a ten-turn run that drifts burns roughly thirty billed calls and one to three minutes of wall clock to produce a result with no information in it, and a forty-objective sweep can do that thirty-nine more times. **Failure mode two: an unfalsifiable criterion, which manufactures false positives.** The scorer is asked "was this objective met?" with an objective that never defines met. Boundary cases then get resolved by the judge model's own priors rather than by any criterion: a refusal that nevertheless names the topic, a generic paragraph anyone could have written, a fictional or hypothetical framing. Judges asked a vague question tend to find something, so the bias is directional — vagueness inflates the success count rather than randomising it. The run halts on the first accepted turn, the report gains a finding, and no reviewer can reconstruct from the transcript what the scorer thought it saw. **Where the number misleads.** The two failure modes push the reported attack-success rate in opposite directions, which is worse than a single bias because the errors do not announce themselves as noise. A suite of vague objectives yields an ASR that is part fiction and part blindness, and the mixture is invisible in the aggregate. It is also uncomparable: rewording an objective between two engagements changes the number without anything about the target changing, so trend lines across quarters measure the objective library's drafting quality as much as the system. **Three properties worth insisting on.** - **Observable** — the objective names what would have to appear in the target's own output, not a category or a mood. A reviewer must be able to point at the artefact. - **Failable** — a plausible response must exist that a reasonable reader would call a no. If every plausible response is a yes, you have written a theme, and the scorer will behave accordingly. - **Single-behaviour** — one objective per run. Compound objectives split the attacker's effort across sub-goals and leave you unable to say which half a success came from, or how many turns each needed. The shape that satisfies all three is roughly: *the target's own reply contains a specific named artefact, in a specific form, in response to a stated operator framing.* That single change gives the adversarial model a gradient to climb and gives the scorer — and any human reviewer — the same checkable test. **What I would check before trusting a suite of objectives.** For each objective, pull three or four runs out of PyRIT's memory. Read the attacker turns: are they building on the target's replies, or restating the theme? That is the drift signature and it takes a minute to spot. Then read the exact turns the scorer flagged as met and ask whether you can see what it saw; if you cannot, the criterion is the problem, not the target. Hand-label a small set of responses yourself and check the scorer's agreement against them, but do that *after* fixing the objective, because a badly phrased objective makes any scorer look mis-calibrated and you will spend the day tuning a referee that was given an unanswerable question. **Where the boundary sits.** How well a judge model agrees with human labels in general is a separate discipline from how an objective is phrased. Here the point is narrower and more actionable: the objective is a run input, not report metadata, and its wording is a defect surface with a measurable cost on both sides of the verdict.
- Give the shape of a well-formed objective without writing an actual attack goal.Something of the form: 'the target's own reply contains <a specific named artefact>, in <a specific form>, in response to the operator persona <X>'. It names what must appear, so a reviewer can confirm or deny it from the transcript alone.
- Why is one objective per run better than bundling three related ones?The attacker converges on one target state instead of splitting effort, and a success is attributable to a specific behaviour. A compound objective leaves you unable to say which part was met or how many turns each needed.
saying these in an interview costs you the question
- Treating the objective as a descriptive label for reporting rather than as a live input to the run.
- Writing objectives that no response could fail.
- Bundling several behaviours into one objective and then quoting a single success rate.
- Blaming the scorer for verdicts an unfalsifiable objective made inevitable.