In PyRIT, an attack strategy object is constructed once with its prompt target and its scorer, while the objective is supplied each time you execute it. What does that split buy you when you have twenty objectives to test against the same endpoint, and what does each execution get of its own?
answer
- strategy built once, objective per execution
- target + scorer + converters = configuration
- one conversation per execution
- loop objectives, not strategies
- objective steers, scorer decides
basics
~20 sYou build the strategy once with its target, scorer and any converters, then execute it in a loop, once per objective. The wiring is shared across all twenty. Each execution carries its own objective and its own conversation, so the transcript and the verdict stay separate per objective.
solid answer
~50 sThe strategy object holds the **run configuration** — which prompt target it sends to, which scorer decides whether an attempt counted, any converters, and the turn cap. That configuration is stable across a campaign, so you construct it once and reuse the instance. The **objective** is per-execution input, not configuration. Each execution starts its own conversation, drives it under that shared wiring, and returns its own result with its own transcript identifier. So a twenty-objective sweep is a loop over objectives against one configured strategy, not twenty configured strategies. The tradeoff is that anything you want to vary *per objective* — a different scorer rubric, a different target, a different converter chain — is not per-execution input, so it forces a second configured strategy. Candidates who put objective-specific scoring into the objective string get a scorer that was never told what to look for.
go deeper
Says the strategy is configured once and executed per objective, and that each execution has its own conversation and result.
Adds which collaborators are constructor-time (target, scorer, converters, turn cap) versus what is execution-time input, and why a per-objective rubric forces a second strategy.
Talks about comparability across a sweep, transcript identity per execution, and how to spot a loop that ended early on an exception rather than returning results.
Frames it as campaign structure — what varies per run must be visible in the results schema, so triage and re-runs can attribute an outcome to one variable.
### What the objects are A PyRIT run is assembled from a few collaborating objects, and the question only makes sense once each one is named. A **prompt target** is the adapter that knows how to reach one endpoint — an API model deployment, a hosted chat app, some other HTTP surface — and return its reply. A **scorer** reads a reply and returns a verdict: a true/false judgement, a scaled number, or a classifier label. A **converter** rewrites an outgoing prompt before it is sent. An **attack strategy** is the driver that owns the send-and-score loop, and it is handed those collaborators when you build it. ### The split: configuration versus per-run input Target, scorer, converter chain and the turn budget are constructor-time arguments. Together they describe *how* you are testing: which endpoint, which judgement, which transforms, how many exchanges you are willing to pay for. The **objective** — a natural-language statement of what you want the target to end up doing — is supplied when you execute the strategy. It describes *what* this particular run is chasing. So a twenty-objective campaign is a loop over twenty strings against one constructed instance, not twenty constructed strategies. Each execution opens its own conversation, drives it under the shared wiring, and returns its own result carrying its own conversation identifier, so the twenty transcripts stay separately readable and separately re-scorable afterwards. What that buys you is **attributability**. Across the sweep exactly one variable moved. If objective seven behaves differently from objective eight, the difference is the objective — not a scorer you happened to re-instantiate with a different threshold, or a target you repointed at another deployment while editing. Constructing the strategy inside the loop is not wrong in itself; it is wrong because it invites two things to change at once and leaves you unable to say which one mattered. ### What it costs Object construction costs nothing worth counting. The money is in calls. A single-send sweep of twenty objectives is twenty target calls plus twenty scoring decisions — and if the scorer is model-backed, which is the normal case for anything more nuanced than a substring match, that is **forty metered calls, not twenty**. The scorer's call is not the cheap half either: it is billed on the prompt plus the whole response it is judging. Add an LLM-backed converter link and each of the twenty prompts costs a further call before it is even sent. Wall-clock runs the other way: single-send executions are independent, so they parallelise up to whatever concurrency the endpoint tolerates, and the practical ceiling is the provider's rate limit rather than the loop itself. ### Where the number misleads The headline this structure produces is “twenty objectives tested”, and that is a *count of executions*, not a measure of coverage. Three specific misreadings: - **The objective is not the rubric.** People write success criteria into the objective text and pair it with a generic scorer. The objective steers the attacker side of the run; the scorer alone decides the verdict. If the two describe different things, every verdict in the sweep is judging something nobody asked about — and the run still reports twenty clean results. - **Twenty objectives is not twenty behaviours.** If the list was produced by paraphrasing a handful of seeds, the denominators are near-duplicates and a single guardrail rule can generate twenty identical outcomes that read as twenty independent pieces of evidence. - **A missing result reads as a pass.** A loop that swallows an exception mid-sweep still prints the objectives it reached, and the count quietly comes back as seventeen. ### What I would check That the number of results equals the number of objectives submitted. That each result carries a distinct conversation identifier and that the transcript can actually be re-read from the memory store rather than only summarised. That the scorer named in the results is the one configured on the strategy, not a default picked up somewhere. And that a human read a sample of transcripts instead of trusting the verdict field. If a rubric genuinely differs for a subset of objectives, that is a second configured strategy with a second scorer — and the report must then say which subset ran under which configuration, because those two subsets are not comparable to each other.
- You need a different success rubric for five of your twenty objectives. Where does that change go?Into a second configured strategy with a different scorer. The rubric is configuration, not per-execution input, so it cannot ride along in the objective text.
- Two executions of the same configured strategy return the same verdict. How do you tell them apart afterwards?By the conversation each execution created — each run's transcript is stored separately and identified, so you re-read the actual exchange rather than trusting the verdict alone.
- Does reusing one strategy instance make the second run aware of the first run's exchange?No. Each execution stands as its own conversation unless you deliberately point the run at earlier history.
saying these in an interview costs you the question
- Believes the objective configures the scorer, so a well-worded objective is enough to make a verdict trustworthy.
- Constructs a fresh strategy per objective and cannot say what that changes, suggesting the object model is opaque to them.
- Assumes reusing one strategy instance continues the previous conversation.
- Cannot name a single collaborator the strategy is configured with.