skip to content

Execution Objects

One run is an attack strategy driving a target through converters into a scorer, and the object you picked wrong caps what the run can ever find. Interviewers probe which one you had to replace.

on this pageshow

explore

questions

23

In PyRIT, an attack strategy object is constructed once with its prompt target and its scorer, while the objective is supplied each time you execute it. What does that split buy you when you have twenty objectives to test against the same endpoint, and what does each execution get of its own?

level: juniorimportance: must knowfreq 62%

answer

  1. strategy built once, objective per execution
  2. target + scorer + converters = configuration
  3. one conversation per execution
  4. loop objectives, not strategies
  5. objective steers, scorer decides

basics

~20 s

You build the strategy once with its target, scorer and any converters, then execute it in a loop, once per objective. The wiring is shared across all twenty. Each execution carries its own objective and its own conversation, so the transcript and the verdict stay separate per objective.

solid answer

~50 s

The strategy object holds the **run configuration** — which prompt target it sends to, which scorer decides whether an attempt counted, any converters, and the turn cap. That configuration is stable across a campaign, so you construct it once and reuse the instance. The **objective** is per-execution input, not configuration. Each execution starts its own conversation, drives it under that shared wiring, and returns its own result with its own transcript identifier. So a twenty-objective sweep is a loop over objectives against one configured strategy, not twenty configured strategies. The tradeoff is that anything you want to vary *per objective* — a different scorer rubric, a different target, a different converter chain — is not per-execution input, so it forces a second configured strategy. Candidates who put objective-specific scoring into the objective string get a scorer that was never told what to look for.

go deeper

for a junior

Says the strategy is configured once and executed per objective, and that each execution has its own conversation and result.

for a middle

Adds which collaborators are constructor-time (target, scorer, converters, turn cap) versus what is execution-time input, and why a per-objective rubric forces a second strategy.

for a senior

Talks about comparability across a sweep, transcript identity per execution, and how to spot a loop that ended early on an exception rather than returning results.

for a principal

Frames it as campaign structure — what varies per run must be visible in the results schema, so triage and re-runs can attribute an outcome to one variable.

### What the objects are A PyRIT run is assembled from a few collaborating objects, and the question only makes sense once each one is named. A **prompt target** is the adapter that knows how to reach one endpoint — an API model deployment, a hosted chat app, some other HTTP surface — and return its reply. A **scorer** reads a reply and returns a verdict: a true/false judgement, a scaled number, or a classifier label. A **converter** rewrites an outgoing prompt before it is sent. An **attack strategy** is the driver that owns the send-and-score loop, and it is handed those collaborators when you build it. ### The split: configuration versus per-run input Target, scorer, converter chain and the turn budget are constructor-time arguments. Together they describe *how* you are testing: which endpoint, which judgement, which transforms, how many exchanges you are willing to pay for. The **objective** — a natural-language statement of what you want the target to end up doing — is supplied when you execute the strategy. It describes *what* this particular run is chasing. So a twenty-objective campaign is a loop over twenty strings against one constructed instance, not twenty constructed strategies. Each execution opens its own conversation, drives it under the shared wiring, and returns its own result carrying its own conversation identifier, so the twenty transcripts stay separately readable and separately re-scorable afterwards. What that buys you is **attributability**. Across the sweep exactly one variable moved. If objective seven behaves differently from objective eight, the difference is the objective — not a scorer you happened to re-instantiate with a different threshold, or a target you repointed at another deployment while editing. Constructing the strategy inside the loop is not wrong in itself; it is wrong because it invites two things to change at once and leaves you unable to say which one mattered. ### What it costs Object construction costs nothing worth counting. The money is in calls. A single-send sweep of twenty objectives is twenty target calls plus twenty scoring decisions — and if the scorer is model-backed, which is the normal case for anything more nuanced than a substring match, that is **forty metered calls, not twenty**. The scorer's call is not the cheap half either: it is billed on the prompt plus the whole response it is judging. Add an LLM-backed converter link and each of the twenty prompts costs a further call before it is even sent. Wall-clock runs the other way: single-send executions are independent, so they parallelise up to whatever concurrency the endpoint tolerates, and the practical ceiling is the provider's rate limit rather than the loop itself. ### Where the number misleads The headline this structure produces is “twenty objectives tested”, and that is a *count of executions*, not a measure of coverage. Three specific misreadings: - **The objective is not the rubric.** People write success criteria into the objective text and pair it with a generic scorer. The objective steers the attacker side of the run; the scorer alone decides the verdict. If the two describe different things, every verdict in the sweep is judging something nobody asked about — and the run still reports twenty clean results. - **Twenty objectives is not twenty behaviours.** If the list was produced by paraphrasing a handful of seeds, the denominators are near-duplicates and a single guardrail rule can generate twenty identical outcomes that read as twenty independent pieces of evidence. - **A missing result reads as a pass.** A loop that swallows an exception mid-sweep still prints the objectives it reached, and the count quietly comes back as seventeen. ### What I would check That the number of results equals the number of objectives submitted. That each result carries a distinct conversation identifier and that the transcript can actually be re-read from the memory store rather than only summarised. That the scorer named in the results is the one configured on the strategy, not a default picked up somewhere. And that a human read a sample of transcripts instead of trusting the verdict field. If a rubric genuinely differs for a subset of objectives, that is a second configured strategy with a second scorer — and the report must then say which subset ran under which configuration, because those two subsets are not comparable to each other.

  • You need a different success rubric for five of your twenty objectives. Where does that change go?
    Into a second configured strategy with a different scorer. The rubric is configuration, not per-execution input, so it cannot ride along in the objective text.
  • Two executions of the same configured strategy return the same verdict. How do you tell them apart afterwards?
    By the conversation each execution created — each run's transcript is stored separately and identified, so you re-read the actual exchange rather than trusting the verdict alone.
  • Does reusing one strategy instance make the second run aware of the first run's exchange?
    No. Each execution stands as its own conversation unless you deliberately point the run at earlier history.

saying these in an interview costs you the question

  • Believes the objective configures the scorer, so a well-worded objective is enough to make a verdict trustworthy.
  • Constructs a fresh strategy per objective and cannot say what that changes, suggesting the object model is opaque to them.
  • Assumes reusing one strategy instance continues the previous conversation.
  • Cannot name a single collaborator the strategy is configured with.

context

open as a page

In PyRIT, what does a prompt converter do to a prompt before it reaches the target, and why is adding a converter on its own not an attack?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A PyRIT prompt converter rewrites the outgoing prompt text after the seed prompt is chosen and before it is sent. It changes the surface form, not the request underneath. On its own it is just a transform: it tests whether that form gets through, never whether the model will actually comply.

open as a page

In PyRIT, what does the conversation memory store record while a run executes, and what do you give up by running with the in-memory store instead of the durable one?

level: juniorimportance: must knowfreq 62%

basics

~20 s

PyRIT's memory records every prompt sent and every response received, turn by turn, with the scores attached, so a run can be resumed, audited and re-scored later. The in-memory option keeps that only for the process lifetime: when it exits, the transcripts are gone and nothing can be re-read or re-scored.

open as a page

In PyRIT, what is a scorer's job during an attack run, and what does its verdict do to the run itself?

level: juniorimportance: must knowfreq 72%

basics

~20 s

In PyRIT, the scorer reads the target's response and decides whether the attack objective was met. In a multi-turn run that verdict is control flow: a positive verdict ends the run early and marks it a success. So the scorer, not you, decides when testing stops.

open as a page

In PyRIT, what does a prompt target object do, and why does one run usually need more than one of them?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A PyRIT prompt target is the adapter that knows one endpoint: it sends a prompt there and returns the response. A run usually needs several of them, because the system under test, the attacker model that drafts prompts, and any model that scores replies are three different endpoints.

open as a page

In PyRIT, what distinguishes a single-send attack strategy from a multi-turn one, and what actually happens when you point a single-send strategy at an objective that only succeeds after several turns?

level: middleimportance: must knowfreq 70%

basics

~20 s

A single-send strategy delivers each prompt once, scores the reply and stops; nothing feeds forward. A multi-turn strategy loops: it reads the last reply, composes the next prompt with an attacker model, and scores each turn until the scorer says achieved or the turn cap runs out. Single-send on a multi-turn goal returns not-achieved, having proved nothing.

open as a page

A PyRIT run is configured with several prompt converters applied to each outgoing prompt. In what order does the chain apply them, and why does swapping two converters change what the target actually receives?

level: middleimportance: must knowfreq 62%

basics

~20 s

The chain runs in series over a single prompt: the first converter's output becomes the second's input, and only the final string is sent. So the converters compose, and composition is not commutative — rewriting an already-transformed string gives a different result than transforming a rewritten one.

open as a page

By default a PyRIT run persists its conversations to an unencrypted local database file. What is actually in that file, and how should it change where and how you run an engagement?

level: middleimportance: must knowfreq 58%

basics

~20 s

By default PyRIT writes the run to a local database file on the machine you launched it from, unencrypted. That file holds attack prompts and the target's worst answers verbatim. Treat it as sensitive evidence: put it on encrypted storage you control, keep it off shared drives and backups, and delete it on schedule.

open as a page

A PyRIT scorer can return a true/false verdict, a scaled numeric score, or a category label. How do you choose between them for a multi-turn attack run, and what extra decision does a scaled scorer force on you?

level: middleimportance: must knowfreq 62%

basics

~20 s

A true/false scorer gives the loop the yes-or-no it needs to stop, so it is the default for objective-driven runs. A scaled scorer returns a degree of compliance and forces you to pick the threshold that counts as success. A category scorer says what kind of harm, not whether the attack worked.

open as a page

In PyRIT, how does earlier conversation history reach the endpoint on turn five of a multi-turn run, and what changes when the prompt target wraps a service that keeps its own server-side session?

level: middleimportance: must knowfreq 58%

basics

~20 s

For a chat-completion style endpoint the target replays the stored turns: it reads the conversation out of memory and sends the whole message list each call. If the service keeps its own session and accepts only the newest message, the history lives on the server, so the tool's copy and the real context can drift apart.

open as a page

Adding an encoding converter to a PyRIT chain triples the success count over the same seed prompts. Before you report that as a jailbreak result, how do you establish whether the target actually complied?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Treat the jump as a scoring hypothesis first. Pull the stored exchanges, decode the replies, and read whether the content is really there. A transform that changes the reply's form also changes what a text-based PyRIT scorer sees, so refusals stop matching and garbled output can read as compliance.

open as a page

Before you let a PyRIT scorer end multi-turn runs on its own, how do you check it against transcripts you labelled yourself, and which responses have to be in that labelled set?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Pull stored conversations from earlier runs, label the responses yourself, then re-score them offline with the candidate scorer and compare. Load the set with the hard cases: refusal-then-comply, partial compliance, in-fiction compliance, and confident nonsense. Only then let it stop runs. Replaying stored text costs no target turns.

open as a page

A batch of PyRIT multi-turn runs reports a high objective-achieved rate, but most transcripts stop on an early turn at a hedged, non-compliant answer. What is happening, and how do the two directions of scorer error differ in what they cost you?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The scorer is firing on responses that are not real compliance, and because a positive verdict ends the run, every false hit is also a truncated attack. False positives inflate the success rate and delete the turns you never sent. False negatives cost the whole turn budget and undercount. Read the stop-turn responses.

open as a page

A multi-turn PyRIT attack run finishes with the objective not achieved. What are the ways that loop can terminate, and why is that outcome not evidence that the target is safe?

level: seniorimportance: should knowfreq 52%

basics

~20 s

It ends three ways: the scorer returns an achieved verdict and the loop stops early, the turn cap is exhausted, or something raised — target error, rate limiting, the attacker side declining to continue. Not-achieved conflates all of the non-success endings, so it says the run stopped, not that the target held.

open as a page

An application team disputes one item in your report from a PyRIT engagement. How do you get from that report item back to the exact stored exchange that produced it, and what does the transcript prove and not prove?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Each stored turn carries identifiers tying it to its conversation and run, so a report item can point at the exact exchange that produced it. Keep those identifiers in the finding. The transcript proves what happened once against that configuration - not that it reproduces, since the target samples and may have changed.

open as a page

A finished PyRIT engagement is stored, and you now want to relabel it with a stricter judgment of what counts as a hit without spending another call against the target. What makes that possible, and what can relabelling stored transcripts not tell you?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Because the run stored every prompt and response, you can point a different scorer at the saved transcripts and relabel them without spending another target call. What that cannot tell you is what the attack would have done under the new judgment: the strategy branched on the old verdicts, so the turns it never explored simply do not exist.

open as a page

Your only access to a support assistant is an authenticated HTTP API with a bespoke JSON body, a rotating session token and a streamed reply. How do you get it under PyRIT as a prompt target, and what does that wrapper silently bound about the run?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Use PyRIT's generic HTTP target: give it the request template with a placeholder where the prompt goes, plus logic that pulls the assistant text out of the reply. You then own auth refresh, retries and rate limiting yourself. Anything the API never returns - system prompt, tool calls, retrieved context - stays untested.

open as a page

You have a fixed call allowance against a metered chat endpoint for one PyRIT engagement. How do you split it between many single-send runs across a broad objective list and a few deep multi-turn runs, and how do you choose the per-run turn cap?

level: principalimportance: should knowfreq 34%

basics

~20 s

Spend the first slice on broad single-send runs to find where the target is soft, then concentrate deep multi-turn runs on those objectives. Set the turn cap from observed successes, not from a guess, and reserve a slice for repeats, because one run against a non-deterministic endpoint is a sample rather than a result.

open as a page

Your team's standing PyRIT suite runs every seed prompt through a matrix of prompt-converter variants. How do you decide how many variants the suite carries, and what does a per-converter success number actually measure when you publish it?

level: principalimportance: should knowfreq 33%

basics

~20 s

Size the matrix by distinct transform classes, not by count: variants multiply target and scorer calls linearly and mostly re-test the same input path. A per-converter number measures which surface form got past the input handling and the judge — not distinct model weaknesses, and not attack-surface coverage.

open as a page

Your team is scaling from one operator running PyRIT locally to several running in parallel on the same engagement. How do you decide between per-operator local stores and one shared conversation store, and what does the shared choice oblige you to build?

level: principalimportance: should knowfreq 33%

basics

~20 s

A local per-operator store is simplest and keeps harmful transcripts on one controlled machine. A shared store lets a team resume each other's runs, deduplicate findings and report across the whole engagement, but concentrates every operator's harmful text in one place that now needs access control, retention rules and a named owner.

open as a page

Your team runs PyRIT campaigns against several internal endpoints on a recurring schedule. What should the scorer be allowed to decide on its own, and what has to reach a human before it counts?

level: principalimportance: should knowfreq 40%

basics

~20 s

Let the scorer decide only inside the run: when to stop a conversation and how to rank what gets read first. Anything leaving the team as a finding needs a human who read the transcript. Keep a fixed labelled transcript set as the scorer's own regression test, re-run whenever the scorer changes.

open as a page

For a PyRIT engagement you can point the prompt target at the raw model endpoint behind a product or at the product's own chat API. How do you decide which surface to wrap, and how does that choice change what your report may claim?

level: principalimportance: should knowfreq 34%

basics

~20 s

Decide from the question being asked. The raw endpoint measures the model's own behaviour with no product defences; the product API measures what a user can actually reach, defences included. Each report claims only its own surface. Where budget allows, wrap both and read the gap between them as the value of the defences.

open as a page