In PyRIT, what does a prompt target object do, and why does one run usually need more than one of them?
answer
- target = the only object that knows the endpoint
- send prompt, return response, write to memory
- three roles: system under test, attacker, scorer
- one turn can be several billed calls
- shared target = attacker inherits refusals
basics
~20 sA PyRIT prompt target is the adapter that knows one endpoint: it sends a prompt there and returns the response. A run usually needs several of them, because the system under test, the attacker model that drafts prompts, and any model that scores replies are three different endpoints.
solid answer
~50 sThe prompt target is the only object in a PyRIT run that knows a concrete endpoint - its address, credentials, request shape and how to pull text out of the reply. Everything else (the attack strategy, the converters, the scorers) talks to it through the same send-and-return contract, which is why the same kind of object can play three roles in one run: the system under test, the adversarial model that writes the next attack turn, and the model a scorer asks for a verdict. Keeping those as separate target objects matters, because they normally differ in credentials, cost and safety configuration. You typically point the attacker role at a deployment permissive enough to draft adversarial text, while the system under test is the guarded production endpoint. Reuse one object for both and the attacker inherits the guarded endpoint's refusals: the run stalls, and the log reads like a robust target rather than a broken harness.
go deeper
Says the target is the thing that sends prompts to a model or endpoint and brings back the answer, and that the attacker model and the model under test are configured separately.
Explains the uniform send-and-return contract, names the three roles a target can play in one run, and notes the per-turn call multiplication and the credential/quota separation.
Diagnoses a zero-hit run from stored conversations, separates quotas per role, and treats an inherited refusal from a shared attacker deployment as a harness defect rather than a result.
Sets the standard for how runs are wired across the team - which deployments are approved for the attacker and judge roles, whose budget each burns, and what a report may claim when the roles shared infrastructure.
**What the object actually is.** PyRIT (the Python Risk Identification Toolkit) splits a red-team run into a few collaborating objects: an *attack strategy* that decides what to send and when to stop, *prompt converters* that rewrite a prompt before it leaves, *scorers* that judge each reply, a *memory* store that records everything, and *prompt targets*. The prompt target is the only one of those that knows the outside world exists. Its contract is deliberately narrow: accept a prompt request, deliver it to one concrete endpoint, return the response, and write both pieces to memory under a conversation id. Everything endpoint-specific lives inside the target and nowhere else - base URL, deployment or model name, the API key or bearer token, the request body shape, which field of the reply holds the assistant's text, and the retry and rate-limit behaviour. Concrete targets differ only in those details. `OpenAIChatTarget` speaks a chat-completions API. `HTTPTarget` takes a raw HTTP request template with a placeholder where the prompt is substituted, plus a callback that parses the reply body. Others cover non-chat surfaces - image generation, speech. Because they all satisfy the same send-and-return contract, a strategy written against "a target" runs unchanged against any of them: repointing a whole suite from a hosted deployment to a locally served open-weights model is a constructor change, not a rewrite. That is the payoff of keeping the boundary narrow. **Why one run needs several of them.** A red-team run is not one endpoint, it is a small fleet of them, and each one is a target object: | role | what it is | typical wiring | |---|---|---| | system under test | the thing you are measuring | the target passed to the attack strategy | | adversarial model | drafts the next attack turn in an automated multi-turn strategy | a second, separate target, usually a permissive deployment | | scoring model | a model-backed scorer asking "did this reply do the thing?" | a third target held by the scorer | | converter backend | any converter that rewrites a prompt with a model (translate, restyle, vary) | a fourth target held by that converter | These are normally different deployments with different credentials, different content filtering, and different quotas. The framework does not object if you pass the same object to all of them; nothing in the type system says you must not. **What it costs.** Do the arithmetic before you launch. One turn of an automated multi-turn strategy with a model-backed scorer is at minimum three billed calls: the attacker drafts, the system under test answers, the judge scores. Add one LLM-backed converter link and it is four. So a ten-turn strategy over twenty objectives is not 200 calls, it is 600 to 800, spread across three separate quotas - and the two calls you are *not* measuring are often the larger share of the bill, because attacker and judge prompts carry long instructions. Wall-clock follows the tightest quota, not the average one. Engineer time goes on wiring credentials for three deployments and on getting approval for a permissive attacker deployment, which in most organisations is the slow part. **Where the number misleads.** The failure that produces a *believable wrong answer* is role sharing. Point the attacker role at the same guarded production deployment as the system under test and the attacker inherits its refusals: it declines to draft the next adversarial turn, so what actually gets sent is an apology or an empty string, the target answers harmlessly, the scorer records a miss, and the run reports "0 of 200 objectives achieved." That reads as a strong defence and is in fact a broken harness. The same shape appears with a shared judge: a judge that shares the target's blind spots agrees with it, and a single outage across the shared endpoint scores as perfect safety. Shared *quota* is the subtler cousin. Scorer and converter traffic throttle the attack, retries stretch the run, and any latency or throughput figure you report then describes your own rate limit rather than the target's resilience. Note also what a target does *not* do: it returns text. It has no opinion on whether that text is a refusal, an error string in a 200 body, or a violation. Every safety judgement happens downstream of it. **What I check before believing a run.** Which target object fills which role, and that each holds its own client and credentials. A direct smoke test of the attacker target - one benign prompt, one that should be refused - reading the stored replies rather than trusting the constructor. Separate rate limits per role, with backoff configured so throttling degrades the run instead of recording failed attempts. A hand-read sample of stored conversations before I look at any aggregate: real adversarial prompts and real refusals mean a robust endpoint, while the attacker declining or empty prompt bodies mean a harness defect. And which memory store is in use, because an in-memory store leaves nothing to re-read or re-score once the process exits.
- If a PyRIT run reports zero successful attempts, how do you tell a genuinely robust endpoint from a misconfigured attacker target?Read the stored conversations. A robust endpoint shows real adversarial prompts and refusals from the system under test; a broken harness shows the attacker itself declining, empty prompts, or transport errors on every turn.
- Why does a model-backed converter also need a target?Because rewriting a prompt with a model is itself an endpoint call. That converter has to be pointed at a deployment, and each link costs a call per prompt on top of the attack and scoring traffic.
Wiring a run is like hiring three contractors who each bill by the hour: one writes the letters, one receives them, one reads the replies - and you only ever look at the invoice from the one you were testing. If you hire the same cautious person for all three jobs, the letters never get written and the file comes back empty.
saying these in an interview costs you the question
- Believing one target object serves the whole run and the framework figures out the rest
- Treating a run with zero hits as evidence of a safe endpoint without reading the stored conversations
- Not knowing that a model-backed scorer or converter also spends endpoint calls
- Pointing the attacker role at the same guarded production deployment as the system under test and being surprised by refusals