skip to content

Prompt Targets

The target is the only object that knows your endpoint, so what it cannot expose — turn history, a system prompt, a non-chat surface — bounds every strategy. Interviewers probe how you wired one.

on this pageshow

explore

questions

4

In PyRIT, what does a prompt target object do, and why does one run usually need more than one of them?

level: juniorimportance: must knowfreq 72%

answer

  1. target = the only object that knows the endpoint
  2. send prompt, return response, write to memory
  3. three roles: system under test, attacker, scorer
  4. one turn can be several billed calls
  5. shared target = attacker inherits refusals

basics

~20 s

A PyRIT prompt target is the adapter that knows one endpoint: it sends a prompt there and returns the response. A run usually needs several of them, because the system under test, the attacker model that drafts prompts, and any model that scores replies are three different endpoints.

solid answer

~50 s

The prompt target is the only object in a PyRIT run that knows a concrete endpoint - its address, credentials, request shape and how to pull text out of the reply. Everything else (the attack strategy, the converters, the scorers) talks to it through the same send-and-return contract, which is why the same kind of object can play three roles in one run: the system under test, the adversarial model that writes the next attack turn, and the model a scorer asks for a verdict. Keeping those as separate target objects matters, because they normally differ in credentials, cost and safety configuration. You typically point the attacker role at a deployment permissive enough to draft adversarial text, while the system under test is the guarded production endpoint. Reuse one object for both and the attacker inherits the guarded endpoint's refusals: the run stalls, and the log reads like a robust target rather than a broken harness.

go deeper

for a junior

Says the target is the thing that sends prompts to a model or endpoint and brings back the answer, and that the attacker model and the model under test are configured separately.

for a middle

Explains the uniform send-and-return contract, names the three roles a target can play in one run, and notes the per-turn call multiplication and the credential/quota separation.

for a senior

Diagnoses a zero-hit run from stored conversations, separates quotas per role, and treats an inherited refusal from a shared attacker deployment as a harness defect rather than a result.

for a principal

Sets the standard for how runs are wired across the team - which deployments are approved for the attacker and judge roles, whose budget each burns, and what a report may claim when the roles shared infrastructure.

**What the object actually is.** PyRIT (the Python Risk Identification Toolkit) splits a red-team run into a few collaborating objects: an *attack strategy* that decides what to send and when to stop, *prompt converters* that rewrite a prompt before it leaves, *scorers* that judge each reply, a *memory* store that records everything, and *prompt targets*. The prompt target is the only one of those that knows the outside world exists. Its contract is deliberately narrow: accept a prompt request, deliver it to one concrete endpoint, return the response, and write both pieces to memory under a conversation id. Everything endpoint-specific lives inside the target and nowhere else - base URL, deployment or model name, the API key or bearer token, the request body shape, which field of the reply holds the assistant's text, and the retry and rate-limit behaviour. Concrete targets differ only in those details. `OpenAIChatTarget` speaks a chat-completions API. `HTTPTarget` takes a raw HTTP request template with a placeholder where the prompt is substituted, plus a callback that parses the reply body. Others cover non-chat surfaces - image generation, speech. Because they all satisfy the same send-and-return contract, a strategy written against "a target" runs unchanged against any of them: repointing a whole suite from a hosted deployment to a locally served open-weights model is a constructor change, not a rewrite. That is the payoff of keeping the boundary narrow. **Why one run needs several of them.** A red-team run is not one endpoint, it is a small fleet of them, and each one is a target object: | role | what it is | typical wiring | |---|---|---| | system under test | the thing you are measuring | the target passed to the attack strategy | | adversarial model | drafts the next attack turn in an automated multi-turn strategy | a second, separate target, usually a permissive deployment | | scoring model | a model-backed scorer asking "did this reply do the thing?" | a third target held by the scorer | | converter backend | any converter that rewrites a prompt with a model (translate, restyle, vary) | a fourth target held by that converter | These are normally different deployments with different credentials, different content filtering, and different quotas. The framework does not object if you pass the same object to all of them; nothing in the type system says you must not. **What it costs.** Do the arithmetic before you launch. One turn of an automated multi-turn strategy with a model-backed scorer is at minimum three billed calls: the attacker drafts, the system under test answers, the judge scores. Add one LLM-backed converter link and it is four. So a ten-turn strategy over twenty objectives is not 200 calls, it is 600 to 800, spread across three separate quotas - and the two calls you are *not* measuring are often the larger share of the bill, because attacker and judge prompts carry long instructions. Wall-clock follows the tightest quota, not the average one. Engineer time goes on wiring credentials for three deployments and on getting approval for a permissive attacker deployment, which in most organisations is the slow part. **Where the number misleads.** The failure that produces a *believable wrong answer* is role sharing. Point the attacker role at the same guarded production deployment as the system under test and the attacker inherits its refusals: it declines to draft the next adversarial turn, so what actually gets sent is an apology or an empty string, the target answers harmlessly, the scorer records a miss, and the run reports "0 of 200 objectives achieved." That reads as a strong defence and is in fact a broken harness. The same shape appears with a shared judge: a judge that shares the target's blind spots agrees with it, and a single outage across the shared endpoint scores as perfect safety. Shared *quota* is the subtler cousin. Scorer and converter traffic throttle the attack, retries stretch the run, and any latency or throughput figure you report then describes your own rate limit rather than the target's resilience. Note also what a target does *not* do: it returns text. It has no opinion on whether that text is a refusal, an error string in a 200 body, or a violation. Every safety judgement happens downstream of it. **What I check before believing a run.** Which target object fills which role, and that each holds its own client and credentials. A direct smoke test of the attacker target - one benign prompt, one that should be refused - reading the stored replies rather than trusting the constructor. Separate rate limits per role, with backoff configured so throttling degrades the run instead of recording failed attempts. A hand-read sample of stored conversations before I look at any aggregate: real adversarial prompts and real refusals mean a robust endpoint, while the attacker declining or empty prompt bodies mean a harness defect. And which memory store is in use, because an in-memory store leaves nothing to re-read or re-score once the process exits.

  • If a PyRIT run reports zero successful attempts, how do you tell a genuinely robust endpoint from a misconfigured attacker target?
    Read the stored conversations. A robust endpoint shows real adversarial prompts and refusals from the system under test; a broken harness shows the attacker itself declining, empty prompts, or transport errors on every turn.
  • Why does a model-backed converter also need a target?
    Because rewriting a prompt with a model is itself an endpoint call. That converter has to be pointed at a deployment, and each link costs a call per prompt on top of the attack and scoring traffic.

Wiring a run is like hiring three contractors who each bill by the hour: one writes the letters, one receives them, one reads the replies - and you only ever look at the invoice from the one you were testing. If you hire the same cautious person for all three jobs, the letters never get written and the file comes back empty.

saying these in an interview costs you the question

  • Believing one target object serves the whole run and the framework figures out the rest
  • Treating a run with zero hits as evidence of a safe endpoint without reading the stored conversations
  • Not knowing that a model-backed scorer or converter also spends endpoint calls
  • Pointing the attacker role at the same guarded production deployment as the system under test and being surprised by refusals

context

open as a page

In PyRIT, how does earlier conversation history reach the endpoint on turn five of a multi-turn run, and what changes when the prompt target wraps a service that keeps its own server-side session?

level: middleimportance: must knowfreq 58%

basics

~20 s

For a chat-completion style endpoint the target replays the stored turns: it reads the conversation out of memory and sends the whole message list each call. If the service keeps its own session and accepts only the newest message, the history lives on the server, so the tool's copy and the real context can drift apart.

open as a page

Your only access to a support assistant is an authenticated HTTP API with a bespoke JSON body, a rotating session token and a streamed reply. How do you get it under PyRIT as a prompt target, and what does that wrapper silently bound about the run?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Use PyRIT's generic HTTP target: give it the request template with a placeholder where the prompt goes, plus logic that pulls the assistant text out of the reply. You then own auth refresh, retries and rate limiting yourself. Anything the API never returns - system prompt, tool calls, retrieved context - stays untested.

open as a page

For a PyRIT engagement you can point the prompt target at the raw model endpoint behind a product or at the product's own chat API. How do you decide which surface to wrap, and how does that choice change what your report may claim?

level: principalimportance: should knowfreq 34%

basics

~20 s

Decide from the question being asked. The raw endpoint measures the model's own behaviour with no product defences; the product API measures what a user can actually reach, defences included. Each report claims only its own surface. Where budget allows, wrap both and read the gap between them as the value of the defences.

open as a page