skip to content

Custom Targets

When nothing ships for your endpoint you write the adapter, and every scorer and converter downstream inherits its bugs. Interviewers ask what your adapter got wrong and how you found out.

on this pageshow

explore

questions

5

You need to run PyRIT against an internal HTTP chat endpoint that PyRIT ships no prompt target for. What does the custom prompt target you write own, and what does the rest of the PyRIT run still do for you?

level: juniorimportance: must knowfreq 55%

answer

  1. adapter = only piece that knows your endpoint
  2. translate, do not editorialise
  3. auth, schema, timeouts, retries
  4. no attack logic in the target
  5. swallowed error looks like a refusal

basics

~20 s

Your target owns everything endpoint-specific: authentication, request and response shape, timeouts and retries, and turning the reply into one response PyRIT can store. Everything else is unchanged. The attack strategy still chooses prompts, converters still transform them, scorers still judge them, and the memory store still records each exchange.

solid answer

~50 s

PyRIT's prompt target is a deliberately narrow adapter: it receives a prompt request from the run and returns the model's response. Chat-shaped targets additionally accept the prior turns of the conversation so a multi-turn attack works at all. Inside that boundary you own auth and token refresh, mapping PyRIT's request into your service's JSON, mapping the reply back, timeouts, retry policy, and any content types beyond text. What you must **not** absorb: prompt selection, prompt transformation (converters), the decision about whether an attempt counted (scorers), or the turn budget (the attack strategy). Pushing any of those into the adapter makes the run unauditable, because the memory store no longer reflects what was actually sent. The single most common beginner bug is swallowing failures — catching an exception and returning an empty string. Downstream, that is indistinguishable from a target that answered with nothing, so a transport outage quietly becomes a column of clean negatives.

go deeper

for a junior

Should say the adapter handles auth, request/response mapping and errors, and that prompts, converters and scoring stay in the framework.

for a middle

Adds the chat-versus-single-prompt distinction, the memory store as the record of truth, and why swallowing errors corrupts results.

for a senior

Frames the adapter as the run's only untrusted translation layer and describes the byte-for-byte verification they do before believing any number from it.

for a principal

Treats adapter conformance as a shared engineering asset — one reviewed template plus a conformance check, so numbers from different services stay comparable.

### What a PyRIT prompt target actually is A *prompt target* is PyRIT's adapter for exactly one endpoint. You subclass `PromptTarget`, or `PromptChatTarget` when the endpoint accepts a conversation rather than a lone string, and you implement one method: `send_prompt_async(prompt_request=...)`. The framework hands that method a `PromptRequestResponse` holding one or more `PromptRequestPiece` objects, and each piece carries the fields the rest of the run is built on — `original_value` (what the attack strategy produced), `converted_value` (what the converters turned it into, and therefore the string that must actually go on the wire), `conversation_id`, `role`, `original_value_data_type`, and `response_error`. You return a `PromptRequestResponse` of your own, normally built with the helper `construct_response_from_request(request=..., response_text_pieces=[...], response_type="text", error="none")`. `PromptChatTarget` adds `set_system_prompt(...)`, which is the only route by which a strategy that carries a system prompt reaches your service at all. Both sides of that exchange are written into PyRIT's memory store — `CentralMemory`, backed by DuckDB on a laptop or a shared SQL database in a team setup. Every artefact produced afterwards is computed from those rows: the transcripts a triager reads, the attack-success rate in the report, a re-score with a different scorer next month. Your adapter is the only component in the run that can make that record disagree with what really happened on the wire. ### The three jobs the adapter owns **Faithful translation.** Map `converted_value` into your service's request schema, and map the reply back, without editorialising. No stripping of system banners, no trimming, no friendly prefix, no lower-casing. A scorer that matches on refusal language reads whatever you fabricate as if the model said it. **Session and conversation identity.** A stateless endpoint needs the prior turns passed on every call; a stateful one needs its session bound to PyRIT's `conversation_id` and torn down afterwards. Getting this wrong silently mixes attempts together. **An honest error contract.** Authentication failure, HTTP 5xx, quota rejection and client timeout are transport events, not model behaviour. They belong in `response_error` (or as a raised exception the framework records as a failed attempt), never as an empty string with `error="none"`. ### The prohibition No attack logic inside the target. Not a hardcoded system prompt, not a filter on what you are willing to send, not a cleanup pass on the answer. Anything the adapter does to the prompt is invisible in `converted_value`, so the memory store stops describing what was tried. ### What the rest of the run still does The attack strategy still selects and escalates prompts and owns the turn budget. Converters still transform `original_value` into `converted_value`. Scorers still decide whether an attempt counted, and can be re-run over stored transcripts. Memory still records every exchange keyed by `conversation_id`. ### What it costs Writing and reviewing a target for a plain HTTP chat endpoint is typically half a day to two days of engineer time; most of that is auth, retries and the error contract, not the happy path. The adapter then sits inside the run's metered cost: a multi-turn red-team attempt spends three separate model calls per turn — the attacker model that composes the next prompt, your target, and the scorer that judges the reply. Sixty seed prompts at a five-turn budget is on the order of 900 calls before any retry, and roughly a third of those go through your adapter. A well-meaning "retry three times on any failure" inside it can triple that third, turning a rate-limited afternoon into an unbudgeted bill and a run whose wall-clock is dominated by backoff. ### Where the number misleads The classic bug is `except Exception: return ""`. The empty string is stored with `error="none"`, so nothing downstream can tell it from a model that answered with nothing. Those attempts stay in the attack-success-rate denominator and count as clean negatives, so the rate falls in proportion to how flaky the transport was that day. Nobody reads a low attack-success rate as "my adapter is broken" — they read it as "the service held", which is precisely the wrong conclusion drawn from the right-looking number. ### What I check before believing anything Send one benign prompt, then pull the stored pieces for that `conversation_id` out of memory and diff them against the same call made by hand with `curl`: `converted_value` must equal the request body, and the stored response text must match the reply byte-for-byte. Then force the unhappy paths — a 500, a 401 after token expiry, a hang past the client timeout — and confirm each lands with a non-`none` `response_error` rather than as a tidy empty answer.

  • Why is a chat-shaped target different from one that takes a single prompt?
    A multi-turn attack needs the prior turns to reach the endpoint. A target that accepts only a lone prompt cannot carry history, so any multi-turn strategy driven through it silently degrades into repeated single-shot attempts.
  • Your endpoint needs a bearer token that expires every fifteen minutes. Where does refresh belong?
    Inside the adapter, transparent to the run. A long engagement outlives the token; if refresh is manual, half the run's failures are auth errors that the report will read as target behaviour.
  • The endpoint returns text plus a moderation verdict field. Where should that verdict go?
    Store it, but do not let the adapter act on it. It is evidence for a scorer or for triage, not a decision the adapter makes about whether the attempt counted.

saying these in an interview costs you the question

  • Catching every exception and returning an empty string or a placeholder as the model's answer.
  • Putting a system prompt, prompt filtering or output cleanup inside the target.
  • Assuming the scorer will notice that a response was an error body rather than a reply.
  • Never checking what the adapter actually wrote into PyRIT's memory store.
  • Reimplementing turn-taking inside the adapter instead of letting the attack strategy drive it.

context

open as a page

You pointed a PyRIT run at an internal chat service through a prompt target you wrote yourself, and the attack-success rate came back near zero. How do you tell a genuinely hardened service from a broken adapter?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Do not trust the summary number yet. Open the responses PyRIT stored: a hardened service returns real refusal text, a broken adapter stores empty strings, error bodies, truncated replies or the same string every time. Then replay two or three of those exact prompts by hand and compare against what the adapter recorded.

open as a page

A custom PyRIT prompt target wraps an endpoint that streams its reply back in chunks. What must the adapter do before that reply reaches PyRIT's scorers, and what goes wrong if it returns early?

level: middleimportance: should knowfreq 38%

basics

~20 s

Consume the whole stream and assemble one complete response before returning, because scorers see a single stored response, not chunks. If you return on the first chunk or bail at a client timeout, a truncated answer is stored as the target's real reply, and a response that was about to comply can score as a refusal.

open as a page

Your custom PyRIT prompt target calls a service that keeps conversation state server-side behind a session id it issues. What breaks within a multi-turn run and across repeated runs, and how do you make each attempt independent?

level: seniorimportance: should knowfreq 34%

basics

~20 s

PyRIT already tracks conversation identity, so reusing one server session makes history arrive twice or bleed between attempts. Mint a fresh session per PyRIT conversation, bind it to that conversation, and tear it down at the end. Otherwise a result depends on which attempts ran before it and stops being reproducible.

open as a page

Four teams each wrote their own PyRIT prompt target for their own service, and each now reports an attack-success rate to you. What do you require before you are willing to compare those numbers across teams?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Require every adapter to pass one shared conformance check: a known refusal, a known compliance, a forced error and a forced timeout, each landing in PyRIT's memory as a distinct expected outcome. Also require the same prompt set, scorer and turn budget. Otherwise the differences measure adapters, not services.

open as a page