You need to run PyRIT against an internal HTTP chat endpoint that PyRIT ships no prompt target for. What does the custom prompt target you write own, and what does the rest of the PyRIT run still do for you?
answer
- adapter = only piece that knows your endpoint
- translate, do not editorialise
- auth, schema, timeouts, retries
- no attack logic in the target
- swallowed error looks like a refusal
basics
~20 sYour target owns everything endpoint-specific: authentication, request and response shape, timeouts and retries, and turning the reply into one response PyRIT can store. Everything else is unchanged. The attack strategy still chooses prompts, converters still transform them, scorers still judge them, and the memory store still records each exchange.
solid answer
~50 sPyRIT's prompt target is a deliberately narrow adapter: it receives a prompt request from the run and returns the model's response. Chat-shaped targets additionally accept the prior turns of the conversation so a multi-turn attack works at all. Inside that boundary you own auth and token refresh, mapping PyRIT's request into your service's JSON, mapping the reply back, timeouts, retry policy, and any content types beyond text. What you must **not** absorb: prompt selection, prompt transformation (converters), the decision about whether an attempt counted (scorers), or the turn budget (the attack strategy). Pushing any of those into the adapter makes the run unauditable, because the memory store no longer reflects what was actually sent. The single most common beginner bug is swallowing failures — catching an exception and returning an empty string. Downstream, that is indistinguishable from a target that answered with nothing, so a transport outage quietly becomes a column of clean negatives.
go deeper
Should say the adapter handles auth, request/response mapping and errors, and that prompts, converters and scoring stay in the framework.
Adds the chat-versus-single-prompt distinction, the memory store as the record of truth, and why swallowing errors corrupts results.
Frames the adapter as the run's only untrusted translation layer and describes the byte-for-byte verification they do before believing any number from it.
Treats adapter conformance as a shared engineering asset — one reviewed template plus a conformance check, so numbers from different services stay comparable.
### What a PyRIT prompt target actually is A *prompt target* is PyRIT's adapter for exactly one endpoint. You subclass `PromptTarget`, or `PromptChatTarget` when the endpoint accepts a conversation rather than a lone string, and you implement one method: `send_prompt_async(prompt_request=...)`. The framework hands that method a `PromptRequestResponse` holding one or more `PromptRequestPiece` objects, and each piece carries the fields the rest of the run is built on — `original_value` (what the attack strategy produced), `converted_value` (what the converters turned it into, and therefore the string that must actually go on the wire), `conversation_id`, `role`, `original_value_data_type`, and `response_error`. You return a `PromptRequestResponse` of your own, normally built with the helper `construct_response_from_request(request=..., response_text_pieces=[...], response_type="text", error="none")`. `PromptChatTarget` adds `set_system_prompt(...)`, which is the only route by which a strategy that carries a system prompt reaches your service at all. Both sides of that exchange are written into PyRIT's memory store — `CentralMemory`, backed by DuckDB on a laptop or a shared SQL database in a team setup. Every artefact produced afterwards is computed from those rows: the transcripts a triager reads, the attack-success rate in the report, a re-score with a different scorer next month. Your adapter is the only component in the run that can make that record disagree with what really happened on the wire. ### The three jobs the adapter owns **Faithful translation.** Map `converted_value` into your service's request schema, and map the reply back, without editorialising. No stripping of system banners, no trimming, no friendly prefix, no lower-casing. A scorer that matches on refusal language reads whatever you fabricate as if the model said it. **Session and conversation identity.** A stateless endpoint needs the prior turns passed on every call; a stateful one needs its session bound to PyRIT's `conversation_id` and torn down afterwards. Getting this wrong silently mixes attempts together. **An honest error contract.** Authentication failure, HTTP 5xx, quota rejection and client timeout are transport events, not model behaviour. They belong in `response_error` (or as a raised exception the framework records as a failed attempt), never as an empty string with `error="none"`. ### The prohibition No attack logic inside the target. Not a hardcoded system prompt, not a filter on what you are willing to send, not a cleanup pass on the answer. Anything the adapter does to the prompt is invisible in `converted_value`, so the memory store stops describing what was tried. ### What the rest of the run still does The attack strategy still selects and escalates prompts and owns the turn budget. Converters still transform `original_value` into `converted_value`. Scorers still decide whether an attempt counted, and can be re-run over stored transcripts. Memory still records every exchange keyed by `conversation_id`. ### What it costs Writing and reviewing a target for a plain HTTP chat endpoint is typically half a day to two days of engineer time; most of that is auth, retries and the error contract, not the happy path. The adapter then sits inside the run's metered cost: a multi-turn red-team attempt spends three separate model calls per turn — the attacker model that composes the next prompt, your target, and the scorer that judges the reply. Sixty seed prompts at a five-turn budget is on the order of 900 calls before any retry, and roughly a third of those go through your adapter. A well-meaning "retry three times on any failure" inside it can triple that third, turning a rate-limited afternoon into an unbudgeted bill and a run whose wall-clock is dominated by backoff. ### Where the number misleads The classic bug is `except Exception: return ""`. The empty string is stored with `error="none"`, so nothing downstream can tell it from a model that answered with nothing. Those attempts stay in the attack-success-rate denominator and count as clean negatives, so the rate falls in proportion to how flaky the transport was that day. Nobody reads a low attack-success rate as "my adapter is broken" — they read it as "the service held", which is precisely the wrong conclusion drawn from the right-looking number. ### What I check before believing anything Send one benign prompt, then pull the stored pieces for that `conversation_id` out of memory and diff them against the same call made by hand with `curl`: `converted_value` must equal the request body, and the stored response text must match the reply byte-for-byte. Then force the unhappy paths — a 500, a 401 after token expiry, a hang past the client timeout — and confirm each lands with a non-`none` `response_error` rather than as a tidy empty answer.
- Why is a chat-shaped target different from one that takes a single prompt?A multi-turn attack needs the prior turns to reach the endpoint. A target that accepts only a lone prompt cannot carry history, so any multi-turn strategy driven through it silently degrades into repeated single-shot attempts.
- Your endpoint needs a bearer token that expires every fifteen minutes. Where does refresh belong?Inside the adapter, transparent to the run. A long engagement outlives the token; if refresh is manual, half the run's failures are auth errors that the report will read as target behaviour.
- The endpoint returns text plus a moderation verdict field. Where should that verdict go?Store it, but do not let the adapter act on it. It is evidence for a scorer or for triage, not a decision the adapter makes about whether the attempt counted.
saying these in an interview costs you the question
- Catching every exception and returning an empty string or a placeholder as the model's answer.
- Putting a system prompt, prompt filtering or output cleanup inside the target.
- Assuming the scorer will notice that a response was an error body rather than a reply.
- Never checking what the adapter actually wrote into PyRIT's memory store.
- Reimplementing turn-taking inside the adapter instead of letting the attack strategy drive it.