For a PyRIT engagement you can point the prompt target at the raw model endpoint behind a product or at the product's own chat API. How do you decide which surface to wrap, and how does that choice change what your report may claim?
answer
- target choice defines the experiment
- raw endpoint = model behaviour, no defences
- product API = real exposure, drifts, no attribution
- gap between the two = value of the defence layer
- report names surface, date, roles, untested channels
basics
~20 sDecide from the question being asked. The raw endpoint measures the model's own behaviour with no product defences; the product API measures what a user can actually reach, defences included. Each report claims only its own surface. Where budget allows, wrap both and read the gap between them as the value of the defences.
solid answer
~50 sThe target choice is the experiment's definition, so pick it from the question. Wrapping the raw model endpoint isolates model behaviour: no system prompt, no input or output filtering, no retrieval. It is reproducible and cheap to rerun after a model swap, and says nothing about user risk. Wrapping the product API measures the surface a real user has, defences and all - the number a stakeholder wants - but it moves under you when the product ships a prompt change, and a failed attempt cannot be attributed to the model, the guard or the system prompt. Running both is the informative option: the difference between the runs estimates what the product layer buys. Whatever you choose, the report names the surface, the date and the target's configuration, because a success rate without them is comparable to nothing, including your own next run.
go deeper
Recognises that testing the model directly and testing the product are different things and that results should say which was done.
Explains that the product surface adds a system prompt, filtering and retrieval, so a lower success rate there does not describe the model.
Runs both surfaces where budget allows, attributes the gap carefully, and records surface, history mode and role endpoints with the number.
Sets the engagement policy: which surface answers which question, what a report may claim, whose quota is spent, whether production may be touched, and how a drifting product surface is retested over time.
**The target choice is the experiment.** The prompt target is the boundary of the instrument. Everything inside it is under test; everything outside it is an assumption. So the decision is not a wiring preference, it is the definition of what the engagement measures - and it should be made from the question the engagement was commissioned to answer. Between the product's chat API and the raw model endpoint behind it sits a stack that is invisible from above: a system prompt, an input classifier, an output moderation pass, retrieval over some corpus, a tool layer, per-user rate limits, session management and response templating. Wrapping the raw endpoint puts all of that outside the instrument. Wrapping the product API puts all of it inside, undifferentiated. *Model-level question* - "is this base model acceptable to build on?" Wrap the raw endpoint. Attribution is clean, because there is nothing between the prompt and the model. The suite is stable across product releases and cheap to rerun when the model version changes. It systematically overstates user-facing risk, because none of the defences a real user meets are in the path. *Product-level question* - "what can a customer actually get out of this assistant today?" Wrap the product API. This is the number that maps to real exposure and the one a stakeholder wants. It is also the fragile one: it moves whenever the product ships a prompt or guard change, and a failed attempt has several possible causes - the model, the system prompt, the input filter, the output filter - that this instrument cannot separate. **What each costs.** The raw-endpoint suite runs on your own quota, prices per token, parallelises to whatever concurrency the provider allows, and finishes a few-hundred-objective sweep in tens of minutes. The product-surface run is the expensive one: it spends the product's quota at product rate limits (often single-digit requests per minute per account), while the attacker and scorer calls spend yours in parallel; it usually needs a test tenant, seeded data, and written sign-off before anything touches production. Those approvals are calendar time, not compute, and in practice they - not methodology - are what decides the surface on most engagements. Running both roughly doubles attacker and scorer spend and adds the wiring cost of a second target. **Reading the pair, and where the number misleads.** Two runs with the same strategies over both surfaces is the only defensible basis for a claim about the defence layer, and even then the claim is narrow. Three specific misreadings to refuse: - **The product rate read as a model statement.** A low product-surface rate says the doorman turned people away; it is not evidence about the model, and the raw-endpoint run is the evidence that contradicts it. - **The gap read as durable coverage.** If the guard keys on surface forms - phrasings, keywords, an encoding - that your particular attacker model happens to produce, changing the converter chain or the phrasing family can collapse the gap. The honest statement is scoped: this generator, these strategies, this date. - **The two runs treated as paired trials.** An adaptive multi-turn strategy generates its next prompt from the previous *response*, so identical configuration does not mean identical prompts across the two surfaces. The runs are comparable in aggregate, not attempt by attempt. Drift is the fourth trap: the product surface six weeks later is a different system, so a delta between two product runs is not a regression signal unless the versions on both sides were pinned and recorded. **What the report must carry.** The surface wrapped and why; whether the target replayed history or referenced a server-side session; which deployments filled the attacker and scoring roles; the converter chain and seed dataset version; the model or product version observed and the date; and an explicit list of channels the run did not exercise - other entry points, tool paths, retrieval sources. A success rate stripped of those is comparable to nothing, including your own next run, and stakeholders will compare it anyway. **Organisationally**, I would rather own a small always-runnable raw-endpoint suite plus a scheduled product-surface run than one large one-off. The first catches model regressions cheaply and continuously; the second is the one leadership reads. Decide up front whose budget each burns and whether production may be touched at all - that constraint usually arrives late and rewrites the plan.
- A product-surface run shows a much lower success rate than a raw-endpoint run. When is that gap misleading?When the defence keys on the surface form of the prompts your generator produces. Change the phrasing family or a converter and the gap can collapse, so the number reflects your generator, not the guard's real coverage.
- Why record the target's configuration alongside the success rate?Because the rate is only meaningful relative to the surface, the history mode and the endpoints filling the attacker and scorer roles. Without them the next run is not comparable, even against itself.
saying these in an interview costs you the question
- Reporting one success rate without naming the surface the target wrapped
- Reading a low product-surface rate as evidence that the model is safe
- Comparing this engagement's number to a previous one across different surfaces or history modes
- Treating a defence-layer gap as durable without varying the phrasings the attacker produces
- Ignoring which quota a run spends and whether production may be exercised at all