When you configure a PyRIT prompt target, what is the difference in what a clean run proves if you wire it at the model provider's raw API versus at your application's own endpoint?
answer
- attribution vs realism
- wrapper supplies its own system prompt
- staging guard stubbed
- four-cell read of the pair
- two targets double the bill
basics
~20 sThe raw API tests the model alone, with no production system prompt, retrieval, tools or guard. The application endpoint tests the whole deployed stack. A clean run at one says nothing about the other: a guard can hide model weakness, and a bare-model result ignores every layer users actually pass through.
solid answer
~60 sA target is a send path plus its configuration: URL, credentials, deployment identifier, and whatever the wrapper adds around your prompt. Which layer you attach it to decides what the number means. Wired at the raw provider API you get **attribution**: a hit is a property of the model with your test-time framing, uncontaminated by anything the app does. Wired at the application endpoint you get **realism**: a hit is something a user could actually produce, with the production system prompt, retrieval context, tool wiring and guard in the loop. The mistake is treating either as a substitute. A clean bare-model run tells you nothing about the app's prompt assembly or its retrieval intake; a clean app run tells you nothing about what the model would do if the guard were misconfigured tomorrow. In practice you wire both as separate targets and read the pair: hit at the model, clean at the app means a control is doing work — and you should know which one, because it can be turned off.
go deeper
Should recognise that pointing at the model API and pointing at the app are not the same test.
Should explain what each layer adds — system prompt, retrieval, tools, guard — and why a clean run at one does not transfer to the other.
Should insist on asserting the target's configuration before a long run and read the model/app pair as an attribution matrix.
Should decide how much of the budget buys attribution versus realism, and require that which layer was wired is recorded with every result.
### What a prompt target actually is A PyRIT prompt target is a send path plus its configuration: a base URL or deployment name, credentials, headers, and whatever the wrapper adds around the text you supply. Two targets can send the identical seed prompt and be talking to two different systems, because the layers sitting between the wrapper and the model weights are part of what is under test. The choice of layer is therefore not a plumbing detail — it decides what the resulting number is a statement about. **Wired at the model provider's raw API**, the target reaches the weights with only whatever framing your script supplies. A hit here is attributable to the model: it is a property of the model plus your test-time framing, uncontaminated by the application. That is genuinely valuable, and it is the only place you can get it. **Wired at your application's own endpoint**, the request passes through everything production puts in the path: the assembled system prompt, retrieved context, tool declarations, output filters, an input or output guard, rate limits, tenant policy. A hit here is something a real user could actually produce. That is realism, and it is the only place you can get *that*. ### The failure mode that eats most of these runs The dangerous case is not choosing the wrong URL — it is a target that looks like production and is not. Concretely: - The target wrapper lets you supply your own system prompt or conversation prefix. If you do, and production assembles a different one, the run has characterised a configuration that exists nowhere. - The target points at staging, where the guard is stubbed for cost, the retrieval index is sanitised, tools are disabled, or the test tenant carries a permissive policy. - The deployment behind the endpoint is pinned to a different model version than the one serving users. Any single one of those turns a production claim into a claim about a different system, and none of them is visible in the run's output. The discipline is to **assert the configuration rather than assume it**: before a long run, send a handful of shape-checking probes and confirm the guard rejects something you know it rejects, that a retrieval-dependent question actually retrieves, and that the deployment identifier echoed in headers or logs is the one production serves. ### Reading the pair Wiring both layers gives a four-cell table, and the table is the reason to spend the extra budget: | model layer | app layer | reading | |---|---|---| | hit | hit | no effective control between them — this is the finding | | hit | clean | a control is carrying the result; name it, because it is now a single point of failure | | clean | clean | weakest evidence; could be either layer, or an under-powered prompt set | | clean | hit | the application added the exposure: prompt assembly, retrieved content, tool output | The second row is the one interviewers push on. A clean application run built on a hit at the model does not mean the weakness is gone; it means one component is standing between the weakness and a user, and the moment that component is bypassed, disabled for latency, misconfigured, or simply not present on a second endpoint, the exposure returns. Recording "clean" without recording which control produced it throws away the actionable half of the result. ### What it costs Each attempt bills the target and, when the scorer is an LLM judge, the scorer too; multi-turn strategies add an adversarial model call per turn. Wiring two layers therefore roughly doubles the metered spend for the same prompt set, plus the engineering time to build and credential a second target. That is the honest reason teams wire one layer and not the other. The proportionate answer is to sample rather than skip: run the full prompt set at the application endpoint for realism, and a reduced, high-signal subset at the raw model for attribution. A partial attribution signal is worth far more than none. ### Where the number misleads The single most common misreport is a rate produced at one layer and described in the language of the other — "the assistant refused 98% of attacks" from a run that never went through the assistant. It is misleading in both directions: a bare-model result understates a deployment whose retrieval intake adds exposure the model alone never sees, and overstates a deployment where a guard is doing all the work. So carry the layer with the number. And note the boundary that survives even a careful application-layer run: it is a claim about *that endpoint*. Every other path into the same model — ingestion, upload, batch, admin — is a separate target and a separate claim.
- You get a hit at the raw model and a clean result at the application endpoint. What do you write down?That a control between them is carrying the result, which control it is, and that the model-level weakness remains latent if the control is ever bypassed, disabled or misconfigured.
- How do you check the target is really talking to the production stack?Send small shape-checking probes first: confirm the guard rejects a known-rejected case, that retrieval-dependent answers actually retrieve, and that the deployment identifier matches the one serving users.
saying these in an interview costs you the question
- Reporting an application-level result while the target was actually the raw model endpoint
- Supplying a test system prompt and presenting the result as production behaviour
- Running against staging with the guard stubbed and calling it a production result
- Assuming a clean app run implies the model is safe