A PyRIT run against your deployed chat assistant finishes with no successful attacks recorded. What has that run actually tested, and what has it not?
answer
- denominator = wired targets
- silence is not hardening
- one target object per surface
- attempts per target
- positive control
basics
~20 sOnly the targets you configured. PyRIT sends prompts to the prompt targets you wired and scores the replies; anything it was never pointed at, such as another endpoint, an upload path or an internal service, is untested rather than proven safe. A clean result describes your target list, not the application.
solid answer
~50 sPyRIT's unit of execution is a **prompt target**: an object that knows how to send a request to one place and hand back the response. The run's denominator is the set of targets you constructed, so *zero successful attacks* means zero on **those** targets, with **those** seed prompts, judged by the scorer you attached. A deployed assistant is normally more than one such place: the public chat endpoint, an ingestion path that pulls documents or tickets into context, a file upload, an internal or admin endpoint, an older API version still routed. Each needs its own wired target or it contributes nothing. The trap is that an untargeted surface and a genuinely hardened one produce the identical artefact in the run: silence. So report the clean run as a **coverage statement** — which targets ran, how many attempts each got — not as an assurance about the product.
go deeper
Should say the run only covers what was pointed at, and that a clean result is not proof of safety.
Should name the target object as the unit of coverage and list concrete intake paths a chat-endpoint-only run misses.
Should distinguish never-wired, wrong-layer and inert targets, and know to verify with per-target attempt counts and a positive control.
Should frame the clean run as a coverage claim owed to a decision maker, and set the expectation that the wired set is inventoried rather than remembered.
### What the framework actually does PyRIT is an orchestration library, not a scanner. You write a short script that constructs three kinds of object and hands them to an attack orchestrator: a **prompt target** — an object that wraps exactly one send path, such as a hosted chat-completion endpoint, an HTTP endpoint of your own application, or a local model; a set of **seed prompts**, the attack content, usually loaded from a dataset file; and one or more **scorers**, the objects that read a response and decide whether it counts as a hit. The orchestrator then loops: take a prompt, optionally push it through converters, send it through the target, write the request and the response into PyRIT's memory store (a local DuckDB file by default), score the response, move on. Nothing in that loop performs discovery. There is no crawl, no route enumeration, no step that goes looking for the other places your model can be reached. The set of surfaces probed is precisely the set of prompt-target objects your script constructed — no more, and never implicitly more. ### The sentence a clean run actually supports "Zero successful attacks" unpacks to: *of N attempts, drawn from prompt set P, delivered through target set T, none was labelled a hit by scorer S.* Four qualifiers, and a report typically preserves only the first. Candidates habitually collapse the target bound into the prompt bound — they will discuss whether the prompt set was strong enough and never ask how many places it was sent to. ### Three ways a surface stays out of the denominator 1. **Never wired.** No target object exists for it: a nightly batch summariser, a webhook or mail consumer, a second tenant, a partner integration, an internal console that talks to the same deployment, an older API version still routed for legacy clients. It cannot appear in any result, and its absence produces no artefact at all. 2. **Wired at the wrong layer.** The target points at the raw provider API while real users hit an application endpoint that adds a system prompt, retrieval and a guard — or the reverse. The run characterises a system that is not the deployed one. 3. **Wired but inert.** Every attempt failed on authentication, TLS, a rate limit or a timeout. If those failures are not separated from genuine refusals, the target appears in the run, occupies rows in memory, and tested nothing. ### What it costs, and why teams under-wire Cost per attempt is metered calls, not requests: the target call, the scorer call when the scorer is itself an LLM judge, plus an adversarial model call on every turn of a multi-turn strategy. A single-turn sweep of a couple of hundred seed prompts against one chat endpoint is therefore several hundred metered calls — minutes to an hour of wall clock and small money. Engineer time is the dominant cost and it is wildly uneven: a second chat-shaped target is an hour of work, while a target for an asynchronous ingestion path that must deposit content and later read back the effect is days. That asymmetry, not risk, is what usually decides which surfaces get wired — the wired set skews toward what is cheap to wire. ### Where the number misleads Attack-success rate is hits over attempts, and attempts are counted only over wired targets. Adding a genuinely hardened target lowers the rate. Forgetting a vulnerable surface also lowers the rate. The two are arithmetically indistinguishable in the aggregate, so an improving trend line can mean better defences or a shrinking denominator. Errored attempts make it worse: an empty body carries no harmful content, so a scorer labels it a non-hit, and absent data sits in the denominator behaving like a survived attack. Two runs are comparable only if the target set and the prompt set were identical; nothing in the artefact tells you whether they were. ### What you check before believing it Pull per-target attempt counts out of PyRIT's memory rather than reading the summary. Classify each attempt into scored response, refusal, transport error, auth rejection, truncation, timeout — a target dominated by anything but the first two contributed no coverage. Confirm the target's base URL, credentials class and deployment identifier match what production serves. Then fire a **positive control**: a benign probe whose production response you can recognise, sent through the same target and the same scorer, proving end to end that requests landed and results came back. Finally, publish the target list and its attempt counts beside the hit count. A clean run is a coverage statement about a named set of surfaces; presented without that set, it is not evidence about the application at all.
- Someone says the assistant is safe because PyRIT found nothing. What do you ask for first?The list of targets the run executed against with attempt counts, and how many attempts errored. Without that denominator the zero is uninterpretable.
- How can a target be wired and still contribute no coverage?If every attempt failed on auth, rate limiting or timeout and those failures were not distinguished from refusals, the target appears in the run while testing nothing.
- Why does adding surfaces cost more than it looks?Each attempt meters an attacker model, the target and the scorer, so the price scales with turns across every wired target, not with the target count alone.
A clean run is a searchlight report, not a darkness report. It tells you nothing moved inside the beam, and says nothing whatsoever about the part of the yard the beam never swept.
saying these in an interview costs you the question
- Reading zero successful attacks as evidence the application is safe
- Quoting a success rate without saying which targets were in the denominator
- Assuming one chat endpoint covers every way text reaches the model
- Counting errored or auth-rejected attempts as failed attacks
- Believing the framework enumerates surfaces on its own