You pointed a PyRIT run at an internal chat service through a prompt target you wrote yourself, and the attack-success rate came back near zero. How do you tell a genuinely hardened service from a broken adapter?
answer
- positive control through the same adapter
- read transcripts, not the summary
- empty / constant / error-body responses
- response times below model latency
- third bucket: did not execute
basics
~20 sDo not trust the summary number yet. Open the responses PyRIT stored: a hardened service returns real refusal text, a broken adapter stores empty strings, error bodies, truncated replies or the same string every time. Then replay two or three of those exact prompts by hand and compare against what the adapter recorded.
solid answer
~50 sA near-zero result is ambiguous by construction, because a broken adapter and a hardened target produce the same summary. The only way through is to leave the aggregate and read the stored transcripts. 1. **Response text.** Real refusals vary in wording. Identical strings, empty strings, JSON error bodies or HTML error pages stored as replies all point at the adapter. 2. **Response lengths.** A flat ceiling means truncation; a flat zero means swallowed errors. 3. **Timing.** Responses returning in milliseconds never reached a model. 4. **The other side.** Endpoint logs: did the requests arrive at all, with the prompt intact after converters? 5. **A hand replay.** Send a stored request verbatim and diff the reply against what was recorded. Also run a positive control — a prompt the service definitely answers. If that comes back empty too, the adapter is broken and no security conclusion is available at all.
go deeper
Should at least say to look at the actual responses rather than trusting the summary.
Adds a positive control and names concrete adapter signatures — empty strings, error bodies, suspiciously fast responses.
Runs the diagnosis systematically across both sides, checks the sent request as well as the response, and fixes the error contract so failures are their own bucket.
Makes canaries and a three-bucket outcome model a standing requirement, so no report on any service can present unexecuted attempts as negatives.
### Why the number is ambiguous by construction A near-zero attack-success rate has two explanations that are observationally identical in the report: the service refused everything, or your adapter never delivered anything. PyRIT computes the rate from what its memory store holds; if the adapter wrote empty strings, error bodies or truncated text with `response_error` set to `"none"`, those rows look exactly like a target that declined. Everyone wants the first explanation to be true, which is why this needs a procedure rather than a judgement call. ### Start with a positive control The fastest discriminator is a prompt the service will obviously answer — a plain factual question — sent through the same adapter, in the same run, with the same auth and the same converters. If that also lands as an empty or non-usable response, the transport is broken and there is no security conclusion available at all: the number is meaningless rather than good. This costs one extra call per run and it is the single highest-value thing in this whole answer. ### Read the stored transcripts for adapter signatures Leave the aggregate and open the rows in memory for a handful of `conversation_id`s. Specific bugs leave specific fingerprints: | What you see stored | What it almost always means | |---|---| | The same string every attempt, or an empty string | An exception swallowed and a placeholder returned | | A JSON error envelope or an HTML error page as the reply | The adapter never checked the status code and stored the error body as model output | | Every response clipped near the same length | Streaming truncation or a client timeout | | Response times below plausible model latency | The request never left, or a cache, WAF or auth proxy answered instead | | Requests absent from the endpoint's own logs | Auth or routing failure inside the adapter | | Sensible replies that ignore earlier turns | Session or `conversation_id` binding is wrong, so turn two never saw turn one | Then check the *request*, not only the response. Compare what the adapter put on the wire against `converted_value` in memory. An adapter that JSON-escapes twice, drops the role field, or silently truncates a long prompt is testing a different string than the one the report claims was tried — and a double-escaped prompt reads to the model as gibberish, which produces exactly the confused non-answers that score as refusals. ### Replay by hand Take two or three stored requests and send them verbatim with `curl`, outside PyRIT, then diff the reply against what was recorded. This separates "the service refuses this prompt" from "the adapter mangled this prompt" in about ten minutes, and it is the evidence you show when someone asks why you do not trust the dashboard. ### What it costs, and what it saves The whole diagnosis is under an hour of engineer time: one canary call, ten minutes reading transcripts, three hand replays. Skipping it costs the far larger thing — a re-run of the entire engagement once the bug is found, at full metered cost (attacker model, target and scorer are three billed calls per turn), plus the credibility of every earlier report produced through the same adapter. ### Where the number misleads, and the fix Even after the adapter is correct, a two-bucket report — hit and non-hit — will keep hiding this. An attempt that failed to execute is not evidence about the model, and leaving it in the denominator drags the rate down in proportion to the day's flakiness. Make the adapter raise on non-success status, on an unterminated stream and on a timeout, so the framework records an execution error, then report three buckets: hit, refused, did-not-execute, with the third excluded from the denominator and its count stated. A run that cannot produce the third bucket cannot be believed in either direction. And a genuinely near-zero rate, honestly measured, still does not mean "the service is safe". It means this prompt set, judged by this scorer, at this turn budget, did not land. Report the coverage and what was not tried alongside it, or the number will be read as a clearance it never earned.
- What single check would you add to every run so this ambiguity never recurs?A canary attempt per run: one benign prompt the service reliably answers and, where safe, one known-weak prompt. If the canaries do not behave, the run is void before anyone reads the number.
- The adapter is fine and the refusals are real. What is still wrong with reporting near-zero as 'the service is safe'?It is one prompt set, one scorer and one turn budget. Near-zero says those attempts did not land; it does not say the surface is covered. Report what was tried and what was not.
- How do you record an attempt that failed to execute so it does not pollute the rate?Let the adapter raise so the framework marks the attempt as an error, then exclude errored attempts from the denominator and report their count separately.
A smoke detector that has never gone off is either protecting a safe building or has a dead battery, and from across the room the two look the same. The positive control is pressing the test button before you conclude anything about the building.
saying these in an interview costs you the question
- Reporting a low attack-success rate as evidence the service is hardened, without opening a single transcript.
- No positive control — never checking that a benign prompt gets a real answer through the adapter.
- Assuming the request left the machine because no exception was raised.
- Only two outcome buckets — hit and non-hit — with no way to record an attempt that failed to execute.
- Comparing against a run on a different adapter or a different prompt set and calling the difference an improvement.