A PyRIT run against a support assistant returned no hits, but an incident later shows the same behaviour was reachable in production. How do you work out whether the target wiring, rather than the attack content, produced the clean result?
answer
- reachability before content
- per-target attempt classification
- errors are absent data
- positive control end to end
- incident path becomes a target
basics
~20 sCheck reachability before blaming the prompts. Confirm the incident path was in the target list at all, replay a positive control through each target, read the stored attempts for empty bodies, errors, auth rejections and timeouts, and compare the target's endpoint, tenant and deployment identifier against what production served.
solid answer
~50 sWork outside-in, cheapest checks first. 1. **Was the path even wired?** Map the incident's entry point — endpoint, tenant, locale, API version, upload or ingestion route — onto the run's target list. Very often it simply is not there, and the investigation ends. 2. **Did attempts reach a live system?** Pull the run's stored conversations and count, per target, how many attempts produced a real response versus an error, an empty body, a rate-limit rejection or a timeout. Inert targets look identical to hardened ones in a summary. 3. **Was it the same system?** Compare base URL, credentials class, deployment identifier and any configuration the target wrapper supplied itself, such as a substituted system prompt, against production. 4. **Positive control.** Re-send a probe whose production response you can recognise, and confirm it comes back through the scorer. That single test separates *not reachable* from *not vulnerable*. Only after all four survive is the attack content — or the scoring of it — the leading hypothesis.
go deeper
Should think to check whether the failing path was covered by the run at all.
Should inspect stored attempts for errors and empty responses rather than assuming every attempt was a real test.
Should order the checks reachability-first, use a positive control, and compare target configuration against production before questioning the prompts.
Should treat the miss as a process defect, requiring the incident surface to become a permanent target and the run to publish per-target attempt classifications.
### Why the ordering is the whole answer A wiring failure and a working defence emit the same artefact: an attempt with no hit. Everything below is an attempt to break that tie using evidence the run already stored, before spending money and days on the expensive hypothesis that the attack content was too weak. Work outside-in, cheapest first, because each step can end the investigation. **1. Was the path even wired?** Map the incident's entry point — endpoint, tenant, locale, API version, upload or ingestion route, admin console — onto the run's target list. In practice this closes a large share of these investigations on the first check, at zero cost. PyRIT performs no discovery, so a surface with no prompt-target object never appeared in any denominator. **2. Did attempts reach a live system?** PyRIT writes every request and response into its memory store, so this is a query, not a re-run. Group attempts per target and classify them: scored response, refusal, transport error, auth rejection, truncation, timeout. A target whose column is dominated by anything other than the first two contributed no coverage. A target with an unexpectedly small attempt count lost work to retries, throttling or a cancelled run. If the harness caught exceptions and recorded a failed call as an ordinary attempt, this is the step where it becomes visible. **3. Was it the same system?** Compare the target's base URL, credentials class, deployment identifier, and anything the wrapper supplied itself against what production served on the incident date. The list of ways these diverge is short and each item is fatal to the claim: a staging host, a test tenant with permissive policy, disabled tools, a sanitised retrieval index, a guard stubbed out for cost, a substituted system prompt, a pinned model version production has since moved off. **4. Positive control.** Send a probe whose production response you can recognise — a refusal the stack reliably issues, a fact only the production index holds, a formatting quirk only the production system prompt produces — through the same target and the same scorer, and confirm it comes back. This is the decisive test, and it is one call. If the control does not surface, the clean result was never evidence about behaviour; it was evidence about plumbing. ### What it costs Steps 1 to 3 are free: a list comparison and two queries against stored data. Step 4 is a handful of metered calls. Re-running the suite with a bigger prompt set or a longer turn budget — the step most teams reach for first — is the expensive one: metered calls scale with attempts times turns times the target and scorer calls each turn makes, and it can burn a day of wall clock and a real budget line to produce another uninterpretable clean result. The cost argument alone justifies the ordering. ### Where the number misleads Two specific misreadings live here. First, **absent data counted as survived attempts**: a timeout or an auth rejection produces an empty or non-harmful body, a scorer labels it a non-hit, and it sits in the denominator of the attack-success rate exactly as a genuine refusal would. The reported rate is then understated, and the more broken the target the better the number looks — a metric that improves as instrumentation degrades. Second, **a summary hit count with no per-target breakdown**: one healthy target carrying the whole run looks identical to five healthy targets, and a run whose vulnerable surface was inert reads as a clean sweep. ### After reachability holds Only once all four checks survive do the content and judgement hypotheses become the leading ones, and they are three separate things with three separate owners: the seed prompt set never expressed the behaviour; the strategy's turn budget stopped short of it; or the scorer saw the behaviour and did not label it a hit. Test the last of those directly by replaying the incident transcript through the scorer — if it scores clean, the problem was never coverage. Then close the loop structurally. The surface the incident used becomes a permanently wired target with the incident as a regression case, and the run publishes per-target attempt classifications from that point on, so the next clean result carries a denominator that is larger and honest rather than smaller and flattering.
- What is a positive control here, concretely?A probe whose production response you can recognise — a refusal the stack reliably issues, or a fact only the production index holds — sent through the same target and scorer to prove the plumbing works end to end.
- The incident path was wired and attempts returned real responses. What next?Move to content and judgement: whether the prompt set expressed the behaviour, whether the turn budget stopped short of it, and whether the object deciding hits recognised it — three different owners.
A positive control is the test button on a smoke alarm. Silence from an alarm nobody has ever pressed the button on is not evidence that there is no smoke.
saying these in an interview costs you the question
- Jumping straight to writing more attack prompts
- Accepting a summary hit count without per-target attempt classification
- Never verifying the target pointed at production configuration
- Treating swallowed transport errors as evidence of a defence