skip to content

Four teams each wrote their own PyRIT prompt target for their own service, and each now reports an attack-success rate to you. What do you require before you are willing to compare those numbers across teams?

level: principalimportance: nice to knowfreq 22%

answer

  1. four adapters = four instruments
  2. conformance stub: refuse, comply, error, hang
  3. three buckets, denominator stated
  4. pin strategy, prompts, scorer, budget
  5. incentive to lose attempts quietly

basics

~20 s

Require every adapter to pass one shared conformance check: a known refusal, a known compliance, a forced error and a forced timeout, each landing in PyRIT's memory as a distinct expected outcome. Also require the same prompt set, scorer and turn budget. Otherwise the differences measure adapters, not services.

solid answer

~50 s

Four adapters written independently are four different measuring instruments, and an attack-success rate is comparable only when the instrument is held constant. What I would require: - **A conformance suite each adapter passes** against a stub endpoint the platform owns: a scripted refusal, a scripted compliance, a 500, a rate-limit rejection, an unterminated stream and a hang. Each must land as its expected outcome — the failures as execution errors, not clean negatives. - **A three-bucket report**: hit, not-hit, did-not-execute, with the third excluded from the denominator and stated. - **The same attack strategy, prompt set, scorer and turn budget**, pinned by identifier in the report. - **Round-trip fidelity**: the stored request equals what went on the wire. - **A declared session and identity model**, or a stated caveat. Without that I compare only within a team over time, across runs with identical pinned components.

go deeper

for a junior

Should notice that different adapters and different prompt sets make the numbers non-comparable.

for a middle

Names the concrete things to pin — prompt set, scorer, turn budget — and asks for an error bucket.

for a senior

Designs the conformance stub and the outcome mapping, and gates adapters on passing it before their numbers count.

for a principal

Owns the instrument as a platform asset, refuses single-number ranking, and names the incentive to lose attempts quietly as the thing the controls exist to catch.

### The unit that looks shared and is not An attack-success rate reads like a shared unit — a percentage, comparable across services the way latency is. It is not. Four teams' numbers can differ for at least six reasons before any real difference in security appears: different adapters, different prompt sets, different scorers, different turn budgets, different error accounting and different session hygiene. Ranking services on that is worse than not measuring, because the ranking rewards whichever team's adapter loses the most attempts. Every attempt that dies in transport and is stored as an empty response with `response_error` left at `"none"` counts as a clean negative and pushes that team's number down, which is to say: up the leaderboard. ### Standardise the instrument, not the number The thing that must be held constant is the measuring apparatus. Concretely, a platform team owns: - **One reviewed `PromptTarget` / `PromptChatTarget` template** carrying the error contract, the retry policy and the streaming assembly. Teams fill in only the endpoint-specific translation — auth, request schema, response parsing. - **One conformance stub**: a fake endpoint the platform runs that can be told to refuse, comply, return a 500, return a 429, stall mid-stream and hang past the timeout. Each case has one expected outcome in PyRIT's memory, and the failures must land as execution errors, not as responses. - **One pinned bundle** of attack strategy, seed prompt set, scorer and turn budget, referenced by identifier in every report. The stub is the highest-leverage artefact in the list, because it is the only one that proves an adapter's outcome mapping is honest without touching production, without spending a metered call, and without waiting for an engagement. Make passing it a gate for any adapter whose numbers enter a shared report. ### The three-bucket report Every report states hit, refused and did-not-execute, with the third excluded from the denominator and its count printed. That single requirement removes the incentive problem: once lost attempts are visible as their own number, an adapter that loses them stops looking better and starts looking broken. Alongside it, pin the bundle identifier, the surfaces exercised and — the part everyone omits — what was explicitly not tested. ### What it costs The stub and the template are roughly a week of one engineer, once, plus a day per team to adopt. Compare that with the alternative: discovering after four engagements that the numbers were never comparable, and re-running all four at full metered cost — an attempt meters the attacker model, the target and the scorer on every turn, so a modest suite per service is thousands of calls. The ongoing cost is a few minutes of CI per adapter change, which is negligible against the standing cost of a report nobody can defend in a review. ### Where the number misleads Two ways, and they compound. First, instrumentally: a lower rate may mean a worse adapter, not a stronger service. Second, semantically: even with a perfectly shared instrument, an attack-success rate compresses *which behaviours were tried* into a single figure, so a service that scored 12% against a hard prompt set is not obviously worse than one that scored 4% against an easy one. And the rate has no natural direction to improve in — a team can lower it by hardening the service, by weakening the prompt set, or by shipping an adapter that drops attempts, and the report looks identical in all three cases. ### What I would require, and what I would accept Require: conformance pass on file, pinned bundle identifier, three buckets with the denominator stated, and evidence of round-trip fidelity — the stored `converted_value` equals what went on the wire — rather than an assertion of it. When a team's rate moves sharply after an adapter refactor, the first question is whether the did-not-execute count and the response-length distribution moved with it; an adapter change that moves the number is an instrument change until proven otherwise. Accept, if none of this exists yet: comparisons within a single team over time, across runs with identical pinned components, plus a written line in every report saying cross-team comparison is not supported. That sentence is cheap and it prevents the specific meeting where four incomparable percentages become a ranked slide.

  • A team's rate improves sharply after they refactor their adapter. What do you ask first?
    Whether the did-not-execute count and the response-length distribution moved with it. An adapter change altering the number is an instrument change until proven otherwise.
  • What is the minimum you would standardise if you can only standardise one thing?
    The conformance stub and its expected outcome mapping. It is cheap, needs no production access, and it is what makes an honest error bucket enforceable.

saying these in an interview costs you the question

  • Ranking services on attack-success rates produced by different adapters and different prompt sets.
  • No did-not-execute count, so lost attempts silently improve a team's number.
  • Treating the rate as the deliverable instead of coverage plus what was not tested.
  • Letting each team pick its own scorer and then comparing across teams.
  • Assuming adapter correctness rather than gating on a conformance run.

context