A red-team finding you are reviewing states only "attack success rate: 15%" for an injection attempt against a hosted chat feature. What else has to sit next to that number before anyone can rate the finding's severity?
answer
- no denominator, no rate
- successes over attempts
- who called it a hit
- which endpoint, which config
- dated: hosted targets move
basics
~20 sThe counts behind it: how many attempts and how many succeeded, because 3 of 20 and 150 of 1000 are not the same evidence. Also what decided that a response counted as a success, which endpoint and configuration it was measured against, and when. Without those, 15% is a number nobody can re-derive.
solid answer
~50 sA bare percentage hides its denominator, and the denominator is most of the story: 3 of 20 is compatible with a true rate anywhere from a few percent to nearly forty, while 150 of 1000 is a rate you can actually argue about. Four things travel with the figure: - **Raw counts** — successes over attempts, not just the ratio. - **The success criterion** — which automated judge or human review call decided a response counted as a hit, and whether partial compliance counted. - **The measurement point** — the exact endpoint, model configuration, sampling settings, system prompt, and whether any input or output classifier sat in front of it. - **The date** — a hosted endpoint can change under you between the run and the triage meeting. A reviewer who cannot re-derive your number cannot argue with your severity, so they either rubber-stamp it or throw the finding out. Both are bad outcomes for the fix queue.
go deeper
Should say the number of attempts and the number of successes must be reported, not just the percentage.
Adds the success criterion and the exact endpoint/configuration, and can explain why 3 of 20 is weak evidence.
Ties each missing item to how it would change the filed severity, and insists the judge's calls were human-sampled.
Makes the four items mandatory fields in the finding template so the fix queue can be compared across teams and tools.
A reported success rate is a compressed artefact of a measurement, and severity work needs the measurement, not the compression. "15%" is the last three characters of a process with at least four decisions in it, and every one of those decisions can move the number by more than the rate itself. ### What the fraction is made of The numerator is the number of attempts that *something* ruled a success. That something is a **judge** — a component that reads a model response and returns a verdict. In this tooling space a judge is a garak detector class, a PyRIT scorer, a promptfoo grader, or a human reading transcripts, and none of those four applies the same standard. The denominator is equally a decision: an "attempt" can be one prompt sent once, one prompt sent several times because the runner was told to repeat it (garak's `--generations` flag re-sends each probe prompt N times), or a whole multi-turn conversation counted as a single try. A bare percentage tells a reader neither half. ### Why the denominator dominates Uncertainty on a proportion shrinks with the square root of the sample. Concretely: | observed | headline | 95% interval (Wilson) | |---|---|---| | 3 of 20 | 15% | roughly 5% – 36% | | 30 of 200 | 15% | roughly 11% – 21% | | 150 of 1000 | 15% | roughly 13% – 17% | The first of those cannot distinguish "rare" from "routine". The third can. They print identically. If your rubric lets the rate move the score at all, the sample size decides how much movement is justified, so omitting it asks the reader to trust a point estimate they cannot bound. ### What the sample costs Every attempt is a metered call. A thousand attempts at a couple of thousand tokens round trip is single-digit to low tens of dollars on a commercial endpoint, and the wall clock is set by the provider's rate limit rather than by the model — at a couple of requests per second with backoff, about ten minutes. That part is cheap. The expensive part is human: validating the judge means a person reads a stratified sample of transcripts in both directions (calls it scored as hits, and calls it scored as misses), which is an hour or two of senior time *per configuration*. And the arithmetic is unkind — halving the width of the interval costs roughly four times the attempts. "Just run more" therefore has a price, and it should be paid once, deliberately, rather than after every argument in a triage meeting. ### Where the number misleads - **Judge drift.** A detector that scores any non-refusal as a hit and one that requires the restricted content to actually appear will differ by a factor of several on identical transcripts. Two testers reporting 40% and 8% on the same feature in the same week are usually reporting two definitions, not two systems. - **Correlated attempts.** If a run repeated each of 20 prompts five times, the report divides by 100 — but those are not 100 independent tests. Successes cluster on the prompts that work, so the effective sample is nearer 20 and the true interval is far wider than a denominator of 100 implies. The rate looks better measured than it is. - **Deduplication.** If near-identical variants were collapsed, the numerator shrank; if they were not, one lucky family of phrasings can carry the entire rate on its own. - **The entry point.** A rate measured against a raw model API with the tester's own permissive system prompt is not the product's rate, because the product's template, field limits and classifiers are not in the path. - **The date.** Hosted endpoints get reversioned and guardrails get rolled out with no announcement, so a rate is a measurement of a moving target and is only meaningful as of a moment. ### What to check before you file Re-derive the percentage from the raw counts and confirm it matches. Confirm which judge issued the verdicts, and that a human sampled its calls both ways. Confirm what one attempt was, and whether repeats of the same prompt were counted separately. Confirm the model identifier, decoding settings, system prompt and any classifier in front of the endpoint, and that none of them moved mid-run. Stamp the date. Then put those items in the finding body, where the reviewer reads them — not in a chat thread that dies next week. A reviewer who cannot re-derive your number cannot argue with your severity, so they either rubber-stamp it or bin the finding, and both outcomes misorder the fix queue.
- Two testers report 40% and 8% on the same feature in the same week. What is the first thing you check?Whether they used the same definition of a success. A judge that counts any non-refusal and one that requires the restricted content to actually appear will disagree by exactly this kind of margin before you look for anything else.
- Why does the date belong on the number?A hosted endpoint can be reversioned, or a guardrail rolled out, without any announcement. A rate is a measurement of a moving target, so it is only meaningful as of a moment.
saying these in an interview costs you the question
- Filing a percentage with no attempt count and treating it as precise
- Not knowing whether an automated judge or a human decided each hit
- Reporting a rate measured against a raw model API as if it were the product's rate
- Rounding 3 of 20 to "about 15%" and dropping the counts entirely