skip to content

Rating Severity

Rating an AI finding means scoring a control that fails only some of the time, so the observed success rate has to sit inside the score. Interviewers ask because a bad rating misdirects the whole fix queue.

on this pageshow

explore

questions

6

An automated red-team run got a hosted chat endpoint to return another user's private record in 3 of 20 attempts. When you write the severity, does the 3-in-20 belong in the impact score, in the likelihood/exploitability score, or in both?

level: middleimportance: must knowfreq 65%

answer

  1. impact from one success
  2. rate lives in exploitability
  3. never multiply impact by rate
  4. 1-(1-p)^n: retries win
  5. throttling is what lowers it

basics

~20 s

Likelihood, not impact. Impact is what one success does, and one success already leaks the whole record — a 15% rate does not make the data 15% leaked. The rate belongs on the exploitability side, and even there it counts for little when an attacker can simply retry cheaply.

solid answer

~50 s

Split the two axes cleanly. **Impact** is scored from a single successful attempt, at its worst realistic outcome. One success discloses one record in full; the observed rate cannot dilute that. Multiplying impact by an observed rate is the classic mistake here, and it systematically buries the findings that matter, because the highest-impact failures are usually the rarest. **Likelihood/exploitability** is where the rate goes, alongside who can reach the entry point, whether authentication is required, and what an attempt costs. But it is weak input: 3 in 20 means an attacker sending a few dozen requests succeeds essentially every session. The rate only pulls the score down meaningfully when something makes retries expensive or visible — per-account rate limits, quotas, alerting on repeated refusals, a human in the loop. The practical consequence: for a probabilistic control, the score is driven by impact plus attacker retry economics, and the measured rate is a modifier inside exploitability, never a discount on consequence.

go deeper

for a junior

Should at least say the rate is about how often, not how bad, and that one success is a full disclosure.

for a middle

Places the rate on the exploitability axis, explains why multiplying impact by it is wrong, and knows retries compound.

for a senior

Computes the retry arithmetic, identifies rate limiting and alerting as the controls that actually move the score, and records that dependency in the finding.

for a principal

Writes the rubric so impact can never be probability-weighted, and audits filed findings for that failure across teams.

This question separates people who have rated AI findings from people who have only rated deterministic ones. With a memory-safety bug the exploit either fires or it does not, and a rate axis never appears; with a probabilistic control it always does, and the first instinct — multiply — is wrong. ### Two axes, two different questions **Impact** answers: *if this succeeds once, what has happened?* It is a legal, regulatory, contractual and customer-trust question, and there is no probability in it. **Exploitability** (called likelihood or reachability in some rubrics) answers: *how easily does an attacker obtain that one success?* Its inputs are who can reach the entry point, whether authentication is required, what one attempt costs the attacker, and — last and weakest — the measured rate. ### Why probability-weighting impact is a category error Multiplying a consequence by an observed rate invents "20% of a data breach", which is not a thing that exists. The record either left the system or it did not; a notification obligation does not scale with how many attempts it took to trigger. Worse, the distortion is not random. It is anti-correlated with what matters: the reliable failures tend to be the shallow ones, because the obvious defects were fixed long ago, so the rarest findings are disproportionately the ones with real consequence. A probability-weighted score therefore demotes exactly the findings that become incidents, while looking rigorous while it does so. ### What the rate actually measures The rate is a per-attempt probability, conditional on your specific attempt set, judge and entry point. An attacker is not limited to one attempt, and independent retries compound as 1-(1-p)^n: ``` p = 0.15 n: 1 5 10 20 50 P: 0.15 0.56 0.80 0.96 0.9997 ``` So the honest reading of 3 in 20 is not "unlikely" but "reliable for anyone who retries". A defender must stop essentially every attempt forever; an attacker needs one. ### Cost is the term that moves the score, not p Twenty attempts against an unauthenticated public endpoint is a few seconds of scripted requests, and the *defender* pays for the inference, not the attacker. That asymmetry is why a low rate on an open endpoint is nearly worthless as a mitigation. The rate only earns weight when something bounds n or makes those attempts visible: a per-account throttle, an authenticated account that can be banned, a paid tier, a proof-of-work or challenge, alerting on bursts of refusals, or a human approving the resulting action. Under a limit of, say, five attempts per hour per account, p=0.15 turns "seconds" into "about an hour and a burned account", and that is a genuine change in exploitability — but you are now scoring the compensating control, and the finding must say so by name, because when the control is disabled or misconfigured the severity moves straight back up. ### Where the number misleads - **Reading per-attempt p as per-attacker p.** "Only 15%" is the single most common misreading in a triage meeting. - **Assuming the retries are interchangeable.** If all three successes came from one phrasing family, an attacker holding that phrasing is not at 15% — they are close to certain, and the aggregate rate understates them. If instead the successes are scattered across many prompts, the compounding argument holds as written. Look at *which* attempts succeeded, not just how many. - **Counting a suppressed output as a failure.** If an output classifier blocked a response the model had already produced, the model failed and a control saved you. That belongs in the finding as its own outcome, not folded into the miss count. - **Using the rate as fix verification without re-running the identical set.** A different attempt set produces a different number and proves nothing about the mitigation. ### What you would check Whether the entry point is authenticated. Whether the per-account and per-IP limits are actually enforced — test them, do not read the config. Whether anything alerts on repeated refusals from one principal. Whether the successes cluster on one prompt family. Then write it down in this shape: impact scored from the single worst success; the rate with its counts on the exploitability line; the retry cost stated explicitly; and every compensating control the reduced score assumes, named, so a reviewer can see exactly which of them a fix would change.

  • At a 15% per-attempt success rate, how many attempts before an attacker is more likely than not to succeed?
    About five. 1-(0.85^5) is roughly 0.56. That is why a rate in this range does almost nothing to the likelihood of an attacker who is allowed to retry.
  • What single fact would most reduce the exploitability score here?
    That the entry point is authenticated and rate-limited per account, with repeated refusals alerted on — that raises the cost and visibility of the retries the attack depends on.
  • Where does the rate matter most in the finding, then?
    In the fix verification: the team needs a measured before-and-after rate on the same harness to show the mitigation actually moved something, even though the severity was driven by impact.

A 15% per-attempt rate is like a lock that opens on roughly one try in seven: the burglar does not give up after the first failure, and nothing about the low rate makes what is in the room smaller.

saying these in an interview costs you the question

  • Scaling the impact score by the observed success rate
  • Calling a 15% rate "low likelihood" without asking what an attempt costs the attacker
  • Treating a rare-but-total data disclosure as lower severity than a frequent cosmetic failure
  • Claiming the finding is not severe because the tester could not reproduce it every time

context

open as a page

Your success rate was measured by calling a model API directly with your harness's own system prompt. Production wraps that same model in a fixed template, a 300-character user field and an input classifier. What does that do to the severity you file?

level: seniorimportance: must knowfreq 50%

basics

~20 s

It means your rate describes your harness's entry point, not the product's. Say so explicitly in the finding, then re-measure through the production path before you finalise the rating. If you cannot re-measure, file the number with the measurement point stated and rate the reachability separately, rather than quietly discounting it.

open as a page

A red-team finding you are reviewing states only "attack success rate: 15%" for an injection attempt against a hosted chat feature. What else has to sit next to that number before anyone can rate the finding's severity?

level: juniorimportance: should knowfreq 50%

basics

~20 s

The counts behind it: how many attempts and how many succeeded, because 3 of 20 and 150 of 1000 are not the same evidence. Also what decided that a response counted as a success, which endpoint and configuration it was measured against, and when. Without those, 15% is a number nobody can re-derive.

open as a page

You reran the same 20-attempt red-team script against the same hosted chat endpoint the next day: the first run succeeded twice, the rerun six times. How should the finding's severity handle a measured success rate that moved from 10% to 30%?

level: middleimportance: should knowfreq 45%

basics

~20 s

Do not rerate on it. At 20 attempts both results sit inside the same wide confidence interval, so the swing is what sampling noise looks like, not evidence the target got worse. Report both runs with their counts, widen the band you quote, and if the rate really drives the score, collect a much larger sample.

open as a page

Two findings against the same chat feature: A succeeds in 18 of 20 attempts and returns the assistant's own refusal-policy wording; B succeeds in 1 of 50 attempts and returns another customer's record. Which do you rate higher, and what do you tell the team about B's 2%?

level: seniorimportance: should knowfreq 55%

basics

~20 s

B, clearly. Impact decides the order: one success in B is a real disclosure of another customer's data, while A leaks the product's own policy text. Tell the team that 2% is not rare for an attacker — around fifty cheap retries make success roughly even odds, and nothing in the rate makes the leak smaller.

open as a page

You own the severity rubric for AI findings across an organisation where every control the findings target fails some fraction of the time. How do you let a measured success rate into the rating, and what do you refuse to let it do?

level: principalimportance: should knowfreq 35%

basics

~20 s

Let the rate modify exploitability only, alongside retry cost and who can reach the entry point. Refuse to let it touch impact, refuse to let it gate whether a finding gets a severity at all, and set a floor so any reproducible success keeps a minimum rating. Require counts, a measurement point and a date on every filed rate.

open as a page