Two findings against the same chat feature: A succeeds in 18 of 20 attempts and returns the assistant's own refusal-policy wording; B succeeds in 1 of 50 attempts and returns another customer's record. Which do you rate higher, and what do you tell the team about B's 2%?
answer
- impact orders the queue
- 18/20 of nothing is still nothing
- 2% x 100 tries = 87%
- reproducible is not severe
- score depends on throttling
basics
~20 sB, clearly. Impact decides the order: one success in B is a real disclosure of another customer's data, while A leaks the product's own policy text. Tell the team that 2% is not rare for an attacker — around fifty cheap retries make success roughly even odds, and nothing in the rate makes the leak smaller.
solid answer
~50 s**Order by impact first.** A is high-frequency and near-zero consequence — arguably a product polish item, not a security finding. B is low-frequency and high consequence: cross-customer data disclosure, with notification implications, from a single success. **Then correct the reading of 2%.** The instinct is to treat 1-in-50 as an edge case. The retry arithmetic says otherwise: at p=0.02, fifty attempts is about a 64% chance of at least one success and a hundred about 87%. On an unauthenticated, unthrottled, unmonitored endpoint those attempts cost nothing. The rate does not shrink the blast radius; it describes how long the attacker waits. **Say what would actually change the rating.** Per-account limits, quotas, alerting on bursts of near-misses, an authenticated session you can revoke. Record that the severity is conditional on them. And note 1 of 50 is a thin sample — an order-of-magnitude statement, not a figure to argue bands over.
go deeper
Should say B is worse because it exposes another customer's data, even though it happens less often.
Adds the retry arithmetic showing 2% is reliably reachable, and separates reproducibility from consequence.
Names the compensating controls that would genuinely lower B, records the rating's dependence on them, and resists rubric arguments over a 1-of-50 estimate.
Ensures the org's rubric cannot invert this pair, and reviews closed findings for cases where demonstrability drove the priority.
The value of this pair is that it forces the two axes apart in a case where a naive score inverts them. One finding is easy to demo and worth nothing; the other is hard to demo and is the one that generates a notification obligation. ### Order by consequence **Finding A** returns the assistant's own refusal-policy wording. That is the product's own boilerplate — text that any user can elicit by being refused. Disclosing it costs the business essentially nothing; it is arguably a product-quality item rather than a security finding at all. **Finding B** returns another customer's record: cross-tenant data disclosure, with regulatory notification clocks, contractual exposure, and a customer-trust cost that starts the moment one success happens. Impact is scored from a single success, so B outranks A and it is not close. ### Why A feels worse than it is A reproduces 18 times out of 20. It screenshots well, it demos live without sweating, and the write-up is easy. That is **demonstrability bias**, and it is one of the strongest distortions in a real fix queue. Reproducibility is a property of your *evidence*, not of the consequence. Exploitability of a worthless outcome is still worthless: a control that reliably fails to protect something nobody wants is a cosmetic defect at a high rate. ### The retry reading of 2% B's rate compounds as 1-(1-p)^n at p=0.02: ``` n: 25 35 50 100 200 P: 0.40 0.51 0.64 0.87 0.98 ``` On an unauthenticated, unthrottled, unmonitored endpoint, 200 scripted requests is seconds of attacker time and costs the attacker nothing — the defender pays for the inference. So "1 in 50" is a *waiting time*, not a barrier. The attacker's real constraint is the throttle, not p, and if there is no throttle there is no constraint. The rate does not shrink the blast radius; it describes how long someone waits for it. ### How thin 1 of 50 actually is The 95% interval on 1/50 runs from well under 1% to around 10%. Arguing a rubric band between "2%" and "5%" on that sample is arguing about nothing. Pinning B's rate to within a point would take thousands of attempts against production: hours to days of throttled traffic, several accounts, a likely abuse-detection incident, and a judge-validation pass over the transcripts. Nobody should buy that precision to rate a finding whose impact already sets the score. Spend those hours on the questions that actually move the rating instead — whose record came back, how the victim was selected, whether an attacker can *target* a chosen victim rather than receiving a random one, and how many records one success yields. Targetability and volume are genuine impact modifiers; a third significant figure on p is not. ### Where the numbers mislead - **"2%, so informational."** A rate threshold used as a gate on severity is exactly the rule that buries rare, high-consequence findings. - **The opposite error:** rating B down because "rate limiting will stop 200 attempts", without testing the limit and without recording the dependency. If the limit is per-IP only, or applies to a different route, or is disabled during an incident, the score was never real. - **Reading 18/20 as strength.** High reproducibility on A is evidence that A is easy, not that A is important. - **Pooling the two rates into a comparison.** They were measured with different attempt counts and possibly different judges; the comparison that matters is between consequences. ### What to hand the team B rated on impact, with the retry arithmetic written out in the finding so nobody can misread 2% as "unlikely". The compensating controls the rating assumes, named, so that removing one triggers a re-rate. The counts, not the percentage, so the thinness of the sample is visible. And A filed at its true, low severity, with an explicit note that it may be fixed first if it is cheap — provided it does not displace B in the queue. Queue order follows consequence; ease of fixing is a scheduling detail, not a severity input.
- Roughly how many attempts before B's attacker is more likely than not to succeed?About 35. 1-(0.98^35) is a bit over 50%. On an unauthenticated, unthrottled endpoint that is a few seconds of scripted requests.
- The team wants to close A first because it is easy. What do you say?Fine if it is cheap and does not displace B, but say plainly that A is a product-quality item and B is the one with a disclosure consequence — the queue order should follow impact, not ease.
- What single mitigation would most reduce B's rating?Making the retries expensive and visible: authentication with per-account throttling plus alerting on repeated near-miss attempts, so the dozens of attempts the attack needs cannot be made quietly.
saying these in an interview costs you the question
- Ranking A above B because it reproduces reliably
- Calling B "informational" or "edge case" purely because of the 2%
- Arguing rubric bands over a 1-of-50 estimate as if it were precise
- Rating B down for the rate limits without recording that the score depends on them