skip to content

You own the severity rubric for AI findings across an organisation where every control the findings target fails some fraction of the time. How do you let a measured success rate into the rating, and what do you refuse to let it do?

level: principalimportance: should knowfreq 35%

answer

  1. impact never probability-weighted
  2. rate = last, capped term
  3. floor for any reproducible success
  4. counts + point + date mandatory
  5. name the compensating control

basics

~20 s

Let the rate modify exploitability only, alongside retry cost and who can reach the entry point. Refuse to let it touch impact, refuse to let it gate whether a finding gets a severity at all, and set a floor so any reproducible success keeps a minimum rating. Require counts, a measurement point and a date on every filed rate.

solid answer

~50 s

Four rules carry most of the weight: 1. **Impact is scored from one success, never probability-weighted.** This stops rare, high-consequence findings from being buried by arithmetic that looks rigorous. 2. **The rate enters exploitability, bounded.** It may move that band by at most one step, and only together with retry cost — authentication, throttling, quotas, alerting. 3. **A floor.** Any reproducible success at a reachable entry point cannot fall below a defined minimum, however low the rate; otherwise the rubric invites arguing a finding to zero by shrinking the sample. 4. **Mandatory metadata.** Counts rather than a bare percentage, the success criterion, the measurement point, the date. A rate lacking them does not enter the score. Then two operational rules: bands coarse enough to survive realistic sample sizes, so teams stop arguing over noise; and any severity that depends on a compensating control must name it, so removing the control triggers a re-rate rather than silent drift.

go deeper

for a junior

Not expected to design a rubric; should know the rate is not a discount on impact.

for a middle

Can state where the rate belongs and what metadata a filed rate needs.

for a senior

Applies the rules under pressure, refuses probability-weighted impact, and records the compensating controls a rating depends on.

for a principal

Designs against the incentives — sample shopping, demonstrability bias, stale ratings — and validates the rubric by reviewing the ordering of closed findings.

The purpose of an organisation-wide rubric is not precision. It is a *defensible and comparable ordering* of a fix queue produced by different people, using different tools, against targets that all fail some fraction of the time. Every design choice below follows from that goal. ### The rules that carry the weight 1. **Impact is scored from a single success and is never probability-weighted.** Write the ban into the rubric text, not just into training, because the multiplication is intuitive and will otherwise reappear every quarter. Define impact levels by *outcome*: cross-tenant data, regulated or dangerous content delivered to a user, an action taken on a user's behalf, reputational-only, cosmetic. 2. **The rate enters exploitability only, as the last and weakest term.** Reachability comes first — who can reach the entry point, authenticated or not — then retry cost, then the measured rate, capped so it can move the band by at most one step and only in conjunction with a retry-limiting control. 3. **A floor.** Any reproducible success at a user-reachable entry point cannot fall below a defined minimum, however low the rate. 4. **Mandatory metadata on any filed rate:** successes over attempts (not a bare percentage), the success criterion and who validated it, the measurement point, the date, and the compensating controls the rating assumes. A rate lacking them does not enter the score at all. ### The incentives you are designing against - **Sample shopping.** If the rate lowers the score, someone will run twenty attempts and file the number that suits their argument. Mandatory counts plus the floor remove the payoff. - **Judge shopping.** The cheapest lever of all, and invisible without rule 4: swap to a stricter detector or grader and the same transcripts yield a lower rate, with no dishonesty required and often no awareness that anything happened. - **Demonstrability bias.** A finding that reproduces on demand feels severe. Make reproducibility a confidence tag on the evidence, explicitly not a score input. - **Stale scores.** Both the model and its wrapper move. Every rating built on a measured rate carries an as-of date and a re-measure obligation before the finding may be closed as fixed. ### What the rubric costs the organisation Each mandatory field is time on every finding, and the bill is not evenly distributed. Counts are free — the tool already has them. A declared measurement point and a date are free. **Judge validation is the expensive one**: a stratified human read of transcripts, per configuration, an hour or two of senior time each. So scope it — require validation only where the rate was load-bearing, meaning it actually moved a band; elsewhere require only that the criterion be *named*. Coarse bands cost you resolution and buy you fewer arguments, and that trade is right whenever samples are small, which on real engagement budgets is always. Calibration reviews cost a couple of hours a quarter and are the only way you learn whether two teams produce the same answer. ### Where a rubric's own numbers mislead A weighted formula that outputs "6.8" implies a precision none of its inputs have: a rate off a 20-attempt run enters the formula and leaves as a decimal that no one can question, because arguing with a decimal feels like arguing with arithmetic. Prefer small ordinal bands over computed scores. And treat cross-team comparability as a claim to be tested rather than an achievement of publishing the document — have two teams rate the same three findings blind and compare, because a rubric everyone reads differently is worse than an explicit convention everyone knows is rough. ### What to accept You will not get statistical rigour on real engagement budgets, and chasing it burns hours better spent widening what was tested. Design for robustness to weak samples instead: the rubric should produce the same ordering whether the tester ran 20 attempts or 200, with only the exploitability nudge differing. ### How you know it works Sample closed findings and ask whether the ordering still looks right in hindsight. Look specifically for inversions — a reliable-but-cosmetic item fixed before a rare disclosure. Count how often a severity was argued down after a rerun, and check whether that rerun was *larger* or *smaller* than the original; smaller is sample shopping and should be treated as a process defect. And verify that findings whose rating named a compensating control were actually re-rated when that control changed, because a dependency nobody revisits is a rating that quietly expired.

  • Why a floor rather than just guidance?
    Because without one, a low rate is an argument anyone can make to talk a finding down, and the cheapest way to get a low rate is a small sample. The floor removes the incentive.
  • How do you keep two teams from producing different severities for similar findings?
    Coarse shared bands, impact levels defined by outcome rather than by judgement words, mandatory metadata, and periodic calibration reviews over a sample of closed findings from both teams.
  • A team argues their finding should drop because a classifier now blocks most attempts. Do you re-rate?
    Yes, but only on a re-measurement at the production entry point, and the new rating must name the classifier as the control it depends on so that disabling it re-opens the severity.

saying these in an interview costs you the question

  • A formula that multiplies impact by the measured success rate
  • A minimum rate required before a finding is eligible for any severity
  • Per-team rate-to-severity mappings, so scores cannot be compared across the queue
  • Rubric bands finer than the sample sizes teams can realistically collect
  • No as-of date or re-measure obligation on a rating built from a measured rate

context