skip to content

Two threats both average 6 in DREAD, one of them scoring Damage 1 and Affected users 10 - why is that ranking untrustworthy?

level: seniorimportance: should knowfreq 40%

answer

  1. the mean deletes the shape
  2. equal weight to unequal dimensions
  3. ordinal ranks, interval arithmetic
  4. same threat, different raters, different digits
  5. a decimal point implies false precision

basics

~20 s

Averaging five ordinal guesses destroys a threat's shape: a trivial issue touching everyone lands on the same digit as a moderate-everywhere one. The ratings are also unanchored, so a later session produces different digits from identical facts.

solid answer

~50 s

There are two separate defects and a strong answer names both. First, aggregation: the arithmetic mean weights all five dimensions equally and dilutes extremes, so `Damage 1, Affected users 10` averages to the same 6 as a threat that is a flat 6 on every axis - yet one is a broad, low-consequence annoyance and the other is a uniformly moderate risk, and you would fund them differently. The mean also blends impact with likelihood into a number that answers neither. Second, reproducibility of the scoring itself: no per-value criteria ship with DREAD, so the same team re-rating the same threat six months later routinely produces different digits from the same facts. A ranking that changes when the raters change is not a ranking. In practice I keep the five ratings visible, never sort on the collapsed mean, and force the damage and blast-radius conversation with the business owner.

go deeper

for a junior

Know that the DREAD score is an average of five estimates, and that an average can hide a very high rating on one dimension behind low ratings on the others.

for a middle

Explain both failures in your own words: equal weighting with dilution of extremes, and unanchored ratings that different people produce differently from the same facts.

for a senior

Demonstrate what you do about it in a real session - keep the vector, sort on damage and blast radius, and never let a decimal-point score travel into a ticket unexplained.

for a principal

Own the organisational consequence: a number nobody can reconstruct cannot survive a challenge from an engineering lead or an auditor, so decide deliberately whether numbers should leave the room at all.

## The scenario A print-shop order portal has two threats on the board after a modeling session. | Dimension | Threat A | Threat B | | --- | --- | --- | | Damage potential | 1 | 6 | | Reproducibility | 9 | 6 | | Exploitability | 8 | 6 | | Affected users | 10 | 6 | | Discoverability | 2 | 6 | | **Mean** | **6.0** | **6.0** | Threat A is a defect that lets any signed-in customer see the *titles* of other customers' print jobs in a shared queue view: almost no damage per record, works every time, and it touches the entire customer base. Threat B is moderate on every axis. DREAD says they are the same. Nobody who read the two sentences would say they are the same, and that gap is the whole lesson. ## Defect one: the arithmetic flattens the shape The mean has three problems here. **It weights every dimension equally.** DREAD asserts, by construction, that Discoverability matters exactly as much as Damage. No organisation actually believes that. A threat that destroys money or audit truth is not equal to one that is merely easy to find, yet a 1 in Damage is fully cancelled by a 10 elsewhere. **It dilutes extremes.** Risk work usually cares most about the tails - the one dimension that is off the scale is often the reason the threat matters. Averaging is exactly the operation that hides a tail. This is why teams that keep the five prompts frequently sort on the *maximum* or on Damage alone, and use the rest as context. **It performs interval arithmetic on ordinal judgments.** The ratings are ranks: a 10 is worse than a 5, but nothing establishes that the gap from 4 to 5 equals the gap from 9 to 10, or that Damage 8 and Exploitability 8 denote comparable amounts of anything. Averaging assumes both. The output has a decimal point and looks like a measurement, which is the dangerous part: false precision travels well into tickets and slide decks, and by the time it reaches a stakeholder the estimate has become a fact. **It mixes two axes.** Damage and Affected users are impact; Reproducibility, Exploitability and Discoverability are about how likely and how easy the attack is. Collapsing both into one digit means a 6 could mean "catastrophic but very hard" or "trivial but certain", and those two demand completely different responses - the first may be accepted, the second must be fixed this week. ## Defect two: the digits do not reproduce This is the failure that ended DREAD's mainstream use, and it is worth stating carefully because the word collides with one of the dimensions. DREAD's **R** is about whether the *attack* repeats reliably. The problem here is whether the *scores* repeat: hand the same threat to two engineers, or to the same team six months later, and you get different numbers. A managed-service provider team that re-scored one threat against the same design half a year on produced a different D-R-E-A-D vector from an unchanged set of facts. Nothing about the system had changed. What changed was who was in the room, what incident they had just lived through, and whether the person holding the pen felt like an 8 or a 6 that morning. Because no published per-value criteria exist - there is no document that says what Exploitability 7 means as opposed to 6 - the digits encode the raters, not the threat. The consequence is that DREAD scores are not comparable across teams, across time, or even across a long session as the group calibrates mid-flight. Sorting a backlog by a number with that property gives you the *feeling* of prioritisation with none of the substance, and it does not survive being challenged: when an engineer asks why their fix is ranked below another, "the average was 6.4" is not an answer anyone can inspect. ## What to do instead in the room - **Do not collapse the vector.** Keep the five ratings visible next to the threat. Most of the value is in the shape - a 10 in Affected users with a 1 in Damage tells a story the mean deletes. - **Sort on what you actually fund by.** For most teams that is damage and blast radius, with ease of attack as a tiebreaker, not an equal fifth of the answer. - **Write per-value criteria if you keep numbers.** If your organisation needs numbers to travel, define in writing what each rating means for *your* assets and services. That fixes the reproducibility problem, at the cost of admitting you have built your own scheme rather than adopted a standard one. - **Record the sentence, not just the digit.** "Every customer can see every other customer's job titles; the content stays private" survives a re-read six months later. A 6.0 does not. - **Ask what decision the number is for.** If it only decides sprint order among a handful of threats, a short ranked discussion is faster and more honest than five estimates and a division.

  • If you must produce a single number from the five dimensions, what aggregation would you defend?
    Taking the maximum, or ranking primarily on Damage with Affected users as a multiplier on it, preserves the tail that actually drives the decision. Either is defensible in a way the mean is not, because both keep an extreme visible. Whatever you choose, publish the rule and the per-value criteria alongside the number, so a reader can reconstruct how the digit was reached.
  • How would you show a sceptical team that the scoring is not reproducible?
    Run the cheapest possible experiment: have three engineers independently rate the same five threats before any discussion, then compare. The spread is usually wide enough to make the point without argument, and the disagreements themselves are useful - they surface that people are holding different assumptions about the asset and the attacker, which is worth more than the scores.
  • Does giving each dimension a written definition fix the problem?
    It fixes the reproducibility half, and it is the right move if numbers must leave the room. It does not fix the aggregation half: even perfectly anchored ratings still get flattened by an equal-weight mean, and the ordinal-to-interval leap remains. You need both a written rubric and a defensible combination rule, at which point you own a bespoke scheme and should say so.
  • Is a threat with Damage 1 and Affected users 10 ever worth prioritising?
    Sometimes, and that judgment is exactly what the mean hides. Broad low-severity exposure can matter for privacy expectations, regulatory notification or reputational trust even when each individual record is dull. The decision belongs to the business owner of the asset, informed by the shape of the threat - not to an averaging step that gives it the same digit as an unrelated moderate risk.

Averaging the five is like grading a building by the mean of five inspections: a structure that is perfect everywhere except the foundations scores the same as one that is merely mediocre throughout.

saying these in an interview costs you the question

  • Defending the mean because it produces one convenient sortable number
  • Confusing score reproducibility with DREAD's Reproducibility dimension
  • Treating an averaged estimate as a measurement of risk
  • Comparing DREAD scores between two teams as if they were calibrated
  • Adding raters to fix a problem caused by having no written criteria

context