skip to content

Why does reporting raw threat count as a threat modeling metric backfire?

level: principalimportance: nice to knowfreq 34%

answer

  1. who chooses the unit being counted
  2. the same analysis can be written many ways
  3. a measure that becomes a target
  4. splitting one entry into nine costs nothing
  5. a low count has two opposite meanings

basics

~20 s

Threat granularity is the analyst's choice, so the count has no fixed unit and is not comparable across teams. Once it becomes a target, the cheapest way to raise it is splitting one threat into several.

solid answer

~50 s

Because the unit is chosen by the person being measured. One replayed-session-token threat can be written as one entry or split into four, and both can be defensible analysis, so counts are not comparable across teams or stable within one. Make it a target and the failure completes itself: at a media company that reported per-team threat counts to the CTO monthly, teams started splitting one threat into nine entries within two cycles - the classic Goodhart effect, where the cheapest way to move a measure is to change the measuring. Triage then fills with near-duplicates and closure rates decay over padded populations. The number is directionally ambiguous too: a low count can mean a clean design or a rushed session. I report coverage of tier-1 designs, high-severity findings closed with linked changes and mean time to close, and escaped issues.

go deeper

for a junior

Understand that threats can be written up at different levels of detail, so counting them is not like counting defects. Be able to say why two teams' counts cannot be compared directly.

for a middle

Explain the mechanism, not just the conclusion: granularity is an analyst choice, so the unit floats, and a target makes splitting the cheapest improvement. Know that the count is also ambiguous in direction even when nobody games it.

for a senior

Show the downstream damage you have to clean up - diluted triage, closure rates computed over padded populations - and describe the replacement set you would report instead, with the sampling that keeps it honest.

for a principal

Own the reporting relationship. Be ready to withdraw a metric an executive already likes, to explain the incentive in one sentence they will accept, and to defend a slower, less flattering metric set because its cheapest path to improvement is the behaviour you want.

Raw threat count - the number of threats a team enumerated - is the metric a new programme reaches for first, because it is trivially available and it goes up. It is also the fastest way to teach an organisation to produce worse threat models. ### The unit is not a unit Threat granularity is an analyst's choice, not a property of the system. *An attacker replays a captured session token* can be written as one threat, or split into capture on the wire, capture from logs, capture from a browser extension, and reuse after logout. Both descriptions can be defensible analysis. That means the metric has no fixed unit: two teams doing equally good work can report numbers that differ by a factor of five, and one team can change its number without changing its analysis at all. Any metric whose unit is set by the person being measured is not comparable across teams and not stable over time. ### Making it a target completes the failure When a media company began reporting each team's raw threat count to the CTO monthly, teams learned within two cycles to split one threat into nine entries. Nothing about the systems changed; the reports improved. This is the effect usually summarised as Goodhart's law: once a measure becomes a target, the cheapest way to move it is to change the measuring, not the thing measured. The damage is not only cosmetic. Inflated splitting has real costs: - Triage load rises with no increase in information, so the serious findings are diluted among near-duplicates. - Follow-through metrics decay, because closure rate and mean time to close are now computed over a population padded with trivia. - Analysts learn that volume is the deliverable, which is the opposite of the judgment the practice needs. - Cross-team comparison becomes noise, and the teams that resisted splitting look worst. ### The direction of the number is ambiguous anyway Even unmanipulated, a threat count cannot be read. A low count may mean a genuinely simple and well-segmented design, a session that was rushed, or a scope drawn too narrowly. A high count may mean a rich attack surface, a scope drawn too wide, or an analyst who enumerates variants. A metric that cannot tell good news from bad news without opening the model is not a metric; it is a prompt to go and read the model. ### The mirror-image gaming Counting *open* threats instead of found threats inverts the incentive but not the problem: the cheapest way to reduce open threats is to close them as accepted risk without changing anything. Any count used as a target will be met by the least expensive available behaviour. If you count open findings, you must sample how findings were closed, or you have simply moved the gaming downstream. ### What to report instead Report the numbers whose cheapest path to improvement is the behaviour you actually want: | Instead of | Report | Because the cheapest way to move it is | | --- | --- | --- | | Threats found | Coverage of tier-1 designs, with a stated denominator | Model something that was not modeled | | Threats found | High-severity findings closed with a linked change | Ship a fix | | Open threats | Mean time to close, per severity | Close things sooner, not looser | | Any count | Escaped issues and what each taught the practice | Actually miss less | Add the input no counter provides: read a sample of models each quarter and judge whether the threats found are the ones a competent attacker with the assumed position would choose. A programme with modest counts and sharp models is in far better shape than the reverse. ### When a count is still useful Volume is legitimate as an internal diagnostic, never as a target and never in a league table. Forty threats out of one session usually means the scope was too large to reason about. Three threats on a service that fronts the internet and holds customer records usually means the session was shallow or the boundaries were drawn too tightly. Used that way the number is a trigger to look, which is the only role it can honestly play. ### Saying this to an executive The argument that lands is not *the metric is philosophically flawed*; it is that the number they are being shown is set by the people reporting it, and it will improve regardless of whether the systems do. Offer the replacement in the same breath: coverage of the systems that matter, closure of the severe findings, and how many design issues escaped. Those move only when the estate moves.

  • Does counting open threats instead of found threats fix the incentive?
    No, it inverts it. The cheapest way to reduce open threats is to close them as accepted risk without changing the system, so the gaming just moves downstream. Any count used as a target gets met by the least expensive available behaviour, which is why I pair a closure metric with a sample review of how findings were actually closed.
  • Is there any legitimate use of threat volume?
    Yes, as an internal diagnostic and never as a target or a league table. Forty threats from one session usually means the scope was too large to reason about; three threats on an internet-facing service holding customer records usually means the session was shallow or the boundaries were drawn too tightly. The number tells you to go and read the model, nothing more.
  • How do you make this argument to an executive who likes the chart?
    Not as a philosophical objection. I say plainly that the number is set by the people reporting it and will improve whether or not the systems do, then offer replacements in the same breath: coverage of tier-1 designs, high-severity findings closed with a linked change, mean time to close, and escapes. Those only move when the estate moves.

Paying a novelist by the word does not produce more story. It produces longer sentences, and the honest writers come last in the ranking.

saying these in an interview costs you the question

  • Treats a rising threat count as evidence the practice improved
  • Compares teams by threats found without questioning granularity
  • Assumes a low threat count means a well-designed system
  • Swaps to counting open threats and calls the incentive fixed
  • Defends the metric because it is easy to collect automatically
  • Reports volume in a per-team league table

context