skip to content

Measuring the Practice

Coverage of services modeled, findings closed and time-to-model show whether the practice works; counting threats found rewards noise. Interviewers ask which number you report upward.

on this pageshow

questions

4

Which metrics show whether a threat modeling practice is actually working?

level: middleimportance: must knowfreq 57%

answer

  1. reach, follow-through, speed, outcome
  2. one number for how far it spreads
  3. closed findings beat found findings
  4. split elapsed time into queue and work
  5. only escapes are a lagging measure

basics

~10 s

Four families: coverage of in-scope designs and services against a stated denominator, findings closed and mean time to close, time-to-model from trigger to owned findings, and escaped issues found after the design passed.

solid answer

~50 s

I measure four things. **Coverage**: what share of in-scope designs or services has a current model, weighted by criticality, with the denominator spelled out. **Follow-through**: how many findings are closed with a linked change, and the mean time to close per severity - this is what proves modeling reaches the code rather than a document. **Time-to-model**: elapsed time from trigger to a finished model with owned findings, decomposed into queue time versus analysis time, because a slow number usually means facilitation capacity, not a heavy method. **Escapes**: design-level issues found later in production or an incident on a system that was modeled. The first three are leading activity signals, escapes are the only lagging outcome. I deliberately do not report raw threat counts, sessions held or hours spent, and I add a qualitative read of a few models a quarter, since no counter tells you whether the threats found were the right ones.

go deeper

for a junior

Be ready to name what gets counted and why: how much of the estate is modeled, how many findings actually got fixed, how long a model takes, and what slipped through. Knowing the four families by name is enough at this level.

for a middle

You are expected to explain what each number exposes and to separate leading activity signals from the one lagging outcome. Practise decomposing time-to-model into queue versus analysis, and explaining why closure matters more than count of findings.

for a senior

Show that you instrument this without building a reporting programme: timestamps on the model record, a link from every finding to its fix ticket, and a quarterly sample read. Be ready to say which number you would act on first when two disagree.

for a principal

Own the choice of what goes on an executive slide and what stays an internal diagnostic. Argue the tradeoff between metrics that are cheap and gameable and metrics that are honest and slow, and defend the qualitative sample as a first-class input rather than a nicety.

A threat modeling practice is a **process**, not an artifact, so the honest question a lead or an executive asks is whether it changes the systems being built. Before the metrics, fix the vocabulary they measure: a **threat** is what could go wrong (an attacker abusing a flow), a **vulnerability** is the concrete flaw that lets it happen, a **risk** is the rated consequence of that threat given the system as it stands, and a **control** is what you do about it. A metric that counts threats measures words produced; a metric that counts closed findings measures systems changed. Four families of measurement cover a practice, and each exposes a different failure. ### 1. Coverage - how far the practice reaches The share of in-scope designs or services that have a current model. It is a *leading* indicator: it tells you whether the practice is being applied at all, long before any outcome shows up. Coverage is only meaningful with a stated denominator (which population of services counts) and a stated definition of *modeled* (a diagram exists, versus threats enumerated, rated and owned). Report it weighted by criticality as well as flat, because the unmodeled remainder is where the interesting question lives. ### 2. Follow-through - findings closed and mean time to close Coverage without closure means the practice produces documents. Track, per severity band, how many findings from models are closed with a linked change, and the mean elapsed time from a finding being written to that change shipping. This is the number that proves modeling output reaches the codebase. Two cautions: closure by silent risk acceptance is not closure, so sample how findings were closed, and a closure rate collected only over findings a team chose to record is self-selecting. ### 3. Latency - time-to-model Elapsed time from the trigger (a design review requested, a new service scaffolded) to a finished model with owned findings. This is the number that predicts whether teams will route around you. It only becomes useful when you **decompose** it into requested -> scheduled, scheduled -> session held, session -> findings written and owned. A fintech that measured eleven elapsed days of time-to-model found that nine of them were queue: teams waited a week and a half for a slot with a facilitator, and the analysis itself took under two days. The headline number does not say *modeling is slow*; it says *facilitation capacity is the constraint*. The remedies differ completely - more trained facilitators or a self-serve path for low-risk changes, versus a lighter method if the analysis itself were the slow part. ### 4. Escapes - issues found after the design passed An **escaped issue** is a design-level security problem discovered in production, in a later test, or in an incident, on a system that had been modeled. This is the only genuinely *lagging outcome* measure of the four, and it is slow, sparse and noisy: a quarter with zero escapes usually means nobody looked, not that the practice is perfect. Track it as a rate over modeled systems and, more usefully, review each escape individually. ### Reading them together | Metric | Type | Fails when it says | | --- | --- | --- | | Coverage | Leading | Practice is not reaching real designs | | Findings closed / mean time to close | Leading | Output never becomes change | | Time-to-model | Leading | Teams will route around the practice | | Escaped issues | Lagging | The method or its scope missed something real | Coverage plus latency describe whether the practice runs; closure plus escapes describe whether it works. Reporting only the first pair is the common trap - it makes an activity look like an outcome. ### Instrumentation None of this needs a measurement programme of its own. Timestamps on the model record (requested, session held, signed off), a link from every finding to the ticket that fixes it, and a small quarterly read of a sample of models for quality give you all four. Add one qualitative input: read five or six models a quarter and judge whether the threats found are the ones a competent attacker would pick, because no counter can tell you that. ### What not to put on the dashboard Raw threat counts, sessions held, hours spent, models written, and any per-team league table built from those. They measure activity, they are trivially inflated, and threat granularity is an analyst choice so the units are not comparable between teams. Use them at most as internal diagnostics - a session that produced forty threats usually means the scope was too big, not that the system is forty times worse. A defensible executive report is three numbers and a sentence: coverage of tier-1 services, closure rate and mean time to close for high-severity findings, and escapes this period, followed by what the escapes taught you about the practice.

  • If you could report only one of these numbers upward, which would you keep?
    Findings closed with mean time to close, per severity band. Coverage can be high while nothing changes, time-to-model measures the process rather than the system, and escaped issues are too sparse to steer by quarter to quarter. Closure is the number that says modeling output reached the running system, and it is the hardest one to inflate without actually shipping a change.
  • How do you separate waiting time from analysis time inside time-to-model?
    Stamp the model record at three points: when the trigger fired or the team requested it, when the session was actually held, and when findings were written and had owners. Requested-to-held is queue, held-to-owned is work. If queue dominates you fix facilitation capacity or offer a self-serve path; if work dominates you lighten the method or shrink the scope of a session.
  • Why is a quarter with zero escaped issues not good news on its own?
    Escaped issues are only counted when someone finds them, so zero usually means no one looked - no penetration testing, no post-incident review reaching design causes, or no habit of tracing an incident back to the model. Treat a zero as a prompt to check your detection paths before you treat it as evidence the practice works.
  • What does a metric dashboard miss entirely?
    Whether the threats being found are the ones that matter. Counters cannot judge analytical quality, so pair them with a qualitative sample: read a handful of models each quarter and ask whether a competent attacker with the assumed position would have picked the same targets. That read is what catches a practice that is busy, fast, well covered and shallow.

It is the difference between measuring how many fire drills you ran and measuring how fast the building actually empties. Both are worth knowing, but only one is the outcome.

saying these in an interview costs you the question

  • Reports number of threats found as the headline metric
  • Quotes a coverage percentage with no stated denominator
  • Counts models produced but never checks whether findings closed
  • Treats coverage and session count as outcome measures
  • Reads a slow time-to-model as proof the method is too heavy
  • Claims zero escaped issues proves the practice is working

context

open as a page

Your dashboard reports 94% of services threat-modeled: how do you check that number is honest?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Interrogate the three hidden choices behind the percentage: which population forms the denominator, what test earns a service a tick in the numerator, and how recent the model must be to count. Then sample and reconcile against an independent inventory.

open as a page

How do you run an escaped-issue review when an incident hits a system you threat-modeled?

level: seniorimportance: should knowfreq 51%

basics

~20 s

Walk a fixed chain and stop at the first no: was it in scope, on the diagram, enumerated, rated correctly, accepted deliberately, controlled, and was the model current. Each stop names a different part of the practice to fix.

open as a page

Why does reporting raw threat count as a threat modeling metric backfire?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Threat granularity is the analyst's choice, so the count has no fixed unit and is not comparable across teams. Once it becomes a target, the cheapest way to raise it is splitting one threat into several.

open as a page