Leadership tracks '4,000 open vulnerability findings' as the security metric - how do you reframe it?
answer
- It measures coverage, not risk
- One upgrade can be hundreds of rows
- Improving coverage makes it look worse
- Count fix actions, not finding rows
- Pre-announce the rise to leadership
basics
~10 sA raw finding count measures scanner coverage and estate size, not risk. It double-counts one shared component across many services, rises when coverage improves, and treats every finding as equal. Report evidence-driven measures instead.
solid answer
~50 sThe number is real but it is not a risk metric - it is a coverage metric wearing a risk metric's clothes. It fails three ways: it **double-counts**, because one shared base layer or transitive library across hundreds of services becomes hundreds of rows for a single upgrade; it **moves the wrong way when you improve**, because onboarding more repositories or resolving deeper graphs makes it jump, which rewards not scanning; and it is **flat**, treating a flaw with observed exploitation on an internet-facing service and the same flaw in an offline batch cluster as one row each. I would report instead: open items with evidence of real-world exploitation and their age, items past their due date, the count of **distinct fix actions**, and coverage as its own separate number. Then warn leadership the raw total will rise as coverage improves, so success does not look like regression.
go deeper
Know that the same vulnerable library appearing in many services produces many findings but usually only one upgrade, so a total count overstates the work.
Explain the three ways a raw count misleads - duplication across services, growth when coverage improves, and treating every finding as equal regardless of exposure.
Propose concrete replacements: confirmed-exploitation entries and their age, overdue items against the likelihood cut, distinct fix actions, and coverage reported separately.
Own the conversation with leadership: pre-announce that the total will rise, keep the old number visible but relabelled as inventory, and refuse a reduction target on a metric whose natural value is zero.
## Why the number feels like a metric A finding count is available, it is precise, it goes down when you work, and it fits in a slide. That is why it survives. The problem is that what it actually measures is **how much of the estate you point a scanner at, multiplied by how deep that scanner looks** - and neither of those is risk. ## Three structural failures **It double-counts.** One vulnerable library pulled transitively into four hundred services is four hundred rows and one upgrade. Meanwhile a genuinely isolated flaw in a single high-value service is one row. Ranked by count, the trivial fan-out dominates the report, and the work implied by 4,000 rows may be a few dozen actual changes. **It moves the wrong way when you improve.** Onboard two hundred more repositories: the number leaps. Turn on deeper transitive resolution: it leaps again. Start scanning the estate's older, unloved services - exactly the ones most likely to be dangerous - and it leaps hardest. A leadership metric that punishes coverage creates a real, rational incentive to scan less, to suppress in bulk, or to quietly exclude the messy corners. The moment a number can be improved by looking away, it will be. **It is flat.** The count says nothing about whether anyone is exploiting these flaws, whether the vulnerable code is reachable, or where the affected service sits. A network-exploitable flaw fronting the public internet and the same flaw inside an air-gapped analytics cluster where no outsider can reach any service both contribute exactly one. In that second environment the printed severity is genuinely the least informative field on the row, and the count inherits that blindness. ## What to report instead A small set, each of which is hard to game and each of which implies an action: | Metric | Why it survives scrutiny | |---|---| | Open findings with evidence of observed exploitation, and the age of the oldest | Should be near zero; any entry is a specific conversation, not a trend line | | Items above the exploitation-likelihood cut that are past their due date | Measures whether the process is keeping its own promises | | Distinct fix actions outstanding (unique component-and-version upgrades) | Reflects the actual work, collapsing fan-out | | Age distribution of the open set, not the total | Ageing is the risk signal; totals are noise | | Coverage - services scanned, ecosystems resolved - reported separately | Improving coverage becomes visible progress rather than apparent regression | The design principle behind all of them: **report the count as inventory and commit to a trend only on the small, evidence-driven set.** You can defend "zero open items with confirmed real-world exploitation, oldest overdue item is nine days" to a board, a customer or an auditor. You cannot defend a total that you know is dominated by duplicates. ## The organisational move, not just the analytical one This is where the question stops being technical. Three things have to happen before the new metrics land: 1. **Pre-announce the rise.** Say explicitly, before it happens, that the raw total will go up as coverage extends, and that this is the programme working. If leadership discovers it themselves in a quarterly review, you will spend the next quarter defending your credibility instead of shipping fixes. 2. **Keep the old number visible, relabelled.** Deleting it looks like hiding. Keep it on the page under the heading it deserves - inventory, or scanner output - and put the small set above it. People trust a metric that survived being reframed more than one that vanished. 3. **Agree what a nonzero looks like.** A small metric only works if everyone accepts that its natural value is zero or near it, and that a single entry warrants a named owner and a date rather than a percentage-improvement target. A small set with a 'reduce by 30 percent' goal attached to it has been turned back into the thing you replaced. ## The failure mode of over-correcting Do not swing to a single composite risk score either. A blended number is just as opaque as a count and much harder to argue with, because nobody can see which input moved. Two or three plain numbers that each name a real thing beat one clever index every time.
- Is a metric of 'open items with confirmed exploitation' not too small to show progress?That is the point - its healthy value is zero, and a single entry is a named conversation rather than a trend line. Pair it with a volume-bearing measure that cannot be gamed the same way, such as the age distribution of the whole open set or the oldest overdue item. You get one metric that means something and one that shows movement, instead of one number pretending to do both.
- You onboard two hundred more repositories and the total jumps forty percent. What do you present?Coverage and risk as two separate lines, with the jump attributed explicitly to the coverage change and dated to it. Normalise where you can - findings per scanned service - so the like-for-like comparison is visible. And point at the small evidence-driven set, which should not have moved much, as the answer to 'did we get more dangerous this month'. The honest framing is that you can now see risk you already had.
- Why not roll everything into one composite risk score for the board?Because it is as opaque as the count and harder to challenge. When a composite moves, nobody can see which input moved it, so the discussion becomes about the formula rather than about the work. Two or three plain numbers that each name a concrete thing - overdue items, confirmed-exploitation entries, oldest open item - survive scrutiny better and drive clearer decisions.
Counting open findings is like measuring a hospital by the number of symptoms recorded in its notes. Hire more doctors who write better notes and the number rises, while the patients are healthier than before.
saying these in an interview costs you the question
- Defends the raw count as a valid risk measure
- Proposes a percentage-reduction target on the total
- Ignores that improving coverage inflates the number
- Replaces the count with an opaque composite score
- Deletes the old metric instead of relabelling it