skip to content

Risk Matrix Pitfalls

A 3x3 or 5x5 matrix squeezes a huge range into a few bands, invites arithmetic on ordinal labels and colours cells by taste. Interviewers probe why two competent raters disagree.

on this pageshow

questions

3

Why is multiplying likelihood and impact band numbers on a 5x5 risk matrix unsound?

level: middleimportance: must knowfreq 58%

answer

  1. check what the numbers actually are
  2. rank order, not measured distance
  3. bands are usually orders of magnitude apart
  4. relabel the bands and re-sort
  5. multiplication needs a ratio scale

basics

~20 s

Risk-matrix band numbers are ordinal ranks, not measured quantities: they say one band is worse than the next, not how much worse. Multiplying ranks yields a score with no consistent meaning, so the ordering it produces is arbitrary.

solid answer

~50 s

The 1-to-5 labels on each axis are ordinal: they encode order only, not distance. Multiplication is a ratio-scale operation, so applying it to ranks manufactures precision that the inputs never contained. The give-away is that the ranking changes if you relabel the same bands 10-50, or 1-2-4-8-16, without changing a single judgment — the order is an artifact of the labels. It also hides real differences: on a payments team's sheet, `Likely (4) x Moderate (3) = 12` and `Possible (3) x Major (4) = 12` sort as equals, even when the second threat loses two orders of magnitude more money. The fix is to stop computing. Assign a severity to each of the 25 cells deliberately and in advance, so the team can say things like "any Catastrophic impact is at least High regardless of likelihood" — a judgment no product can express.

go deeper

for a junior

Be ready to say what the two axes of a risk matrix mean and that the 1-to-5 labels are rankings, not measurements. Knowing that a computed score is not automatically more objective than a judgment is enough at this level.

for a middle

Explain the mechanics: ordinal versus ratio scales, which operations each supports, and why multiplying ranks invents distances. Be able to run the relabelling test out loud to show the ranking depends on the labels chosen.

for a senior

Show what you do about it on a real team: replace the formula with an agreed cell-to-severity lookup, keep the raw cell coordinates attached to every rating, and spot the moment a discussion turns into an argument about a score rather than about a threat.

for a principal

Own the choice of instrument. Decide where the organisation stays qualitative and where a threat is important enough to warrant real quantities, and defend that boundary against pressure for one uniform number that rolls up neatly into a slide.

## What the numbers on a matrix actually are A qualitative risk matrix has two axes: a likelihood scale (something like Rare, Unlikely, Possible, Likely, Almost Certain) and an impact scale (Negligible through Catastrophic). Teams number the bands 1 to 5 because it is convenient to write "4" in a spreadsheet cell. That number is a **label for a rank position**. The scale is *ordinal*: it tells you that Likely sits above Possible, and nothing at all about the size of the gap between them. ## Which operations a scale supports Measurement theory sorts scales by the operations they support: | Scale | Meaningful operations | Example | |---|---|---| | Nominal | equality | threat category | | Ordinal | order, median, mode | severity bands, rank positions | | Interval | differences, addition | calendar dates | | Ratio | ratios, multiplication | money, probability, elapsed time | Likelihood and impact bands are ordinal. Multiplying them applies a ratio-scale operation to ordinal data. The result is a number, and numbers look authoritative, but no information has been added — and the operation has quietly invented distances that were never measured. ## The relabelling test The cleanest way to show the defect in an interview: keep every rating identical and relabel the bands. With 1-5 on both axes, `4 x 3 = 12` ties `3 x 4 = 12`. Relabel the impact bands 1, 2, 4, 8, 16 to reflect that each step is roughly an order of magnitude of loss — a very common reality — and the two threats are no longer tied at all; one is now several times the other. **If a ranking flips when you rename the bands without re-judging anything, the ranking was a property of the labels, not of the risk.** ## What it costs in practice On a payments team, band boundaries are usually roughly logarithmic on both axes: an impact of 3 might be tens of thousands of pounds of fraud loss and a 4 millions; a likelihood of 3 might be once a year and a 4 once a month. Ranking modelled fraud threats by the product then does two damaging things at once. It **ties threats that are nothing alike**, because many different cells share a product. And it **flattens the extremes** the team most needs to see, because a rare-but-ruinous threat scores the same as a frequent nuisance. It also invites a second failure: once a score exists, people argue about the score rather than about the threat. Two teams grading the same account-takeover threat 4x3 and 3x4 will report "we both said 12" and never discover that they disagreed profoundly about how often it happens and how much it costs. ## What to do instead 1. **Make the cell-to-severity map an explicit decision, not a calculation.** Fill in all 25 cells once, by agreement, and write it down. This lets the map be deliberately asymmetric — for example, every cell in the top impact row is at least High no matter how unlikely, which is exactly the judgment a payments or safety organisation wants and which no product can express. 2. **If you genuinely want arithmetic, get quantities.** A probability per year and a loss in money are both ratio scales, so multiplying them to get an expected annual loss is a legitimate operation. The estimates will be wide ranges rather than points, and saying so honestly is better than a false-precision integer. 3. **Use the matrix for what it is good at**: sorting a modelled threat list into a small number of action tiers that a room can agree on and communicate, then ordering within a tier by criteria you declare in advance. ## What an interviewer is listening for Junior answers say "multiplying is fine, it is what everyone does". A strong answer names the scale type, states which operations it supports, gives the relabelling test as proof, and — crucially — offers a workable replacement rather than just objecting. Interviewers also watch for the opposite over-correction: refusing all qualitative rating because it is not measurement. Qualitative bands are useful and defensible; it is the arithmetic performed on them that is not.

  • If the product is unsound, how should a matrix cell become a severity?
    By a lookup table the team fills in once and agrees on: every cell of the grid is assigned a severity deliberately, before any live threat is rated. That lets the map be asymmetric where judgment demands it — for instance, the whole top impact row rated at least High regardless of likelihood — and it removes the arithmetic argument entirely, because there is nothing to compute.
  • When is multiplying likelihood by impact legitimate?
    When both inputs are real quantities rather than ranks. An annual probability multiplied by a loss expressed in money is an expected annual loss, and both are ratio scales, so the operation is valid. The estimates are still uncertain, so express them as ranges rather than single numbers — but uncertainty in an input is a different problem from performing an invalid operation on it.
  • Two payments teams both rate an account-takeover threat 12, one as 4x3 and one as 3x4. What does the equal score hide?
    That they disagree about everything that matters. One team thinks the attack is frequent and moderately costly, the other rare and severe. Those imply different controls: rate limiting and monitoring versus hard transaction limits and recovery planning. The product erases the disagreement and presents a false consensus, which is why the raw cell coordinates should always travel with any severity, not just the derived label.

Race finishers are numbered 1, 2 and 3, but first place times three is not third place. The numbers order the runners; they say nothing about the gaps between them.

saying these in an interview costs you the question

  • Says the product is fine as long as both raters use the same scale
  • Treats two threats scoring 12 as equally urgent
  • Assumes band 4 is twice as bad as band 2
  • Adds the bands instead of multiplying and calls that safer
  • Calls the computed score objective because it is a number

context

open as a page

How do you anchor risk-matrix impact bands so two raters land on the same cell?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Replace adjective labels like Major with concrete, observable outcomes in the product's own vocabulary — for a photo-sharing service, Impact 4 means any cross-account read of another user's private media. Write the anchors before the rating argument, then calibrate them on past threats.

open as a page

A mobile operator's risk matrix rates seventy percent of modeled threats amber — what do you change?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

A matrix that rates most threats alike has stopped discriminating, so it cannot sequence work. Diagnose why first — vague bands, raters avoiding the extremes, or a colour region drawn across too many cells — then re-cut the bands against the real portfolio.

open as a page