Why rate all threats from one design-time model in a single sitting rather than as each one surfaces?
answer
- the output is an order, not a number
- ordinal scales have no fixed unit
- anchoring, recency, fatigue
- enumerate first, rate second, one pass
- two or three worked anchor threats
basics
~20 sRating the whole set in one pass keeps every threat on the same scale. Rated piecemeal over weeks, the bar drifts: what the group called High early becomes Medium later, so the resulting order is not comparable.
solid answer
~50 sThe output of a threat model is an **order**, not a set of absolute numbers, and an order only means something if every row was judged against the same bar. Home-grown High/Medium/Low scales are ordinal and drift with mood, recency and whoever is in the room: rate twenty threats as they surface across three sessions weeks apart and the early ones cluster High while later, similar threats land Medium — or the reverse, once the group is tired of saying High. So I enumerate first and rate second, in one pass, with the same people. I pick two or three anchor threats up front ("this is what a High looks like here, this is a Low") and score everything relative to them, rather than accepting fifteen Highs. If ratings really must be merged from separate sittings, I re-rate a sample of the older rows in the new session and shift the old batch to match.
go deeper
Know that a threat model ends with a rated list, and that the ratings exist to put the threats in order. Be able to say that scores from different sessions may not mean the same thing.
Be ready to explain why an ordinal band drifts — anchoring, recency, fatigue — and to describe the mechanics of a batched rating pass: enumerate first, agree anchors, score in one sitting, force a spread.
Show you have run the session. Talk about who is in the room, how you break a band where everything landed High, and how you re-calibrate an inherited batch before merging new rows into it.
Own the tradeoff between comparability and cadence: strict batching conflicts with modeling continuously as a design evolves. Be able to defend a scheme that keeps ratings comparable across many teams without turning rating into a ceremony nobody schedules.
## What "batching" means here A design-time threat model produces a list — often fifteen to forty threats for one service. Each needs a rating so the team can decide what to fix first. Batching means: finish enumerating the threats, *then* rate the whole list in one sitting, against one bar, with one group of raters. The alternative — scoring each threat the moment it is raised, or letting the list accumulate over three workshops weeks apart — is what most teams do by default, and it is where the ranking quietly becomes meaningless. ## Why the scale drifts Risk ratings on a threat model are usually **ordinal**: High, Medium, Low, or a 1–5 scale on a few dimensions. Ordinal scales have no fixed unit. "High" means only "higher than the things I have called Medium", and that reference set changes as the session goes on. Three effects do the damage: - **Anchoring.** The first threat rated sets the mental yardstick. If the session opens with a catastrophic one, everything afterwards looks small; if it opens with a trivial one, ordinary threats inflate. - **Recency and context.** A threat rated the week after a noisy outage in the same component gets a harsher likelihood than the identical threat rated a month later. - **Fatigue and grade inflation.** A group that has said "High" eight times starts reaching for Medium to keep the list looking actionable — or, in the opposite failure, marks everything High so nothing has to be argued about. None of these are visible in the artifact. You end up with a tidy table of bands where a High from session one and a High from session three were produced by different bars, and the sort order that comes out of it is an artifact of scheduling, not of risk. ## What comparability actually requires Three things make the ratings comparable: 1. **One pass, one group.** Rate the whole enumerated set in a single sitting with the same raters. If the model itself was built over several workshops, that is fine — hold the rating until enumeration is done. 2. **Explicit anchors.** Before scoring, agree on two or three worked examples from *this* system: one that everyone accepts is High, one clearly Low. Every subsequent row is judged against them. Anchors turn an abstract band into a comparison, which humans do far better than absolute estimation. 3. **A forced spread.** If more than roughly a third of the list lands in the top band, the bar is wrong, not the system. Re-rate by comparison — take the two Highs and ask which you would fix first if you could only do one. Pairwise comparison is often faster and more reliable than assigning numbers, and it produces the ranking directly. ## Published scales help, but only partly A scoring framework reduces drift but does not remove it. CVSS base metrics are deliberately written to describe a flaw's *intrinsic* characteristics so two teams score the same flaw the same way — but a base score is environment-independent by design; it does not know that your instance is internet-facing or that the affected data is a regulated record. That is what CVSS's environmental metric group exists to correct, and skipping it is a common mistake. DREAD's five dimensions — damage, reproducibility, exploitability, affected users, discoverability — are more subjective still: two engineers routinely differ by two points on discoverability alone, which is precisely why DREAD fell out of favour as a precision instrument. Whatever the scale, the discipline is the same: same scale, same sitting, same anchors, and treat the output as a rank order rather than a measurement. ## Merging batches when you have to Sometimes you inherit ratings from an earlier session and must slot new threats in beside them. Do not append. Re-rate a sample of the old rows — five or six spanning the bands — inside the new sitting. If they land where they used to, the bars agree and you can merge. If they systematically land a band lower, the old batch was inflated, and you shift it before merging. This costs twenty minutes and is the difference between a queue people trust and one that quietly puts the wrong work first. ## The point of the exercise Precision is not the goal. Nobody needs to know whether a threat is a 7.2 or a 7.6. What the team needs is a defensible statement that *this* threat is worse than *that* one, made on evidence they can restate weeks later. Batched rating with anchors delivers that; drip-fed rating does not.
- Two thirds of the list came out High. What do you do?Treat it as a broken bar, not a dangerous system. I take the top-band rows in pairs and ask which one I would fix if I could only fix one; that pairwise comparison re-splits the band and produces the ranking directly. If a whole cluster genuinely is severe, it usually means several rows share one root cause — which is a signal to look for the shared fix rather than to keep twelve separate Highs.
- Does using a published scoring framework remove the need to batch?It reduces drift but does not remove it. A framework fixes the dimensions and their wording, so two raters argue over less; it still leaves judgment calls — how discoverable, how many users affected — that move with mood and context. And a base score is written to be environment-independent, so it cannot rank *your* threats without local adjustment. Same-sitting rating with local anchors is still what makes the set comparable.
- Who should be in the room when the ratings are set?The people who can dispute both halves of a rating: someone who knows what the data and the transaction are actually worth, and someone who knows how hard the path really is to reach in this deployment. A rating set by security alone drifts toward theoretical severity; one set by the delivery team alone drifts toward whatever is cheap to fix. Keep the group stable across the sitting.
Grading twenty exam papers in one sitting against two sample answers gives a fair ranking; grading them one a week for five months does not, even though every paper got a mark.
saying these in an interview costs you the question
- Treats High/Medium/Low as an objective measurement
- Rates each threat immediately as it is raised, never revisiting
- Appends new ratings to an old batch without re-calibrating
- Accepts a list where nearly everything is High
- Argues about decimal precision instead of relative order
- Assumes a base score already reflects their environment