skip to content

Qualitative Judgment

Rating a threat that has no exploit yet: likelihood as attacker effort, impact as worst credible loss, matrices and bug bars. Interviewers watch whether two raters can ever agree.

on this pageshow

explore

questions

10

What is a bug bar in an SDL, and what does it decide about a security finding?

level: middleimportance: must knowfreq 60%

answer

  1. written down before the argument
  2. bands defined by example, not adjective
  3. every band carries a clock
  4. some bands stop the release
  5. lookup, not negotiation

basics

~20 s

A bug bar is a written table, agreed before any finding exists, that defines each severity band by the concrete kinds of flaw that belong in it, the deadline to fix each band, and which bands block shipping.

solid answer

~50 s

A bug bar is the severity rulebook a product team writes **in advance**, so that rating a finding is a lookup rather than a negotiation. Each band — Critical, High, Moderate, Low — carries three things: anchored example rows describing the class of flaw that lands there ("any read of another tenant's data", "unauthenticated takeover of the device from the local network"), a fix clock for that band, and whether the band blocks a release or is merely tracked. The anchoring is the whole point: bands described as adjectives ("severe impact") are re-argued every time, while bands described as concrete flaw classes let two different reviewers land on the same answer. Some rows deliberately override any likelihood estimate — a B2B HR product may declare *any* cross-tenant read Critical by definition, because the company has decided that impact class alone determines the response.

go deeper

for a junior

Be ready to say what the artifact is: a table of severity bands written in advance, each with examples of what belongs there and how fast it must be fixed. Knowing that it exists before the finding, not after, already puts you ahead.

for a middle

An interviewer expects you to explain the three columns — anchored class definitions, fix clock, release consequence — and why definitions by example beat definitions by adjective. Be able to sketch two or three plausible rows for a product you have worked on.

for a senior

Show you have used one under pressure: how you map a real finding to a row, what you do when no row fits, and why the fix clock and the ship gate are separate decisions. Mention that some rows deliberately override likelihood.

for a principal

Own the tradeoff in writing the bar itself: categorical rows buy repeatability at the cost of occasionally over-fixing, a gate that catches everything catches nothing, and a bar shared across teams is worth more than a locally optimal one. Say who signs it and how often it is revisited.

## The artifact A bug bar is a short written document — almost always a table — that a product team agrees to **before** it has any findings to argue about. Each row is a severity band, and the row states three separate things: 1. **What belongs in this band** — described as concrete, recognisable classes of flaw, not adjectives. 2. **The fix clock** — how long the team has to remediate something rated into this band. 3. **The release consequence** — whether a finding in this band blocks the build from shipping, blocks only a major release, or is simply tracked. The practice comes from SDL-style secure development lifecycles, where it exists for one reason: to turn severity from a matter of taste into a lookup. ## Anchored by class, not by adjective An illustrative shape (the exact rows and clocks are always organisation-specific): | Band | Example rows | Fix clock | Release effect | |---|---|---|---| | Critical | Unauthenticated remote code execution; any read of another tenant's data; authentication bypass | 7 days | Blocks the build | | High | Authenticated low-privilege user alters financial records; unauthenticated takeover of the device from the local network | 30 days | Blocks a major release | | Moderate | Sensitive value written to a log readable by all operators; missing rate limit on a credential-checking endpoint | 90 days | Tracked, not gating | | Low | Verbose error text revealing a framework version | Best effort | Tracked | The left-hand descriptions are what make the bar work. "Severe business impact" is not a definition — two people will place the same finding in two different bands and both will be defensible. "Any read of another tenant's data" is a definition: either the flaw lets tenant A read tenant B's records or it does not. ## Why it must be written first Severity arguments are structurally unfair. The person who would have to do the remediation work has an incentive to argue the band down; the person who found it has an incentive to argue it up; and both arguments get louder as a ship date approaches. Writing the definitions before any specific finding exists removes the incentive, because nobody yet knows whose feature the rule will land on. That is the sentence to say in an interview: **the definitions are agreed before the argument starts, precisely so the argument has somewhere to end.** ## What the bar is not - It is **not a scoring formula**. A bar does not compute a number; it classifies. Where a team also produces numeric scores, the bar is what turns a score into an obligation, and the bar's own class rows normally win where the two disagree. - It is **not a per-finding rating exercise**. Estimating one threat's likelihood and worst credible impact is a separate activity; the bar is where that estimate stops being negotiable. - It is **not the incident severity scale**. A bug bar governs a defect found before or after release and the clock to fix it; classifying a live incident and paging people is a different scale with different owners. ## Bands that ignore likelihood on purpose A common and defensible design is the categorical row: a class of flaw is assigned a band regardless of how likely the team thinks exploitation is. A multi-tenant HR SaaS vendor holding employee personal data may write "any cross-tenant read is Critical" and mean it literally — a curious or malicious tenant administrator poking at an identifier is cheap, the harm is regulated personal data, and the company would rather over-fix that class than re-litigate likelihood in every design review. The cost of the choice is real: some genuinely hard-to-reach flaws get an expensive clock. The benefit is that the class can never be talked down by whoever is under the most schedule pressure. ## Worked example: shipping firmware A consumer Wi-Fi router vendor finds, during design review, that the configuration-restore endpoint accepts an uploaded config blob from anyone on the LAN and applies it without authenticating the uploader — which includes overwriting the administrator credential. If the bar's Critical row reads "unauthenticated takeover of the device from the local network", the mapping takes seconds and the firmware build does not ship. Nobody has to invent a policy at 2am, and nobody has to win an argument about how many customers really have hostile guests on their Wi-Fi. ## How bug bars fail - Bands described as adjectives, so every finding is re-argued. - Severity assigned but no fix clock, so a High sits open for a year. - Everything gates the release, so the gate is quietly ignored and stops meaning anything. - No row matches the finding, and the gap gets resolved in favour of whoever is loudest instead of being fed back into the bar. - Each team keeps its own bar, so severity depends on which team owns the code. - The bar is never revisited, so it encodes the risk appetite of a company that no longer exists.

  • Why anchor each band to concrete flaw classes instead of words like 'severe' or 'moderate'?
    Because adjectives are not testable. Two reviewers reading "severe business impact" will place the same finding differently and neither can be shown wrong, so the band becomes whatever the more senior or more stubborn person says. "Any read of another tenant's data" is checkable against the finding itself, which is what makes ratings repeatable across reviewers, teams and quarters.
  • What do you do when a finding matches no row in the bar?
    Rate it by the nearest row's consequence, say explicitly that you are doing so, and open a change to the bar. The gap is information: it usually means the product has grown a new asset class or a new attacker position the bar never contemplated. Fixing the bar afterwards is legitimate; quietly picking a convenient band because no row fits is not.
  • Does a bug bar's severity band and its fix deadline have to be the same decision?
    They are separate columns and should be. Severity says how bad the class of flaw is; the clock says what the organisation has committed to doing about it; the release column says whether it also stops a build. A Moderate can carry a 90-day clock and gate nothing, while a Critical carries a 7-day clock and blocks shipping. Conflating them produces either a gate that blocks everything or clocks nobody honours.

It is the building code, not the inspector's opinion: the rules about railing height are published before your house exists, so the inspection is a check rather than a debate.

saying these in an interview costs you the question

  • Describes a bug bar as a formula that computes a severity score
  • Thinks the bar rates one finding rather than defining the bands in advance
  • Leaves bands as adjectives like severe or moderate with no anchored examples
  • Assigns severity bands but no fix deadline or release consequence
  • Lets the team that owns the code choose the band case by case
  • Confuses it with the incident severity scale used for paging on-call

context

open as a page

In a threat model, how do you rate the likelihood that a warehouse picker's handheld can post fraudulent stock write-offs?

level: middleimportance: must knowfreq 70%

basics

~20 s

Rate likelihood from what the abuse requires, not from how clever it is: who already holds the access, how much effort and skill it takes, and whether anyone would notice. Every picker already has the scanner, so likelihood is high.

open as a page

Why is multiplying likelihood and impact band numbers on a 5x5 risk matrix unsound?

level: middleimportance: must knowfreq 58%

basics

~20 s

Risk-matrix band numbers are ordinal ranks, not measured quantities: they say one band is worse than the next, not how much worse. Multiplying ranks yields a score with no consistent meaning, so the ordering it produces is arbitrary.

open as a page

How do you assign a bug bar severity to a modeled threat that has no proof of concept?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Band the threat on the worst credible consequence if it is real, using the bar's flaw-class rows. Absence of a working exploit is not evidence of low severity; genuine doubt about whether the flaw exists becomes a time-boxed verification task, not a downgrade.

open as a page

How do you set a modeled threat's impact rating without inflating every finding to critical?

level: seniorimportance: should knowfreq 56%

basics

~20 s

Rate the worst credible loss this one threat delivers, not the worst story you can tell. Name what is reached, cap it at what the attacker's position gives them, then translate that into business and regulatory harm.

open as a page

How do you anchor risk-matrix impact bands so two raters land on the same cell?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Replace adjective labels like Major with concrete, observable outcomes in the product's own vocabulary — for a photo-sharing service, Impact 4 means any cross-account read of another user's private media. Write the anchors before the rating argument, then calibrate them on past threats.

open as a page

With one sprint of capacity, do you fix the likely-but-bounded threat or the remote-but-catastrophic one?

level: principalimportance: should knowfreq 40%

basics

~20 s

Reject the framing that one rating decides both. Ship the cheap fix for the likely bounded threat this sprint, and treat the catastrophic one as an architecture decision with a named owner, a date and a written acceptance.

open as a page

How does OWASP Risk Rating break likelihood and impact into factors you can grade?

level: middleimportance: nice to knowfreq 38%

basics

~20 s

OWASP Risk Rating splits likelihood into four threat-agent factors and four vulnerability factors, and impact into four technical and four business factors. Each is scored 0 to 9 and averaged within its group, then read as low, medium or high.

open as a page

As the owner of a bug bar, how do you keep fix deadlines and the ship gate credible under launch pressure?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Agree the bands, clocks and gate before any finding exists, keep the gating set small enough that a block is believable, route acceptance to an accountable business owner rather than the security engineer, and change the bar only in scheduled reviews, never mid-argument.

open as a page

A mobile operator's risk matrix rates seventy percent of modeled threats amber — what do you change?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

A matrix that rates most threats alike has stopped discriminating, so it cannot sequence work. Diagnose why first — vague bands, raters avoiding the extremes, or a colour region drawn across too many cells — then re-cut the bands against the real portfolio.

open as a page