skip to content

Which facts make a low-severity defect a higher priority than a crash?

level: middleimportance: must knowfreq 71%

answer

  1. Exposure beats drama
  2. Three inputs, not one
  3. How many, how often, workaround grade
  4. Ask who performs the workaround
  5. Measured share, not adjectives

basics

~20 s

Reach, frequency and the absence of a workaround. A cosmetic defect on a path every user crosses, hit on every attempt, with nothing the user can do about it, outranks a crash reachable only through a configuration almost nobody runs.

solid answer

~50 s

Priority is driven by exposure rather than by how dramatic the failure looks. Three inputs dominate: **blast radius** — how many users, accounts or requests cross the broken path; **frequency** — how often the failure fires when that path is taken; and **workaround** — whether the user or an operator can get the outcome another way, and how costly that detour is. Add deadlines: a defect on a flow named in a contract or a launch is urgent regardless of severity. So a trivial-severity wrong label on the first screen of a signup flow, seen by every prospective customer on every visit, with no way around it, can outrank a critical-severity crash sitting behind an unreleased flag that fires perhaps once in several thousand runs and that an operator can retry past. Severity still records the technical damage; it just is not the whole ranking.

code

pseudocode · 11 lines
pseudocode
severity  = impact_of_observed_failure(defect)      # blocker .. trivial

priority  = rank_against_queue(
    severity      = severity,
    blast_radius  = share_of_traffic_on_path(defect),
    frequency     = failures_per_attempt(defect),
    workaround    = workaround_grade(defect),        # user / support / operator / none
    deadline      = commitment_touched(defect)
)

assert priority is not derived_only_from(severity)

go deeper

for a junior

Be ready with one concrete example of each mixed corner — a cosmetic flaw everyone sees, and a crash almost nobody reaches. Naming reach, frequency and workaround as the reasons is what turns the example into an answer.

for a middle

Explain the mechanics of each input: what reach is measured against, why frequency per attempt matters separately, and why a workaround's grade — user, support or operator — changes how much urgency it removes.

for a senior

Demonstrate judgement with numbers you actually have, including saying when you do not have them and how you would get them cheaply. Show that you never launder an urgency argument through the severity field.

for a principal

Own the ranking inputs as policy: which exposure measures the organisation trusts, how operator workarounds get priced as recurring cost rather than treated as free, and how the queue avoids being reordered by whoever argues loudest.

## The grid, and why the mixed corners matter Because severity and priority are independent axes, four combinations exist, and interviewers ask about the two diagonal ones because they are where the thinking shows. **High severity, high priority.** The uncontroversial case. A core flow is broken for everyone, no workaround, fix it now. **Low severity, low priority.** Also uncontroversial. A misaligned column on an internal report screen used twice a quarter. **Low severity, high priority.** Nothing is technically broken in a serious way, but the defect sits where everyone can see it and there is no way around it. Classic shapes: wrong wording on a legal or pricing statement; a misspelled brand or product name on a public entry page; a label that tells the user the wrong thing about what is about to happen; a wrong figure shown in a summary even though the stored data is correct. **High severity, low priority.** The behaviour is genuinely bad — a crash, lost in-flight work — but almost nobody can get there. Shapes: a path behind a flag that has never been enabled outside a test environment; a failure that requires an input combination no real client produces; a deprecated flow that will be removed in the next release anyway. ## The three inputs that move priority **Blast radius.** How much of the population crosses the defective path. Not the whole install base — the population that actually reaches the code. Express it as a measured share where you can: "this path is exercised by 63% of submissions" carries an argument; "lots of users" carries none. **Frequency.** Given that a user takes the path, how often does the failure fire? A defect that fires every single time is a different proposition from one that fires occasionally, even at identical severity, because the intermittent one may be survivable by retrying while the deterministic one is a wall. **Workaround.** Is there another route to the same outcome, and what does it cost? Three grades are worth distinguishing: a workaround the user discovers unaided; one that requires support to talk them through it; and one that requires an operator to intervene in the back end. Only the first genuinely lowers urgency much. "There is a workaround" is a weak claim until you say who performs it and how often. ## A worked example A video-transcoding queue accepts uploads, splits each into segments and writes the finished renditions back. Two defects are open. Defect A is a partial-failure rollback bug: when one segment of a multi-segment job fails, the cleanup removes the successful segments' outputs but leaves the job's index rows behind, so the job reports success with unplayable output. Data is effectively lost — critical severity. Measurement shows it fires on jobs with more than 40 segments, which are 0.4% of the 3,180 jobs a day, and only when a worker is evicted mid-job: 11 occurrences observed in three weeks. An operator can re-queue the job, and the outputs are regenerated correctly. Defect B is a label on the upload confirmation panel reading "your file will be ready in about 5 minutes" when the measured median is 27 minutes. Nothing malfunctions — trivial severity, a text string. Every uploading customer sees it, on every upload, and there is no workaround: the user simply believes the wrong thing, opens a support ticket at minute eight, and the support queue absorbs it. Defect A is far more severe. Defect B is very plausibly the higher priority: it touches 100% of a daily population, fires every time, has no user-side workaround, and its cost is continuous. Defect A is rare, operator-recoverable, and can be scheduled. Note what the argument is made of — measured reach, measured frequency, an identified workaround and its owner — and not of adjectives. ## What does *not* move priority Several things feel like priority arguments and are not. Fix effort is not one: a defect is not urgent because it is easy, though effort legitimately affects sequencing once priority is set. Age in the backlog is not one, though a persistently high-priority defect that never gets fixed is evidence that the ranking is not believed. Who found it is not one, except insofar as a customer report is evidence of real reach. And the reporter's frustration is not one at all. ## Keeping the two fields honest The discipline that makes the mixed corners work is refusing to launder a priority argument through the severity field. If the label defect is trivial, file it as trivial and argue the priority on reach and frequency. Inflating it to major to force attention wins the round and costs the team its severity data: escape analysis, per-severity reporting and release criteria all read the severity field, and once it encodes urgency instead of impact none of them mean anything.

  • Give an example of a high-severity defect you would rank low, and say what would flip it back up.
    A crash that loses in-flight work but is reachable only behind a flag never enabled outside a test environment: severity critical, priority low because reach is effectively zero. It flips the moment the flag is scheduled for enablement, or the moment anyone reports reaching it another way. That is worth writing into the deferral: the trigger that reopens the ranking, so the decision is revisited by an event rather than by whoever remembers.
  • How do you argue reach when you have no usage measurement for the broken path?
    Say so explicitly rather than guessing confidently. Then bound it: name the entry points that reach the code, the preconditions required, and the nearest measurement you do have — traffic to the containing flow, or counts of the input shape that triggers it. Offer the cheapest measurement that would settle it, such as a counter on the path. An honest bound with a stated route to certainty is more persuasive in triage than a fabricated percentage.
  • Does the existence of a workaround always lower priority?
    No, and the grade matters. A workaround the user finds unaided lowers urgency meaningfully. One that needs a support agent to walk them through converts the defect into a recurring support cost, so the priority drop is small and time-limited. One that needs an operator to intervene in the back end is not really a workaround at all — it is unpriced manual labour that scales with the defect's frequency, and it usually argues for fixing sooner rather than later.

saying these in an interview costs you the question

  • Ranks by how dramatic the failure sounds, not by exposure
  • Says a crash is always the top priority
  • Treats any workaround as fully cancelling urgency
  • Argues reach with adjectives instead of a measured share
  • Inflates severity instead of arguing the priority inputs
  • Confuses fix effort with urgency

context