skip to content

Severity Levels & Incident Declaration

How teams decide something is an incident and how bad it is. Interviewers ask about severity criteria because over- and under-declaring both have real costs, and calibrated judgment here marks operational maturity.

on this pageshow

questions

5

Your team is writing a SEV1–SEV4 severity matrix for its production services. Which dimensions should decide an incident's severity, and why must the matrix be written in terms of user impact rather than which component broke?

level: middleimportance: must knowfreq 72%

answer

  1. impact, not component
  2. how many, which journey
  3. workaround buys you a tier
  4. data and money have their own axis
  5. one worked example per tier

basics

~20 s

Grade on observable user impact: what share of users or requests is affected, how critical the blocked journey is, whether a workaround exists, and whether data or money is at risk. Which component failed predicts none of those.

solid answer

~50 s

A severity grade is a routing decision, not a description — it decides how many people drop what they are doing, how fast, who outside engineering hears about it, and what review the incident earns. So the matrix should only contain facts an on-call engineer can establish in five minutes: **breadth** (share of requests, users or tenants affected — ideally readable off a graph), **criticality of the blocked journey** (checkout versus avatar upload), **whether a workaround exists** (usually worth one tier), and **data or money at risk**, which is its own axis because corruption, loss or exposure is often irreversible and starts legal clocks even at zero visible errors. Component identity fails in both directions: a lost database replica with automatic failover has no impact, while a "minor" flag service can gate login. Give every tier one observable threshold plus a real worked example, and a default rule: if you cannot choose between two tiers, take the higher one.

code

yaml · 17 lines
yaml
severity_matrix:
  sev1:
    user_impact: "critical journey unavailable, or >10% of requests failing"
    data_or_money: "confirmed loss, corruption, exposure, or mis-posting"
    example: "2026-02-11 checkout returned 5xx for all users for 22 minutes"
  sev2:
    user_impact: "critical journey degraded, or one region/large tenant down"
    workaround: "exists and can be communicated"
    example: "2026-01-04 EU search latency 8s, US unaffected"
  sev3:
    user_impact: "non-critical journey, or <1% of requests"
    example: "2025-12-19 avatar uploads failing for iOS clients"
  sev4:
    user_impact: "none observed"
    risk: "redundancy or protection lost (one AZ, failed backup)"
    example: "2026-03-02 nightly backup job silently skipped"
tie_break: "if torn between two tiers, take the higher one and adjust later"

go deeper

for a junior

Know the shape of a severity matrix and that SEV1 is the most serious. Be able to say that severity is judged by how many users are affected and how important the broken feature is, and know that when in doubt you take the higher tier.

for a middle

Be ready to name the dimensions — breadth, journey criticality, workaround, data and money at risk — and explain why a component-keyed matrix fails in both directions. Show that you can write a tier definition an on-call engineer could actually apply at 3 a.m.

for a senior

Show you have used one under pressure: give a real grading call you made, including a case where the graphs looked fine but data was wrong. Explain how you keep tier definitions from rotting and how each tier maps to concrete response obligations.

for a principal

Own the incentive design. Argue how tier boundaries and their attached obligations change declaration behaviour across teams, why baking a responder roster into a tier drives under-declaration, and how you would normalise breadth thresholds between a payments service and an internal tool.

## What a severity grade is actually for A severity is not a description of how bad the system feels; it is a routing decision made under time pressure. The grade determines how many people stop what they are doing and how quickly, whether anyone outside engineering is told, and what kind of review the incident earns afterwards. If changing the grade changes none of those things, the matrix is decoration. And because it is decided in the first minutes by one tired person, the only inputs that belong in it are facts that person can establish quickly and defend later. ## The dimensions that belong on the matrix **Breadth of impact.** What share of requests, users, or tenants is affected. Prefer a request-share or error-rate threshold over a headcount, because the responder can read it off an existing graph instead of estimating. Whole-segment failures count as breadth too: one region, one mobile platform, one large tenant. **Criticality of the blocked journey.** Not all functionality is equal. Ranking your user journeys *before* an incident — checkout and login at the top, secondary features below — turns a 3 a.m. argument into a lookup. **Availability of a workaround.** A documented, communicable workaround usually moves an incident down a tier: the impact is real but bounded, because affected users have somewhere to go. **Data and money.** This is a separate axis and it does not scale with user count. Confirmed data loss, corruption, or exposure — or money posted incorrectly — belongs at or near the top tier even when error rates and latency graphs look perfectly normal, because the damage is often irreversible, silently accumulating, and may start a regulatory clock. **Direction and elapsed time.** A small impact that is growing with no known bound outranks a static one of the same size, and impact integrates over time: "2% of requests for four hours" is a different incident from "2% for four minutes." **Loss of protection with no user impact yet.** Losing one of two availability zones, or a failed backup job, may be invisible to users and is still worth declaring at a low tier — you are now one failure away from a large one, and nobody will notice on their own. ## Why component identity is the wrong axis It is tempting to write "primary database down = SEV1, cache down = SEV3." This fails in both directions. A database replica lost behind working automatic failover produces no user impact at all; meanwhile a component nobody thinks of as critical — a configuration service, a feature-flag evaluator, an auth token issuer — can block every login. Component-keyed matrices also rot as the architecture changes, and they push responders into arguing about which box is broken instead of measuring what users are experiencing. An impact-based matrix survives a re-architecture untouched. The same logic rules out two other tempting axes. Do not grade by **how hard the fix looks** — difficulty is a property of the repair, not the impact, and a one-line rollback can end a genuine SEV1. Do not grade down because the **cause is a third party**; your users cannot tell the difference, and outsourcing the cause does not outsource the impact. ## Making the tiers usable at 3 a.m. The most common defect in a real matrix is circular wording: "SEV1 — a critical production system is severely impacted." That gives a responder nothing. Each tier needs at least one observable threshold and at least one worked example taken from a real past incident, because calibration comes from examples far more than from prose. A workable skeleton: - **SEV1** — a critical user journey is unavailable or broadly failing, or there is confirmed data loss, corruption, exposure, or financial mis-posting. - **SEV2** — major degradation of a critical journey, or a whole segment (a region, a large tenant) affected, with a workaround or partial service remaining. - **SEV3** — limited or partial impact: a small user segment or a non-critical journey. - **SEV4** — no current user impact, but protection or headroom is lost, or the issue is cosmetic. Keep the tier count small; four is common, and beyond five teams spend more time arguing about boundaries than responding. State the tie-break explicitly — take the higher tier and adjust — and attach each tier to a short, concrete list of obligations so the grade means something operationally. ## Where matrices go wrong in practice Baking a fixed responder roster into the tier ("SEV1 means these six names") makes people grade down to avoid disturbing those six. Treating internal-only outages as automatically low ignores that deploy pipelines, dashboards, and paging systems are load-bearing during an incident. And a matrix nobody ever re-grades against closed incidents drifts within months into whatever each team privately believes.

  • Would you ever declare a high severity when no user is affected at all?
    Yes, when protection is gone rather than service. Confirmed data corruption with normal error rates is the clearest case — irreversible and often invisible. Losing redundancy (one of two zones, or backups silently failing) is usually a lower tier but still a declared incident, because you are one failure from a large one and nobody discovers it spontaneously.
  • An outage only affects an internal tool. Can that still be a high severity?
    It depends entirely on what the tool carries. An internal wiki is low. A deploy pipeline, a paging system, or the dashboards you triage with are load-bearing during incidents — losing them removes your ability to respond to everything else, so grade them by the impact of that loss, not by the fact that no customer sees the tool.
  • How many severity tiers should a matrix have?
    Four is the common shape and works well. Fewer than three cannot distinguish "wake everyone" from "handle it in hours." More than five and the boundaries stop being distinguishable at 3 a.m., so responders debate the grade instead of responding. If people routinely argue between two adjacent tiers, the tiers are not separated by any real difference in obligation.

A hospital triage nurse grades on what the patient presents with — bleeding, airway, consciousness — not on which department will end up doing the surgery.

saying these in an interview costs you the question

  • Severity is decided by which service or component failed
  • Anything happening in production is automatically a SEV1
  • Grades the incident by how hard the fix looks
  • Third-party cause means we grade it lower
  • Internal-only outages can never be high severity

context

open as a page

In production operations, what is the difference between a monitoring alert firing and an incident being declared, and what happens in between?

level: juniorimportance: should knowfreq 55%

basics

~20 s

An alert is an automated signal that some condition was met. An incident is a human declaration that real or suspected impact warrants coordinated response, with a severity, an owner and a written record. Triage is the step between them.

open as a page

It is 02:10 and you have one ambiguous signal: elevated 5xx errors on one of six API instances, no customer reports yet. Make the case for declaring an incident now on suspicion versus waiting for confirmation, and say how you would make that call repeatable for your team.

level: seniorimportance: should knowfreq 62%

basics

~20 s

Declare now. The costs are asymmetric: an over-declaration is undone with one message and a few interrupted minutes, while a late declaration adds unmitigated impact plus the response ramp-up you could have run in parallel. Declare low and adjust.

open as a page

You declared a SEV3 for elevated checkout errors. Forty minutes in, you discover the same bug also wrote incorrect balances to a subset of accounts. How do you handle severity mid-incident, and what are the rules for downgrading one?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Upgrade immediately on discovery, not at the next sync, and restate the grade explicitly with the time. Downgrade only after impact has actually stopped and recovery is verified — never to quiet the response — and record the peak severity, not the closing one.

open as a page

Across forty engineering teams, incident severity grades are wildly inconsistent — one team's SEV1 is another's SEV3, and one team has not declared anything above SEV3 in a year. As the lead of the reliability practice, how would you calibrate severity across the organisation?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Treat it as an incentive problem, not a wording problem. Diagnose whether teams are inflating or deflating, calibrate with worked reference incidents rather than longer definitions, re-grade a sample of closed incidents regularly, and decouple punishing process from high tiers.

open as a page