Across forty engineering teams, incident severity grades are wildly inconsistent — one team's SEV1 is another's SEV3, and one team has not declared anything above SEV3 in a year. As the lead of the reliability practice, how would you calibrate severity across the organisation?
answer
- it is an incentive problem, not wording
- diagnose inflation versus deflation first
- calibrate on cases, not definitions
- re-grade closed incidents together
- separate impact from response urgency
basics
~20 sTreat it as an incentive problem, not a wording problem. Diagnose whether teams are inflating or deflating, calibrate with worked reference incidents rather than longer definitions, re-grade a sample of closed incidents regularly, and decouple punishing process from high tiers.
solid answer
~60 sRewriting the definitions is the intervention that never works, because people are not misreading the words — they are responding to consequences. First I diagnose direction per team: inflation usually means severity is being used to get attention or resources that are hard to obtain otherwise, while deflation almost always means high tiers carry heavy or career-adjacent process, so teams grade down to avoid it. Then I calibrate with examples rather than prose: a published library of real incidents with their grade and the reasoning, which is how humans actually converge. I add a periodic cross-team review that re-grades a sample of closed incidents and feeds disagreements back, which turns calibration into a measurable, ongoing thing. Structurally, I make high tiers cheap to declare and expensive only in review depth, and I separate impact severity from response urgency so a large-but-contained problem does not require a 3 a.m. all-hands. Finally I normalise breadth thresholds by service criticality tier — a payments service and an internal tool cannot share one absolute user count.
go deeper
Know that severity tiers are meant to mean the same thing across teams, and that grading down to avoid extra process defeats the purpose. If your team's grades feel inconsistent with others', that is worth raising rather than quietly working around.
Be able to explain why a team might systematically under-declare, and why worked examples of past incidents calibrate people better than longer written definitions do.
Show you have seen the incentive dynamics first-hand: name what a high grade actually costs a team in your organisation and how that shapes what gets declared. Be able to argue for re-grading closed incidents as a routine practice.
Own the whole system: diagnose inflation versus deflation from data, redesign the consequences attached to each tier, separate impact from urgency, normalise thresholds by service criticality, and be honest that trust rebuilds over quarters rather than with a policy announcement.
## The wrong first move The instinctive fix is to rewrite the severity matrix with more precise language. It almost never works, because the teams are not confused about what the words mean. A team that has not declared above SEV3 in a year is not misreading a definition — it is responding rationally to what happens when it declares a SEV2. Severity grading is a behaviour, and behaviours follow consequences. Start with the consequences. ## Diagnose the direction, per team The two failure directions have opposite causes and opposite fixes, and an organisation usually has both at once. **Inflation** — everything is a SEV1. Typically severity has become the only lever that reliably gets attention: the only way to pull in a specialist team, get an exception to a freeze, or make a dependency owner respond. The severity dial is being used as a priority dial. Fix the underlying access problem and inflation subsides on its own. **Deflation** — nothing is ever bad. Almost always the high tiers carry weight that feels punitive: a mandatory review with an audience, executive visibility, a mark that follows the team into performance conversations, or a reliability metric the team is measured on. Sometimes it is simpler and sadder — declaring high wakes people up, and the team does not want to be the team that does that. Either way, the team is paying for honesty, so it stops being honest. The evidence is in the data. Look at grade distribution per team over time, compare it to the same teams' user-facing impact measured independently, and look at how many incidents were declared at all. A team with normal outage evidence but a compressed severity distribution is deflating. ## Calibrate with examples, not prose Humans calibrate against cases far better than against definitions. The single highest-leverage artefact is a **library of worked reference incidents**: real incidents from across the organisation, each with its final grade and two or three sentences of reasoning, including the near-misses where the grade was genuinely arguable. "This looked like a SEV1 but had a documented workaround and stayed at SEV2, here is why" teaches more than a paragraph of tier definition ever will. Pair it with a recurring **re-grading review**: a small cross-team group takes a sample of closed incidents from the last month and independently grades them, then compares. Disagreements are the product — they tell you exactly which boundary is ambiguous and which team is drifting, and the discussion itself is the calibration mechanism. Publish the disagreements, not just the conclusions. ## Change the structure, not just the guidance **Decouple burden from tier.** Make declaring a high severity organisationally cheap and only *analytically* expensive. A serious incident should earn a thorough review; it should not earn a hostile audience. If a tier triggers something the team experiences as punishment, you have built a machine that manufactures under-declaration, and no amount of encouragement will out-argue it. **Separate impact severity from response urgency.** These are genuinely different axes and collapsing them into one number forces bad choices. A large data-correctness problem that has already stopped may deserve the top severity for review and notification purposes while needing no overnight response at all; a moderate but rapidly growing outage is the reverse. Two fields, each with its own meaning, removes the incentive to misgrade impact in order to control who gets woken. **Normalise breadth by service criticality.** An absolute user-count threshold cannot be shared between a payments service and an internal reporting tool. Establish criticality tiers for services first, then express severity thresholds relative to that tier. This is also what makes cross-team comparison meaningful, because you are comparing like with like. **Keep the taxonomy small and centrally owned.** Per-team matrices guarantee divergence. One matrix, owned by the reliability practice, with a documented path to propose changes. ## What to accept Some variance is legitimate and chasing it is waste. Domains genuinely differ: a research-facing service and a regulated payments flow will apply the same principles to different consequences. Perfect inter-rater agreement is not the goal — the goal is that a given grade means roughly the same *urgency and obligation* everywhere, so that a person from another team reading it knows what to do. And be honest about the timeline. This is a trust problem, and the deflating team will keep deflating until it has watched a couple of high-severity declarations happen to other teams without consequence. Removing the punishment is necessary but not sufficient; the evidence has to accumulate publicly before the behaviour moves.
- How would you tell inflation from deflation in the data?Compare each team's grade distribution against independent evidence of user impact — availability indicators, support ticket volume, customer-visible downtime. A team with normal outage evidence and a distribution compressed at the bottom is deflating; a team whose top-tier count far exceeds its measurable impact is inflating. Also check total declarations, since deflating teams often under-declare entirely rather than just grading low.
- Why separate impact severity from response urgency?Because collapsing them forces people to misgrade one to control the other. A large data-correctness problem that has already stopped may deserve a top grade for review and notification while needing nobody woken; a growing partial outage is the opposite. Two fields let a responder be honest about impact without deciding, as a side effect, who loses a night's sleep.
- Isn't perfect consistency across forty teams unrealistic?Yes, and chasing it is waste. The goal is that a grade conveys roughly the same urgency and obligation everywhere, so someone from another team can read it and know what to do. Legitimate domain differences remain — regulated payments and an internal reporting tool apply the same principles to different consequences — which is why you normalise thresholds by service criticality rather than by absolute counts.
- How long does this take to actually change behaviour?Longer than the policy change. The deflating team is protecting itself, and it will keep protecting itself until it has watched other teams declare high severities publicly with no adverse consequence. Removing the punishment is necessary but not sufficient; the counter-evidence has to accumulate in view before grading behaviour moves, typically over several quarters.
saying these in an interview costs you the question
- Fix inconsistency by rewriting the definitions more precisely
- Let each team define its own severity matrix
- Require senior approval for every high-severity declaration
- Under-declaration means the team needs more training
- One absolute user-count threshold fits every service