You declared a SEV3 for elevated checkout errors. Forty minutes in, you discover the same bug also wrote incorrect balances to a subset of accounts. How do you handle severity mid-incident, and what are the rules for downgrading one?
answer
- severity is a dial, not a label
- data damage escalates on its own axis
- the silent upgrade is the real failure
- downgrade only when impact has stopped
- peak severity is what gets recorded
basics
~20 sUpgrade immediately on discovery, not at the next sync, and restate the grade explicitly with the time. Downgrade only after impact has actually stopped and recovery is verified — never to quiet the response — and record the peak severity, not the closing one.
solid answer
~50 sSeverity is a live variable, not a label set at declaration. Confirmed data corruption changes the grade regardless of what the error-rate graph says, so I upgrade the moment I know, and I do it explicitly in the incident channel with a timestamp — a silent upgrade, where the team quietly behaves like it is a SEV1 while the record still says SEV3, is the worse failure because everyone outside the channel is still calibrated to the old number. Downgrading has one legitimate trigger: impact has genuinely stopped and I have verified recovery against the service's own indicators, not against the absence of new complaints. I never downgrade because responders are tired or the pager is noisy. And downgrading is not resolving — the incident stays open until it is actually closed, while the *recorded* severity stays at its peak, because that is what the review and reporting obligations should key off.
go deeper
Know that an incident's severity can change while it is running, and that you raise it as soon as you learn the impact is worse. Say the new grade out loud in the incident channel rather than just working faster.
Explain what triggers a change in each direction: new impact evidence or unexpected duration raises it, verified recovery lowers it. Be clear that downgrading is not the same as resolving and that data damage escalates independently of error rates.
Show you have made these calls under pressure. Name the silent upgrade as a failure mode, explain why 'the fix is deployed' is not grounds to downgrade, and articulate the peak-severity convention and the incentive it protects.
Own the policy. Explain how you would build duration-based re-grade checkpoints into the process, how peak-severity recording removes the incentive to downgrade early, and how you would push back when a senior stakeholder wants a lower number in the record.
## Severity is a live variable The most common mistake with severity is treating it as a label applied once at declaration. It is a dial, and moving it is a normal part of running an incident. You declared under uncertainty; the whole point of declaring early is that you expected to learn more. What matters is that the dial moves *deliberately and visibly* rather than drifting. ## Upgrading: on discovery, out loud In this scenario the trigger is unambiguous. Elevated checkout errors are a request-availability problem — bounded, self-healing once fixed, and visible on a graph. Corrupted stored balances are a different class: the damage persists after the bug is fixed, it accumulates silently while you investigate, it may require a reconciliation and customer contact afterwards, and it can start a regulatory or contractual clock. Data correctness escalates on its own axis, independent of how many requests are currently failing. So the grade goes up. Three things about *how* to upgrade: **Immediately, not at the next checkpoint.** Waiting for a scheduled sync to change the number means every decision made in between is made against a stale grade. **Explicitly, with a time.** "14:52 — raising to SEV1: confirmed incorrect balances written to an unknown number of accounts since roughly 14:05." That single line does three jobs: it re-routes the response, it gives everyone outside the channel the new calibration, and it becomes a load-bearing entry in the timeline. **With the obligations attached to the new tier actually firing.** An upgrade that changes a field but not who is engaged is theatre. If the higher tier means additional responders, a different notification set, or specialist involvement, that happens now. The failure mode to name here is the **silent upgrade**: the team recognises the severity in their behaviour — more people pile in, urgency rises — but nobody restates the grade. Stakeholders reading the record, and anyone deciding whether to engage, are still working from SEV3. The gap between how the incident is being run and how it is recorded is where surprised executives and inaccurate external communication come from. A second upgrade trigger worth having explicitly: **duration**. An incident that has run far longer than expected at its tier, with no bounded understanding of impact, deserves a re-grade even when the numbers have not moved. "We still don't know" after an hour is itself information. ## Downgrading: one legitimate trigger Downgrading is where judgment goes wrong, because the pressure to do it is real — responders are tired, the channel is loud, and the graphs look better. Only one trigger is legitimate: **impact has actually stopped, and recovery is verified against the service's own indicators.** Not "we deployed the fix" — deploying is not recovering. Not "complaints stopped" — the absence of new reports is weak evidence, since reporting lags impact badly at night and during weekends. And note the asymmetry with data damage: for a corruption incident, stopping the *writes* is not the same as ending the *impact*. Wrong balances still sit in the database. The grade should not fall until the extent is known and the remediation path is established, even if no new bad records are being created. Illegitimate reasons, all of which show up in real incidents: - Responder fatigue or a noisy channel — that is a staffing and handover problem, not a severity problem. - The cause turned out to be a third party — your users cannot tell, and neither should your grade. - Someone senior finds the number uncomfortable — this corrupts the data everything downstream depends on. - A desire to stop the clock on a response-time target — grading to a metric is the fastest way to make the metric meaningless. ## Downgrade is not resolve, and peak severity is what is recorded Two conventions keep the record honest. First, **downgrading is not resolving**. A SEV1 that becomes a SEV3 is still an open incident with an owner, still being worked, still being communicated. Teams that conflate the two stop the response while remediation — reconciling the bad data, confirming no further accounts are affected — is still outstanding. Second, most organisations record the **peak severity** reached, not the severity at close. The reason is purely practical: the obligations attached to a tier — the depth of review, who is informed, what reporting happens — should key off the worst the incident actually was. If the closing grade governed, then downgrading before closure would quietly discharge those obligations, and you would have built an incentive to downgrade early. Keeping the peak also makes the incident history usable: a quarter's SEV1 count means something only if it counts every incident that was genuinely a SEV1 at some point.
- What is wrong with a team that behaves like it's a SEV1 but never changes the recorded grade?That's a silent upgrade, and it is worse than being slow to raise. Everyone inside the channel is calibrated correctly; everyone outside — people deciding whether to join, stakeholders reading the record, whoever handles external communication — is still working from the old number. The record and the reality diverge exactly when accuracy matters most.
- Is deploying the fix enough to downgrade?No. Deploying is not recovering. Downgrade when the service's own indicators show impact has actually stopped and stayed stopped, not when a pipeline goes green. And for data damage the bar is higher still: stopping the bad writes does not undo the bad rows, so the grade should hold until the extent is known and remediation is under way.
- Why record peak severity rather than the severity at close?Because the obligations attached to a tier — review depth, who is informed, what gets reported — should reflect the worst the incident actually was. If the closing grade governed, downgrading before closure would quietly discharge those obligations, creating a direct incentive to downgrade early. Peak severity also makes the incident history comparable across quarters.
- Can duration alone justify an upgrade when the numbers haven't moved?Yes. An incident running far beyond what its tier normally implies, with no bounded understanding of impact, has told you something: your model of the failure is wrong. Impact also integrates over time, so a small rate sustained for hours can exceed a larger one that lasted minutes. Build a duration-based re-grade checkpoint in rather than relying on someone noticing.
saying these in an interview costs you the question
- Severity is fixed once the incident is declared
- Downgrade to make the pager stop
- Downgrading is the same as resolving
- The fix is deployed, so it's over
- Third-party cause, so we lower the grade