You applied a mitigation five minutes ago and the dashboards look calmer. How do you decide the incident is actually mitigated, and what commonly makes an incident look recovered when it is not?
answer
- check the user-facing signal, not the cause graphs
- errors fell or traffic fell?
- the window still holds the incident
- the queue is still draining behind you
- mitigated is not resolved
basics
~20 sVerify against the user-facing SLI at the granularity that failed, not against a calmer dashboard or a cleared alert. The classic traps are an error ratio that fell because traffic fell, a metric window that has not turned over yet, and a backlog still draining behind a healthy-looking front door.
solid answer
~50 sI check the signal that defines user impact — request success rate and latency for real user traffic — at the granularity that actually failed. Global averages hide one shard, one region or one large tenant still broken, so if the failure was regional I verify regionally. Then I check the traps. An error *ratio* can improve because the denominator collapsed: clients gave up, retries exhausted, or an upstream is shedding, so I always read throughput next to error rate. A metric computed over a five-minute window needs a full window before it reflects the new state, so "it looks better" ninety seconds after the lever is not evidence. And a front door can look healthy while the damage is still queued behind it — check queue depth, replication lag, retry backlog and consumer lag, and confirm the trend is *decreasing*, not just non-alerting. Finally I separate **mitigated** from **resolved**: impact stopped is not cause fixed, and side effects like unsent messages or unprocessed payments still need reconciliation.
go deeper
Know that an alert clearing is not the same as users being fine, and that the signal to check is the success rate and latency of real user requests after the mitigation.
Name the false recoveries and why each fools you: a collapsed denominator, a metric window that still holds the incident, an undrained backlog, and cold caches after a restart. Say how you would check each.
State the recovery criterion out loud before reading the graph, verify on the dimension that actually failed rather than a global average, prefer a probe outside the affected path, and keep the incident open for reconciliation of side effects.
Own the distinction organisationally: define mitigated, resolved and reconciled as separate states with separate owners, and make the error budget consumed by the incident a recorded output that drives what the team does next.
## Verify against the user, not the dashboard The dashboard you were staring at during the incident was chosen for diagnosis, and it is usually full of causes: CPU, queue depth, connection counts, error logs. Those going quiet is suggestive, not conclusive. The question that decides whether the incident is mitigated is narrower: **are users getting correct responses in acceptable time again?** That means the service level indicator — the proportion of valid requests served successfully, and the latency distribution of real user traffic — read at the granularity that failed. Granularity is where most false recoveries hide. If a single shard, one region, or one very large tenant is still failing, a global success rate can look almost perfect while a real population remains completely down. Whatever dimension the failure was scoped to — region, shard, customer, endpoint, client version, device platform — verify along that same dimension. And prefer a signal measured as close to the user as you can get: an independent synthetic probe or client-reported success beats a server-side counter, especially when the failure is in the network path or in the telemetry pipeline itself. ## The four classic false recoveries **1. The denominator collapsed.** Error ratio is errors divided by requests. If clients time out, retries exhaust, a mobile app stops retrying, or an upstream starts shedding, the request count falls and the ratio improves without a single user being better off. The fix is a habit: never read an error ratio without reading throughput beside it. A sudden drop in traffic during an incident is a symptom, not a relief. **2. The window has not turned over.** A rate or ratio computed over a five-minute window carries five minutes of history. Ninety seconds after the mitigation, most of what you are seeing is still the incident. The same lag applies to the alert that fires and clears on those metrics. Wait past one full evaluation window before treating an improvement as real — and remember that any alert clearing tells you the alert's condition is no longer met, which is a weaker claim than users being fine. **3. The backlog is still draining.** Mitigation often restores the front door while the damage sits behind it: a queue that grew to millions of messages, replication lag measured in minutes, a retry backlog that is about to arrive all at once, a consumer group far behind the tip. New requests succeed, so the SLI looks healthy, while the work accepted during the incident is still unprocessed or is about to re-overload the recovered service. Verify that backlog metrics are not merely bounded but **decreasing**, and estimate the drain time from the rate of decrease before you close anything. **4. The system is still cold or fragile.** After a rollback, restart or scale-up, caches are empty, connection pools are refilling and JIT-style warmup has been lost. Latency can remain well above baseline, sometimes badly enough to breach the objective on its own, and the system is more likely to fall over again under a load it used to handle. Watch latency percentiles, not just error rate, and expect a recovery curve rather than a step change. ## Say what "mitigated" means before you declare it A usable bar, agreed out loud in the incident channel: the user-facing SLI is back inside its objective, measured on the affected dimension, sustained for longer than one full metric window, with backlog trending down and no lever still masking the problem in a way that will expire. Naming the criterion before you look at the graph keeps optimism from doing the deciding. Then keep the two states distinct. **Mitigated** means impact has stopped — this is what ends the urgency, the paging and the customer-facing severity. **Resolved** means the cause has been addressed, which usually happens days later through the postmortem's action items. Between the two sits a third obligation people forget: **reconciliation**. Requests that failed during the incident may have left side effects — unsent notifications, half-completed orders, payments captured but not recorded, duplicated writes from retries. Those need identification and repair, and the incident is not finished just because the graphs are green. ## And keep the budget in view The error budget consumed during the incident does not come back when the mitigation lands. That number — how much of the allowance for the window the incident spent — is what turns "we recovered" into a decision about what the team does next, and it belongs in the incident record alongside the recovery time.
- Right after a mitigation the error ratio drops sharply. Why is that not sufficient evidence?Because the ratio has a denominator. If clients timed out, retries exhausted, or an upstream started shedding, request volume falls and the ratio improves while users are no better off — sometimes worse off, since they have stopped even trying. Always read throughput next to the ratio; a traffic drop during an incident is a symptom, not a recovery.
- Your SLI looks healthy but a downstream queue holds two hours of backlog. Is the incident mitigated?Not for the users whose work sits in that queue. New requests succeeding says the front door is open; it says nothing about the accepted-but-unprocessed work. I would keep the incident open, track the drain rate to estimate completion, and watch for the backlog itself re-overloading the recovered service as it flushes.
- After a rollback, latency stays well above baseline even though errors are gone. What is happening?Most likely cold start effects: empty caches, refilling connection pools, and lost warmup. It usually recovers along a curve rather than a step, but it can breach the latency objective on its own and it leaves the service fragile under load it previously handled. Watch the percentiles until they settle before declaring mitigation, and consider ramping traffic back gradually.
saying these in an interview costs you the question
- Treats a cleared alert as proof users are fine
- Reads error ratio without looking at request volume
- Declares recovery within one metric window
- Ignores queue depth and replication lag
- Closes the incident without reconciling failed work