skip to content

Mitigation-First Response

Stop the bleeding before finding the root cause. Interviewers test this instinct directly — candidates who start debugging instead of mitigating fail the scenario, because users are still down while you investigate.

on this pageshow

questions

6

You are on call. Five minutes after a routine deploy, your service's error rate jumps from 0.1% to 12% and users are seeing failures. What is your first action, and why is "open the logs and find the bug" the wrong one?

level: juniorimportance: must knowfreq 80%

answer

  1. users are broken while you read logs
  2. deploy timing is a strong suspect
  3. undo first, understand later
  4. mitigated is not the same as resolved
  5. previous build was serving fine

basics

~20 s

Roll back to the last known-good release first, then investigate. Restoring users is the goal during an incident, the deploy timing is strong enough evidence to act on, and reading logs leaves users broken for however long the debugging takes.

solid answer

~50 s

My first action is to revert to the previous release. A failure that starts minutes after a deploy is correlated strongly enough with that deploy to act on, and I do not need to know *which* line broke in order to undo it. Debugging first is the classic scenario failure: every minute spent reading logs is a minute of user-visible errors that a one-command rollback would already have stopped. Mitigation and diagnosis are two separate phases, and mitigation comes first. Before I revert I take a few seconds of cheap evidence so the postmortem is still possible, and I announce the action in the incident channel with a timestamp. After the rollback I verify against the actual user-facing error rate, not just the alert clearing — and if the errors do **not** drop, that itself is useful information: the deploy was not the cause, and my hypothesis needs to widen.

go deeper

for a junior

Say the words "roll back first, debug after" and mean them. Know that a failure starting minutes after a deploy makes that deploy the prime suspect, and that you do not need to identify the bug to undo the change.

for a middle

Explain why the decision works: the revert is cheap, fast and reversible, so being wrong costs two minutes while being slow costs continuous user impact. Be ready to say what evidence dies with the process and what survives centrally.

for a senior

Show that you announce the action with a timestamp, keep one failing instance out of rotation as a specimen, and verify against the user-facing success rate rather than a cleared alert. Say what a failed rollback tells you about your hypothesis.

for a principal

Own the conditions that make this reflex viable at all: reverts measured in single-digit minutes, on-call engineers authorised to pull the lever without approval, and an organisation that treats an occasional unnecessary rollback as cheaper than minutes of downtime.

## The two phases of an incident An incident has two goals that compete for the same minutes: **stop the impact** and **understand the cause**. Mitigation-first response says those are ordered, not simultaneous. You stop the bleeding, then you find out what cut you. This is the single instinct interviewers probe most directly with a scenario like this one, and candidates who start narrating a debugging session — "I'd grep the logs for the stack trace" — have already failed it, because in their story users are still getting errors the whole time. The distinction to hold is between **mitigated** and **resolved**. Mitigated means user-visible impact has stopped. Resolved means the underlying defect is gone. A rollback usually mitigates without resolving anything: the bug still exists, it is just no longer running in production. That is a perfectly good outcome at minute five. The permanent fix can be written calmly tomorrow with tests, by someone who is not sleep-deprived. ## Why the deploy timing is enough to act on You do not have proof that the deploy caused this. You have correlation: the system was healthy, a change landed, the system broke. In a production system that is one of the strongest signals available, because deploys are the most common change vector — most incidents are traceable to something a human deliberately changed shortly before. Crucially, the rollback decision does not require you to be right. It requires the action to be **cheap, fast and reversible**. Reverting to a build that was serving fine ten minutes ago has a very low downside: if the deploy was the cause, you are recovered; if it was not, you have eliminated a variable and lost a couple of minutes. Compare that with the cost of being wrong in the other direction — you read logs for twenty minutes, discover it *was* the deploy, and hand your users twenty minutes of failure you could have prevented. ## What you do in the seconds before you pull the lever Mitigation-first is not evidence-destroying recklessness. Two habits keep both goals alive: - **Capture what dies with the process.** Logs, metrics and traces that are already shipped to a central system survive a rollback. In-process state does not: thread stacks, heap contents, in-memory queues, local files on the instance. If a dump takes seconds, take it. - **Keep a specimen.** Instead of restarting or replacing every instance, pull one bad instance out of rotation and leave it running. It serves no users, so it costs nothing, and it preserves a live, reproducible copy of the failure. If the capture will take more than about a minute or requires thought, skip it. User impact outweighs postmortem quality, every time. ## Announce the action Say what you are about to do, in the incident channel, before you do it: "rolling back web to build 4412, started 14:07." Two responders independently mitigating in different directions is a genuine failure mode, and the timeline you are writing as you go is what makes the postmortem reconstructable. ## When this changes The reflex is not unconditional. It weakens or inverts when: - **The deploy is not recent.** A change that landed three days ago is a much weaker suspect — unless it was rolled out progressively and has only just reached the affected population, which is exactly the kind of detail worth checking before you dismiss it. - **The rollback is not actually safe or fast.** If an irreversible data migration ran as part of the release, the old build may no longer be able to read the data, and reverting makes things worse rather than better. - **The previous build has the same defect.** If the release merely exposed a latent bug, you will roll back into the same failure. Those cases push you toward a different lever — a feature kill switch, a failover, or a deliberate fix-forward — but note what has not changed: you are still choosing a **mitigation**, not opening a debugger. The question is only which lever, never whether to mitigate before diagnosing. ## Verifying, and what a non-recovery tells you After the rollback, watch the user-facing success rate, not the alert. If the error rate returns to baseline, you are mitigated and the incident moves into diagnosis. If it does not, resist the urge to redeploy the new build "to see": you have learned something valuable. Either the deploy was not the cause and you should widen to dependencies, traffic shifts and infrastructure, or the rollback did not fully undo the change — configuration, a schema change, or bad data the new build already wrote can all outlive the binary.

  • Does your answer change if the deploy was forty minutes ago rather than five?
    The correlation is weaker, so I would not revert on timing alone — but I still would not start with logs. I would check whether the release was rolled out gradually and has only now reached the failing population, and look for other change vectors: config pushes, dependency deploys, traffic shifts. Rollback stays on the menu as the cheapest reversible lever if nothing better presents itself quickly.
  • The rollback completes and the error rate does not drop. What now?
    Two possibilities. Either the deploy was not the cause, so I widen the hypothesis to dependencies, traffic and infrastructure and pick a different lever. Or the rollback did not fully undo the change — a config value, a schema migration, or corrupted data written by the new build can all survive reverting the binary. I check that the old version is actually serving before concluding anything, and I do not redeploy the suspect build to test the theory.
  • Everyone agrees the deploy is guilty, but the author says a one-line fix is ready. Do you take it?
    Only if it is genuinely faster and safer than the rollback, and only with a deadline. A revert restores a build that demonstrably worked; a fresh fix is untested code written under pressure and can start a second incident. If the fix is truly minutes away and the rollback is slow or blocked, I take it — with an agreed cutoff after which we revert or use a cruder lever instead.

saying these in an interview costs you the question

  • Starts by reading logs while users are still failing
  • Refuses to revert until the root cause is proven
  • Pages the change author and waits for their analysis
  • Thinks rollback means the bug is fixed
  • Dismisses a five-minute-old deploy as coincidence

context

open as a page

An incident is active and users are affected. You have several generic levers available: roll back the last release, fail over to another region or replica, flip a feature kill switch, shed load, or scale up. How do you choose between them in the first few minutes?

level: seniorimportance: must knowfreq 65%

basics

~20 s

Match the lever to the most likely change vector — code, config, a dependency, traffic, or capacity — then prefer whichever is fastest to take effect, smallest in blast radius, and easiest to undo. Apply one lever at a time so the user-facing signal tells you which one worked.

open as a page

Rolling back, restarting, or replacing instances during an incident usually destroys the state you would need to explain the failure later. What do you capture before you pull the mitigation lever, and how do you keep that capture from delaying the mitigation itself?

level: middleimportance: should knowfreq 45%

basics

~20 s

Capture only what dies with the process — thread stacks, heap state, local files, in-memory queues — and prefer quarantining one failing instance out of rotation over restarting them all. Anything already shipped to central logging, metrics or tracing survives the mitigation, so do not wait for it.

open as a page

You applied a mitigation five minutes ago and the dashboards look calmer. How do you decide the incident is actually mitigated, and what commonly makes an incident look recovered when it is not?

level: middleimportance: should knowfreq 50%

basics

~20 s

Verify against the user-facing SLI at the granularity that failed, not against a calmer dashboard or a cleared alert. The classic traps are an error ratio that fell because traffic fell, a metric window that has not turned over yet, and a backlog still draining behind a healthy-looking front door.

open as a page

Rollback is the usual first mitigation, but sometimes fixing forward is genuinely the less risky choice during a live incident. Give the concrete conditions under which you would fix forward, and how you would bound that decision.

level: seniorimportance: should knowfreq 55%

basics

~20 s

Fix forward when rollback is impossible, ineffective, or slower than the fix: an irreversible migration has already run, the previous build carries the same defect, or reverting takes forty minutes while a one-line change takes four. Bound it with a hard deadline and a cruder fallback mitigation.

open as a page

Across your organisation, incidents are detected in about three minutes but the median time to mitigate is around forty-five. As the engineering lead, what would you change so responders can stop user impact faster?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Attack the two things that fill those forty-two minutes: decision latency and lever latency. Give each service a short menu of mitigations with measured times, pre-authorise on-call engineers to pull them without approval, and treat a slow rollback or an undrilled failover as a defect to fix.

open as a page