You are on call. Five minutes after a routine deploy, your service's error rate jumps from 0.1% to 12% and users are seeing failures. What is your first action, and why is "open the logs and find the bug" the wrong one?
answer
- users are broken while you read logs
- deploy timing is a strong suspect
- undo first, understand later
- mitigated is not the same as resolved
- previous build was serving fine
basics
~20 sRoll back to the last known-good release first, then investigate. Restoring users is the goal during an incident, the deploy timing is strong enough evidence to act on, and reading logs leaves users broken for however long the debugging takes.
solid answer
~50 sMy first action is to revert to the previous release. A failure that starts minutes after a deploy is correlated strongly enough with that deploy to act on, and I do not need to know *which* line broke in order to undo it. Debugging first is the classic scenario failure: every minute spent reading logs is a minute of user-visible errors that a one-command rollback would already have stopped. Mitigation and diagnosis are two separate phases, and mitigation comes first. Before I revert I take a few seconds of cheap evidence so the postmortem is still possible, and I announce the action in the incident channel with a timestamp. After the rollback I verify against the actual user-facing error rate, not just the alert clearing — and if the errors do **not** drop, that itself is useful information: the deploy was not the cause, and my hypothesis needs to widen.
go deeper
Say the words "roll back first, debug after" and mean them. Know that a failure starting minutes after a deploy makes that deploy the prime suspect, and that you do not need to identify the bug to undo the change.
Explain why the decision works: the revert is cheap, fast and reversible, so being wrong costs two minutes while being slow costs continuous user impact. Be ready to say what evidence dies with the process and what survives centrally.
Show that you announce the action with a timestamp, keep one failing instance out of rotation as a specimen, and verify against the user-facing success rate rather than a cleared alert. Say what a failed rollback tells you about your hypothesis.
Own the conditions that make this reflex viable at all: reverts measured in single-digit minutes, on-call engineers authorised to pull the lever without approval, and an organisation that treats an occasional unnecessary rollback as cheaper than minutes of downtime.
## The two phases of an incident An incident has two goals that compete for the same minutes: **stop the impact** and **understand the cause**. Mitigation-first response says those are ordered, not simultaneous. You stop the bleeding, then you find out what cut you. This is the single instinct interviewers probe most directly with a scenario like this one, and candidates who start narrating a debugging session — "I'd grep the logs for the stack trace" — have already failed it, because in their story users are still getting errors the whole time. The distinction to hold is between **mitigated** and **resolved**. Mitigated means user-visible impact has stopped. Resolved means the underlying defect is gone. A rollback usually mitigates without resolving anything: the bug still exists, it is just no longer running in production. That is a perfectly good outcome at minute five. The permanent fix can be written calmly tomorrow with tests, by someone who is not sleep-deprived. ## Why the deploy timing is enough to act on You do not have proof that the deploy caused this. You have correlation: the system was healthy, a change landed, the system broke. In a production system that is one of the strongest signals available, because deploys are the most common change vector — most incidents are traceable to something a human deliberately changed shortly before. Crucially, the rollback decision does not require you to be right. It requires the action to be **cheap, fast and reversible**. Reverting to a build that was serving fine ten minutes ago has a very low downside: if the deploy was the cause, you are recovered; if it was not, you have eliminated a variable and lost a couple of minutes. Compare that with the cost of being wrong in the other direction — you read logs for twenty minutes, discover it *was* the deploy, and hand your users twenty minutes of failure you could have prevented. ## What you do in the seconds before you pull the lever Mitigation-first is not evidence-destroying recklessness. Two habits keep both goals alive: - **Capture what dies with the process.** Logs, metrics and traces that are already shipped to a central system survive a rollback. In-process state does not: thread stacks, heap contents, in-memory queues, local files on the instance. If a dump takes seconds, take it. - **Keep a specimen.** Instead of restarting or replacing every instance, pull one bad instance out of rotation and leave it running. It serves no users, so it costs nothing, and it preserves a live, reproducible copy of the failure. If the capture will take more than about a minute or requires thought, skip it. User impact outweighs postmortem quality, every time. ## Announce the action Say what you are about to do, in the incident channel, before you do it: "rolling back web to build 4412, started 14:07." Two responders independently mitigating in different directions is a genuine failure mode, and the timeline you are writing as you go is what makes the postmortem reconstructable. ## When this changes The reflex is not unconditional. It weakens or inverts when: - **The deploy is not recent.** A change that landed three days ago is a much weaker suspect — unless it was rolled out progressively and has only just reached the affected population, which is exactly the kind of detail worth checking before you dismiss it. - **The rollback is not actually safe or fast.** If an irreversible data migration ran as part of the release, the old build may no longer be able to read the data, and reverting makes things worse rather than better. - **The previous build has the same defect.** If the release merely exposed a latent bug, you will roll back into the same failure. Those cases push you toward a different lever — a feature kill switch, a failover, or a deliberate fix-forward — but note what has not changed: you are still choosing a **mitigation**, not opening a debugger. The question is only which lever, never whether to mitigate before diagnosing. ## Verifying, and what a non-recovery tells you After the rollback, watch the user-facing success rate, not the alert. If the error rate returns to baseline, you are mitigated and the incident moves into diagnosis. If it does not, resist the urge to redeploy the new build "to see": you have learned something valuable. Either the deploy was not the cause and you should widen to dependencies, traffic shifts and infrastructure, or the rollback did not fully undo the change — configuration, a schema change, or bad data the new build already wrote can all outlive the binary.
- Does your answer change if the deploy was forty minutes ago rather than five?The correlation is weaker, so I would not revert on timing alone — but I still would not start with logs. I would check whether the release was rolled out gradually and has only now reached the failing population, and look for other change vectors: config pushes, dependency deploys, traffic shifts. Rollback stays on the menu as the cheapest reversible lever if nothing better presents itself quickly.
- The rollback completes and the error rate does not drop. What now?Two possibilities. Either the deploy was not the cause, so I widen the hypothesis to dependencies, traffic and infrastructure and pick a different lever. Or the rollback did not fully undo the change — a config value, a schema migration, or corrupted data written by the new build can all survive reverting the binary. I check that the old version is actually serving before concluding anything, and I do not redeploy the suspect build to test the theory.
- Everyone agrees the deploy is guilty, but the author says a one-line fix is ready. Do you take it?Only if it is genuinely faster and safer than the rollback, and only with a deadline. A revert restores a build that demonstrably worked; a fresh fix is untested code written under pressure and can start a second incident. If the fix is truly minutes away and the rollback is slow or blocked, I take it — with an agreed cutoff after which we revert or use a cruder lever instead.
saying these in an interview costs you the question
- Starts by reading logs while users are still failing
- Refuses to revert until the root cause is proven
- Pages the change author and waits for their analysis
- Thinks rollback means the bug is fixed
- Dismisses a five-minute-old deploy as coincidence