A runbook's remediation has just been fully automated, so the fix now runs without a human. What should happen to the runbook itself?
answer
- the page moves, it does not vanish
- the human now arrives after the machine failed
- how do I see it ran, and how do I stop it?
- unexercised steps rot silently
- drill it or label it unvalidated
basics
~20 sIt changes role rather than disappearing: it now documents what the automation does, how to tell whether it ran, how to disable it, and the manual fallback for when it is off or broken. That fallback is never exercised, so it must be deliberately drilled or plainly marked unvalidated.
solid answer
~50 sDeleting it is the tempting move and the wrong one, because a human still gets paged — just later and for a harder case. The runbook is rewritten around the new reality: what the automation is triggered by and what it does, how to confirm whether it ran and what it changed, how to disable it and what happens when you do, and the manual procedure for when it is disabled or has failed. The trap is that the manual path is now dead code: nobody executes it any more, so it rots silently and its first real use is during an incident where the automation already let you down. Either exercise it deliberately, or label it plainly as unvalidated so nobody mistakes it for a tested path. The diagnostic sections usually need to get stronger too — the residual pages are the cases the automation could not handle.
go deeper
Understand that automating a fix does not mean nobody gets paged for it any more — the page just happens when the automation does not do its job, and someone still needs written guidance for that case.
Explain what the rewritten document has to answer: what the automation does, how to confirm whether it ran and what it changed, how to disable it, and what the manual path is.
Show you have been on the wrong side of this. Talk about the fallback rotting because nobody executes it, the choice between drilling it and labelling it unvalidated, and the residual pages being harder than the ones automation removed.
Own the systemic version: as more remediations are automated, human competence at the manual path decays across the team. Decide what gets exercised, what is accepted as unvalidated, and how much practice time that is worth.
## Automation moves the page, it does not remove it When a remediation is automated, the frequent, well-understood instance of the failure stops reaching a human. What remains is a smaller set of pages: the automation did not fire, fired and did not fix it, fired and made things worse, or is deliberately disabled. Those are strictly harder cases than the one the original runbook covered, and they arrive rarely enough that nobody has recent practice. So the question is not whether to keep documentation, but what documentation the remaining pages need. ## What the rewritten runbook contains **What the automation does, in operational terms.** Its trigger condition, the actions it takes, its limits (how many times it will act, over what window), and its blast radius. A responder arriving at a half-fixed system needs to know what already happened to it before they add their own changes. **How to tell whether it ran.** Where its execution is recorded, what a successful run looks like, and what state the system is in after a partial one. Without this, the responder's first move is often to repeat an action that has already been taken — which is how a restart loop becomes two restart loops. **How to disable it, and what follows.** The single most important operational control, and the one people most often cannot find under pressure. State the command or toggle, who may use it, and what the system does once the automation is off — including whether the failure it was suppressing now becomes visible immediately. **The manual fallback.** The original procedure, still needed for when the automation is disabled or broken. **Escalation.** More prominent than before, because the residual cases are the ones nobody predicted. ## The rot problem, stated honestly The original runbook stayed roughly correct because people executed it and fixed what was wrong. Once the automation takes over, that repair mechanism is gone: the manual steps are no longer exercised by anyone, while the system underneath them keeps changing. Six months later the fallback references a flag that was renamed and a dashboard that was retired — and the moment it is needed is an incident where the automated path has already failed. That is the worst possible time to discover it. There are only two honest responses: 1. **Exercise it deliberately.** Run the manual procedure during a scheduled exercise, on a non-production or drained instance if possible, and treat every step that no longer works as a finding. This costs real time and it is the only way to keep the fallback genuinely usable. 2. **Mark it unvalidated.** If you will not exercise it, say so on the document: last executed date, plus a plain statement that the steps have not been verified since. A responder who knows they are on an unverified path proceeds carefully and escalates sooner. A responder who assumes it is current does not. What you must not do is leave it looking authoritative while nobody has run it in a year. ## A related decision: does the automation's failure still page? This is worth raising because it is where teams quietly lose reliability. Automation that silently retries and silently gives up removes the symptom from view without removing the problem. The runbook is the natural place to record what happens when the automation exhausts its attempts — and if the answer is "nothing", that is a gap to name rather than a design. ## The interview signal An answer that says "we automated it, so we deleted the runbook" tells the interviewer the candidate has not been on the other side of an automation failure. The answer that lands describes the role change — from *how to fix it* to *what the machine does, how to see it, how to stop it, and what to do when it is off* — and is honest about the fallback decaying. Conceding that the drill costs time you may not always spend is more credible than claiming the fallback stays fresh on its own.
- Why does the responder need to know what the automation already did before acting?Because they are arriving at a system someone else has partially changed. Without a record of what ran and what state it left behind, the natural move is to repeat the action — restarting something that has already been restarted twice, or failing over a component the automation just failed over. Knowing what happened is a precondition for every step that follows.
- Should the runbook document how to disable the automation, given that turning it off could make the outage worse?Yes — an operational control nobody can find under pressure is not a control. Document the toggle, who is authorised to use it, and explicitly what happens after: whether the underlying failure becomes immediately visible, and what the responder is now on the hook for. The risk you named is an argument for stating the consequence, not for hiding the switch.
saying these in an interview costs you the question
- Deletes the runbook because the fix is automated now
- Keeps the manual fallback and assumes it still works
- Cannot say how to turn the automation off
- Ignores that the remaining pages are the harder cases
- Treats an unexercised procedure as a validated one