skip to content

Your team's runbooks have gone stale — commands reference a decommissioned host and several document alerts that no longer exist. How do you make runbook freshness a property of the process rather than a one-off cleanup project?

level: seniorimportance: must knowfreq 58%

answer

  1. attach maintenance to moments that already happen
  2. whoever used it, fixes it — same shift
  3. the link belongs to the alert
  4. last-validated date is information, not paperwork
  5. a wrong runbook beats no runbook only in theory

basics

~20 s

Tie maintenance to use: whoever follows a runbook during a page repairs it in the same shift. Bind each runbook to the alert that links it so orphans are visible, review them at on-call handoff, stamp owner and last-validated date, and delete more than you write.

solid answer

~50 s

A cleanup sprint fixes the symptom and the documents rot again, so I attach maintenance to the moments that already happen. First, **touch-on-use**: whoever follows a runbook during a page fixes what was wrong with it before the shift ends — a five-minute edit counted as part of handling the page, not as follow-up work. Second, **bind runbooks to alerts**: the link lives in the alert definition, so deleting an alert surfaces its orphaned runbook and a page arriving with no runbook is a visible gap. Third, use the **on-call handoff** as the recurring forcing function — the outgoing on-caller reports which runbooks were used and which were wrong. Fourth, put **owner and last-validated date** on every one, so a responder can weigh a 14-month-old procedure appropriately. Fifth, **prune hard**: forty exercised runbooks beat four hundred aspirational ones. Runbooks that never fire only get validated by drilling them deliberately.

go deeper

for a junior

Know that a runbook you just followed is the best time to fix it — note the wrong hostname or the missing step while you still remember, rather than assuming someone else will notice.

for a middle

Explain the mechanisms rather than the aspiration: correction attached to use, runbook links stored with the alert definition, owner and last-validated metadata as a trust signal for the next responder.

for a senior

Show that you have watched documents rot. Argue why cleanup projects only reset the clock, why a wrong runbook is more dangerous than a missing one, and how the on-call handoff becomes the recurring forcing function without becoming a formality.

for a principal

Own the cost side across an org: how much incident time you are willing to spend on maintenance, how you keep rarely exercised procedures honest, and how you avoid a coverage metric that rewards volume of documents over documents that work.

## Why cleanup projects do not work A runbook rots because the system underneath it changes continuously while the document changes only when someone chooses to edit it. A cleanup sprint resets the clock; it does not change the rate. Six months later you are back, with the added problem that people now trust the documents slightly more than they should. The fix has to be structural: make correction happen at moments that already occur, and make staleness visible without anyone auditing anything. ## Touch-on-use The single highest-value rule. Whoever follows a runbook during a real page corrects it in the same shift — a wrong hostname, a renamed dashboard, a step that no longer applies, a missing precondition they had to work out themselves. Two things make it work in practice: - **Count the edit as incident work, not as a follow-up.** If it becomes a ticket, it ages in a backlog behind feature work. Five minutes at the end of the page, while the detail is fresh, is when it is cheapest. - **Make the edit path trivial.** If updating requires a pull request, two reviewers and a docs pipeline, a tired responder will not do it and you have designed the rot in. Whether runbooks live in a wiki or in the repo, the friction has to be near zero for the person holding the pager. This rule also has a pleasant property: it concentrates maintenance exactly on the runbooks that get used, which are the ones whose accuracy matters most. ## Bind the runbook to the alert When the runbook link is part of the alert definition rather than a separate wiki page, two failure modes become detectable without an audit: - An alert with no runbook link shows up as a gap the moment it pages. - A runbook nobody links to is either documenting an alert that was deleted, or documenting a failure mode nothing watches — both worth knowing. The practice tree owns the policy here, not the tooling: whichever alerting system you run, the point is that the association is data, not folklore. ## The handoff as the forcing function On-call handoff already happens on a schedule, already involves the two people who most recently held the pager, and already reviews what fired. Adding "which runbooks did you use, and which were wrong" costs a couple of minutes and produces a list from the only reliable source: actual execution. It also spreads knowledge of what has been changing. A quarterly documentation review meeting, by contrast, tends to review the runbooks that were easy to review rather than the ones that matter, and it decays into a formality. ## Metadata as a decay signal Every runbook carries an owner and a last-validated date. The date is not bureaucracy — it is information the responder uses in the moment. Seeing "last validated three weeks ago" versus "last validated 14 months ago" changes how much you trust the commands and how carefully you check preconditions before running them. The owner is who repairs it when a responder finds it wrong and has no context to fix it themselves. ## Prune, and prune aggressively The instinct is to keep everything because deletion loses knowledge. Weigh that against the real cost: a wrong runbook is worse than a missing one, because a missing one produces caution and escalation, while a wrong one produces confident action. A responder who cannot tell which of four hundred documents is current effectively has none. So delete runbooks for alerts that no longer exist, for services that were decommissioned, and for failures that architecture changes made impossible. If something feels too valuable to delete, that is a signal to validate it rather than to keep it unexamined. ## The runbooks that never fire The hardest category is the procedure for the rare event — a regional failover, a data restore. It cannot be maintained by use because it is never used, and it is exactly the one you cannot afford to have wrong. The only honest options are to exercise it deliberately, or to mark it clearly as unvalidated so nobody mistakes it for a tested path. Choosing neither and letting it sit is how teams discover during an incident that the documented restore procedure references tooling that was retired. ## What this costs Touch-on-use spends incident time. Handoff review spends meeting time. Pruning risks losing something that turns out to matter. State the costs in an interview — the answer that concedes nothing sounds like it came from a blog post. The reason the trade is worth taking is that the alternative cost is paid at 3am by someone following instructions that are wrong.

  • Touch-on-use asks a responder to edit documentation during or right after an incident. How do you keep that from competing with the response itself?
    Mitigation always wins — nobody edits a wiki while users are down. The edit happens once the incident is stable, in the same shift while the detail is fresh, and it is explicitly counted as incident work rather than a follow-up ticket. It also has to be a two-minute change, not a pull request with reviewers, or it will not happen at all.
  • How do you maintain a runbook for something that has never happened, like a regional failover?
    Use cannot maintain it, so either exercise it deliberately on a schedule and treat what breaks as findings, or label it plainly as unvalidated so nobody mistakes it for a tested path. What you must not do is leave it looking authoritative — the failure mode is discovering mid-incident that the documented procedure references retired tooling.
  • Is there any value in a periodic audit of every runbook?
    Limited, and it degrades fast. Audits review what is easy to review, produce a green report, and give false confidence; they also cost real hours. If you run one at all, scope it to the highest-consequence procedures and make it an execution test rather than a read-through — the question is whether the steps still work, not whether the page looks tidy.

saying these in an interview costs you the question

  • Proposes a quarterly documentation review and stops there
  • Keeps every runbook because deleting might lose knowledge
  • Files runbook corrections as backlog tickets after the incident
  • Treats a stale runbook as harmless because it is still just documentation
  • Assumes rarely used runbooks stay correct without exercise

context