Leadership mandates that every service must have runbooks covering its top failure modes before launch, and compliance is tracked as a runbook count per service. Would you adopt that, and what would you commit to instead?
answer
- counting documents measures writing, not capability
- the denominator should come from real pages
- validated by someone who did not write it
- an honest blank beats a confident wrong page
- gate hard only where an unguided response is catastrophic
basics
~20 sNo — counting documents measures writing, not response capability, and produces procedures written from imagination that nobody has executed. Commit instead to coverage defined against the alerts that actually page, at least one execution by someone who did not write it, and a documented escalate-to-owner entry as a legitimate answer.
solid answer
~50 sI would push back on the metric, not the intent. A count rewards volume, so you get runbooks written from imagination by the person who built the service, never executed, stale by the time they are needed — and false confidence is more expensive than an honest blank, because a wrong procedure is followed while a missing one triggers escalation. What I would commit to instead: define coverage against something real, namely the alerts that page and the failure modes postmortems actually produced, rather than a brainstormed top-five. Require validation — each runbook exercised once by an engineer who did not write it, in a drill or by following it through a real page. Measure use and outcome, not existence: was it opened, was it followed, did it work. And allow "no runbook, escalate to owner X" as a documented, legitimate entry. I would keep a hard launch gate only where an unguided response is catastrophic, and accept the delay there.
go deeper
Know that a runbook is only useful if it works for the person following it, so a document written but never tried is not evidence that the team can handle the failure.
Be able to explain why an unexecuted runbook has an unknown defect rate, and why the person who built the service is the worst judge of whether their own procedure is followable by someone else.
Argue the concrete alternative: coverage measured against alerts that page, validation by a non-author, and use-and-outcome data from real incidents rather than existence counts.
Own the negotiation. Agree with the intent, name the specific way the proposed metric fails, offer a countable replacement leadership can report, and be explicit about what your alternative costs in drill time and owner load.
## What the mandate gets right The intent is sound. Services do launch with no operational documentation, the team that built them absorbs every page personally, and the knowledge stays in one head until that person leaves or goes on holiday. Wanting a floor is reasonable. ## Why the metric breaks it Counting artifacts measures the act of writing. Everything that follows comes from that: - **Runbooks written from imagination.** "Top failure modes" before launch means a brainstormed list. The failures a service actually has are usually not on it — they involve a dependency behaving in a way nobody modelled, or a saturation point nobody predicted. So the documents cover hypotheticals and miss the real thing. - **Written by the wrong person for the wrong reader.** The service's author writes them, and the author cannot see their own assumed context. Steps read fine to them and are unusable by an on-caller from another team, which is the reader that matters. - **Never executed, therefore never falsified.** A runbook nobody has run has an unknown defect rate. It looks identical to a good one. - **Stale on a schedule.** Volume and currency trade off directly. Four hundred documents cannot be kept current by a team that can genuinely maintain forty. - **False confidence is the expensive part.** A missing runbook makes a responder cautious and makes them escalate. A wrong one authorises confident action. The metric optimises away the honest blank and replaces it with the dangerous document. This is a specific instance of a general failure: measuring the artifact rather than the capability the artifact is supposed to produce. ## What I would commit to instead **Define coverage against reality.** The denominator should not be a brainstorm. Two good sources: the alerts that actually page a human — every one of them should either link a runbook or be a candidate for deletion — and the failure modes that postmortems have already produced. Both lists are grounded in things that happened. **Require validation, not authorship.** A runbook counts once someone who did not write it has followed it end to end: in a scheduled exercise, or by being the responder on a real page and confirming it worked. This is the single change that converts the metric from writing to capability, and it is the one leadership tends to accept because it is still countable. **Measure use and outcome.** Was the runbook opened during the incident? Did the responder follow it? Did it work, or did they have to deviate? That data is a far better health signal than a document count, and it points at exactly which runbooks need repair. **Make the honest blank legal.** "No runbook for this failure mode — page the service owner" should be an acceptable, documented answer. Some failures genuinely require a diagnosing human rather than a procedure, and forcing a document into that slot produces theatre. Naming the gap explicitly is more useful than filling it badly. **Keep a hard gate only where it is earned.** For a service where an unguided 3am response risks data loss, a regulatory breach, or a safety consequence, a validated procedure before launch is a reasonable blocker and the launch delay is worth paying. Applying the same gate to a low-traffic internal tool buys nothing and teaches teams that the gate is a formality — which is how the gate stops working for the cases that mattered. ## The costs of my alternative — say them out loud Validation is expensive: someone's time to run a drill on every documented failure mode, and that time competes with delivery. Alert-derived coverage misses failures that nothing alerts on yet, so it is a floor rather than a guarantee. Allowing escalate-to-owner as an answer concentrates load on the owners, which is a real burden on a small team. And any policy with a defined bar will eventually be gamed by someone under launch pressure. An interview answer that concedes none of this sounds rehearsed. The honest position is that these costs are smaller than the cost of an organisation that believes it is prepared because it has documents. ## How I would land it with leadership Not as a refusal. Agree with the goal, name the specific failure mode of the proposed metric, and offer a replacement that is still measurable — validated-runbook coverage against paging alerts, plus use-and-outcome data from incidents. Leadership usually wants an assurance they can report; the negotiation is about what the number counts, not about whether there is one.
- Leadership still wants a single number to report. What number would you give them?Validated runbook coverage over paging alerts: of the alerts that wake a human, what share have a procedure that someone other than its author has executed end to end within the last N months. It is countable, it moves for the right reasons, and it cannot be inflated by writing more documents — which is exactly why it will be argued about.
- Is there any case where you would block a launch on runbook coverage?Yes, where an unguided response could be catastrophic — data loss with no restore path, a safety or regulatory consequence, or a failure whose wrong mitigation is irreversible. There the launch delay is cheaper than the exposure. Applying the same gate everywhere is what turns it into a formality and hollows it out for the cases that genuinely needed it.
- What is the downside of accepting "escalate to the owner" as a documented answer?It concentrates load on the people who already carry the most context, which on a small team is a burnout path and a single point of failure. It is acceptable as an explicit, visible entry — because it is honest and it makes the concentration measurable — but a growing list of them is a signal to invest in either documentation or a broader rotation, not a stable end state.
saying these in an interview costs you the question
- Accepts a document count as evidence of readiness
- Treats runbook coverage as a compliance exercise
- Assumes a runbook written by the author works for an outsider
- Applies the same launch gate to every service regardless of stakes
- Prefers a speculative runbook to an honest documented gap