skip to content

Runbooks & Playbooks

Written procedures that let any on-caller handle a known failure at 3am. Interviewers ask what a good runbook contains — and expect you to treat runbooks as the staging ground for automation, not an end state.

on this pageshow

questions

6

You are writing a runbook for a failure that pages the on-call engineer at 3am. What sections should it contain, and what does each section have to do for someone who did not build the service?

level: middleimportance: must knowfreq 68%

answer

  1. written for someone who did not build it
  2. each command needs a baseline and a blast radius
  3. how do you know it worked?
  4. permission to stop and escalate
  5. the disconfirming clause at the top

basics

~20 s

A usable runbook states its trigger and the user impact, then gives diagnostics with expected output, remediation with preconditions and blast radius, a verification signal that proves the fix worked, an escalation path with a time-box, and metadata naming the owner and last validation date.

solid answer

~50 s

I write seven parts. **Trigger**: which alert or symptom this applies to, and an explicit "if you see X instead, this is not your runbook" — that stops misapplication. **Impact**: one line on who is affected, so the responder can judge urgency without reverse-engineering it. **Diagnostics**: the exact commands or dashboards plus what a healthy result looks like, because a responder who does not own the service has no baseline. **Remediation**: copy-pasteable steps, each with its preconditions, its blast radius, and whether it is reversible. **Verification**: the specific signal that proves the fix worked — a metric back at baseline, a probe passing — never "looks fine". **Escalation**: who to page and after how long, so a tired person has permission to stop. **Metadata**: owner, last-validated date, and the link back from the alert. The bar is that a competent engineer from another team can execute it correctly at 3am.

go deeper

for a junior

Be able to name the parts — trigger, diagnostics, remediation, verification, escalation — and say who the reader is. Saying that the fix step is not the last step will already put you ahead of most answers.

for a middle

Explain why each section exists rather than listing it: expected output because an outsider has no baseline, blast radius because the responder cannot judge it themselves, verification because a silent alert is not proof of a fix.

for a senior

Show the judgment: the disconfirming clause that stops misapplication, destructive steps gated or escalated instead of documented, and a time-boxed escalation that gives a tired responder permission to stop. Mention how you validate the runbook with someone who did not write it.

for a principal

Own the standard across teams — what a runbook template mandates, how currency is enforced, and where you accept that no runbook should exist because the failure needs a diagnosing human rather than a procedure.

## The design constraint Every choice in a runbook follows from one assumption: the reader is impaired, is not the service's author, and is spending user-visible error budget while they read. That rules out prose, rules out background, and rules out anything the reader has to interpret. The document is an instrument, not an essay. ## Section by section ### 1. Trigger and scope Name the alert or the observable symptom this procedure answers, and — the part most runbooks omit — the disconfirming condition. "Use this when the checkout error-rate alert fires **and** the payment provider's status page is green; if the provider is degraded, go to the provider-outage runbook instead." The disconfirming clause exists because the expensive failure mode of runbooks is not a missing runbook, it is a responder confidently applying the right procedure to the wrong failure. A restart that fixes a wedged worker can extend an outage if the real cause is a corrupt config being reloaded on start. ### 2. Impact One or two lines: who is affected, in what way, and whether the effect is user-visible. The responder uses this to decide how hard to push — whether to mitigate immediately, whether to pull in help, whether this warrants waking anyone else. Without it, they either under-react to something serious or escalate a background job's retry storm as if customers were down. ### 3. Diagnostics with expected output Give the command, query or dashboard, and state what healthy looks like. "`kubectl get pods -n checkout` — expect all pods Running; more than two in CrashLoopBackOff means…" The expected-value part is what makes it usable by an outsider: numbers mean nothing without a baseline, and at 3am nobody is going to go find last week's graph to compare. Keep this section short and ordered by discriminating power. The first check should be the one that most cheaply splits the possible causes, not the one the author happened to run first. ### 4. Remediation with preconditions and blast radius For each step, three things travel with the command: - **Precondition** — what must be true before running it. - **Blast radius** — what it affects. Does it drop in-flight requests? Does it affect one shard or the fleet? - **Reversibility** — can this be undone, and how? Steps that are irreversible or wide-blast (deleting data, failing over a database, flushing a cache the whole fleet depends on) should require an explicit second person or be excluded from the runbook and replaced with "escalate to the service owner". A 3am solo responder should not be able to cause permanent damage by following instructions. ### 5. Verification The most commonly missing section. State the signal that proves the remediation worked: error rate back under the threshold for a sustained window, the synthetic probe passing, the queue draining rather than merely stopping its growth. Without this, responders declare victory on the absence of a new alert, which is not the same thing — alerts can be suppressed, flapping, or slow to clear. Include what to do if verification fails: that path normally leads to escalation, not to trying the same step again. ### 6. Escalation Who, and after how long. A time-box ("if not mitigated within 15 minutes, page the service owner") is the section's real content: it converts "should I bother someone?" — a social question a tired person answers badly — into a rule they can follow. The runbook should also say what information to hand over, so the escalation is not a cold start. ### 7. Metadata Owner, last-validated date, and a link to the alert definition. The date is the decay signal: a responder seeing "last validated 14 months ago" treats the commands with appropriate suspicion. The owner is who fixes it when it turns out to be wrong. ## What to leave out Background theory, architecture diagrams, historical incident narrative, and anything that is really an explanation. Every paragraph the responder must read but cannot act on lengthens the outage and increases the chance they skim past the step that matters. If the context is genuinely needed, link it rather than inline it. ## The test that matters A runbook is finished when an engineer from a neighbouring team, who has never operated the service, can follow it correctly without asking anyone. The only reliable way to know is to have someone like that execute it — during a drill, or by watching what happens the next time it is used for real.

  • Which section do most real runbooks omit, and what goes wrong when it is missing?
    Verification. Without an explicit "this proves it worked" signal, responders treat the absence of a new alert as success — but alerts can be suppressed during the incident, flapping, or slow to clear. The result is a prematurely closed incident that re-pages twenty minutes later, and a second responder who now inherits a muddier picture.
  • How would you handle a remediation step that is destructive or irreversible?
    Either keep it out of the runbook and escalate to the service owner at that point, or gate it explicitly: state the precondition, require a second person to confirm, and describe the recovery path if it goes wrong. A solo responder half-asleep should not be able to cause permanent data loss by following instructions correctly.
  • Why put a time-box on the escalation step rather than leaving it to the responder's judgment?
    Because judgment about whether to wake someone is exactly what degrades under fatigue and social pressure — people tend to keep trying alone for too long. A stated limit ("if not mitigated in 15 minutes, page the owner") converts a social decision into a rule, and it makes late escalation a visible process failure rather than a personal one.

saying these in an interview costs you the question

  • Lists commands with no expected output or baseline
  • Ends at the fix, with no verification signal
  • Includes destructive steps with no precondition or approval
  • Buries the procedure in architecture background
  • No owner or last-validated date, so nobody notices it rotted
  • Assumes the reader knows which alert this runbook answers

context

open as a page

Your team's runbooks have gone stale — commands reference a decommissioned host and several document alerts that no longer exist. How do you make runbook freshness a property of the process rather than a one-off cleanup project?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Tie maintenance to use: whoever follows a runbook during a page repairs it in the same shift. Bind each runbook to the alert that links it so orphans are visible, review them at on-call handoff, stamp owner and last-validated date, and delete more than you write.

open as a page

A team keeps a detailed architecture wiki and argues it does not need runbooks. What does a runbook give an on-caller at 3am that architecture documentation does not?

level: juniorimportance: should knowfreq 52%

basics

~20 s

A runbook is a procedure for one known failure: the symptom it matches, the exact diagnostic and remediation steps, how to verify the fix, and when to escalate. Architecture documentation explains how the system works and leaves the responder to derive the response.

open as a page

A responder followed a runbook's "restart the service" step for a symptom that resembled the documented one but had a different cause, and made the outage worse. How prescriptive should a runbook be, and what has to guard a copy-paste command?

level: middleimportance: should knowfreq 40%

basics

~20 s

Be maximally prescriptive about mechanics and explicit about the conditions under which each step is valid. Every copy-paste command needs a stated precondition, its blast radius, whether it is reversible, and an escape hatch that says stop and escalate when the symptoms do not match.

open as a page

A runbook's remediation has just been fully automated, so the fix now runs without a human. What should happen to the runbook itself?

level: seniorimportance: should knowfreq 34%

basics

~20 s

It changes role rather than disappearing: it now documents what the automation does, how to tell whether it ran, how to disable it, and the manual fallback for when it is off or broken. That fallback is never exercised, so it must be deliberately drilled or plainly marked unvalidated.

open as a page

Leadership mandates that every service must have runbooks covering its top failure modes before launch, and compliance is tracked as a runbook count per service. Would you adopt that, and what would you commit to instead?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

No — counting documents measures writing, not response capability, and produces procedures written from imagination that nobody has executed. Commit instead to coverage defined against the alerts that actually page, at least one execution by someone who did not write it, and a documented escalate-to-owner entry as a legitimate answer.

open as a page