A team keeps a detailed architecture wiki and argues it does not need runbooks. What does a runbook give an on-caller at 3am that architecture documentation does not?
answer
- reference material versus a procedure
- who is reading it, and when
- scoped to one known failure
- trigger, steps, verification, escalation
basics
~20 sA runbook is a procedure for one known failure: the symptom it matches, the exact diagnostic and remediation steps, how to verify the fix, and when to escalate. Architecture documentation explains how the system works and leaves the responder to derive the response.
solid answer
~50 sArchitecture docs are reference material — they describe components, data flows and design decisions, and they are read at leisure by someone trying to understand the system. A runbook is an operational procedure scoped to one known failure: it names the symptom or alert that triggers it, gives the diagnostic commands and what their output should look like, the remediation with its preconditions, a verification step that proves the fix worked, and an escalation path if it did not. The difference matters because at 3am the responder is sleep-deprived, may not have built the service, and is spending user-visible error budget every minute they spend reasoning from first principles. A runbook converts one person's hard-won knowledge into something a generalist can execute under stress. Both artifacts are worth having — but reading the architecture wiki during an outage is a symptom that the runbook is missing.
go deeper
Be able to say plainly that a runbook is a step-by-step procedure for one known failure, aimed at whoever is on-call, while architecture docs explain how the system works. Mentioning verification and escalation as parts of it will stand out.
Explain what makes the procedure usable under stress: a trigger that confirms you are in the right document, expected output for each diagnostic, and reversibility noted on each remediation step. Be ready to say why length is a defect here.
Show you have followed one during a real page. Talk about runbooks being linked from the alert, about the responder often not owning the service, and about a stale runbook being more dangerous than a missing one because it is executed with confidence.
Own the tradeoff between coverage and currency across many teams: which failure modes justify a documented procedure at all, what runbooks cost in maintenance, and how you keep responders capable of diagnosing the novel incidents that no procedure covers.
## The two artifacts answer different questions Architecture documentation answers *how does this system work?* It describes services and their dependencies, data flow, storage choices, the reasoning behind a design, and the failure domains that exist. It is written for someone with time: a new joiner, a designer of the next feature, a reviewer weighing a change. Its natural reading mode is exploratory. A runbook (many organisations use *playbook* interchangeably) answers a much narrower question: *this specific thing is happening right now — what do I do?* It is scoped to one known failure mode, usually one that has already happened at least once or that an alert explicitly watches for. Its natural reading mode is execution: read a line, do it, read the next line. ## Why the distinction shows up hardest on-call Three conditions make an outage a bad time to read reference material: 1. **The responder is impaired.** They were asleep ten minutes ago. Working memory is poor, and the cost of a wrong keystroke is high. 2. **The responder may not own the service.** Modern rotations frequently page a team that operates many services, or a platform on-caller covering a wide surface. They have never seen this component's internals. 3. **The clock is running against users.** Every minute of reasoning-from-first-principles is measurable in failed requests. Reducing time-to-mitigate is the whole point. Under those conditions, the useful artifact is not an explanation but a procedure. ## What a runbook adds concretely - **A trigger.** "This runbook applies when the checkout latency alert fires and the upstream payment provider dashboard is green." The responder can confirm they are in the right document before touching anything — and the trigger is also what stops them from applying it to a superficially similar failure. - **Diagnostics with expected output.** Not just "check the queue depth" but the command and what a healthy versus unhealthy result looks like. Someone who has never seen this system does not know what normal is. - **Remediation with preconditions and blast radius.** The exact commands, what they will affect, whether they are reversible, and what must be true before running them. - **Verification.** The specific signal that proves the fix worked — a metric returning to baseline, a synthetic probe passing — rather than "looks OK now". - **Escalation.** Who to call, and when to stop trying. A runbook that ends with "if this did not work within ten minutes, page the service owner" is doing real work: it gives a tired person permission to stop. ## What a runbook is not It is not a replacement for understanding. A team that only has runbooks builds responders who can execute but cannot diagnose anything novel, and novel is what real incidents mostly are. It is also not a dumping ground: a runbook that grows into a system tour stops being executable, and long documents are skimmed under stress, which is how the wrong section gets applied. The healthy relationship is that the architecture doc explains the machine, the runbook handles the failures you already know about, and anything the runbook cannot handle escalates to a human who understands the machine. ## The honest counterargument The team in the question is not entirely wrong about one thing: runbooks cost maintenance, and a wrong runbook is worse than no runbook, because it is followed with confidence. That is an argument for keeping few runbooks and keeping them current — attached to the alerts that actually page, repaired by whoever uses them — not an argument for writing none and asking a half-awake engineer to read a design document instead. ## Saying it in an interview The answer an interviewer is listening for is the shift in audience and mode: reference material for someone with time and context, procedure for someone with neither. If you can add that a runbook is scoped to a trigger and ends in verification and escalation, you have shown you have actually followed one at 3am rather than only read about them.
- If a runbook is so valuable during an outage, why not write one for every conceivable failure?Because coverage is not free and unread runbooks decay. Most real incidents are novel, so speculative runbooks written from imagination are usually wrong by the time they are needed — and a wrong procedure is followed with confidence. Write them for the failures that actually page or that postmortems surfaced, and keep the set small enough to stay current.
- Where should a runbook physically live so an on-caller can find it at 3am?Linked directly from the alert that triggers it, so it arrives in the notification rather than needing a search. Beyond that: somewhere available when your own systems are down — not behind the service that is broken, and reachable from a phone. Teams also keep a short offline copy of the handful of procedures needed when the primary tooling is unavailable.
saying these in an interview costs you the question
- Says good architecture docs make runbooks unnecessary
- Treats a runbook as a full system tour
- Writes remediation steps without any verification step
- Assumes the responder built and understands the service
- Leaves out escalation, so the responder keeps trying alone