At 03:00 you are paged by an alert whose entire body reads "HighErrorRate firing — checkout-service". What should that paging alert's payload have carried instead, and why is a runbook link treated as a mandatory field rather than a nicety?
answer
- read half-awake, on a phone
- impact in user terms, not rule names
- who owns it, where to go next
- link is a forcing function at authoring time
- a link that restates the alert fails
basics
~20 sA page should state the user-visible impact, how bad it is, the owning service and team, and link a runbook that gives the first diagnostic step and the available mitigations. Without that link every responder re-derives the response under time pressure.
solid answer
~50 sThe body I actually want tells me four things in one screen: what is broken in user terms rather than metric terms, how bad it is and since when, which service and team own it, and where to go next. "Where to go next" is the runbook link, plus a dashboard and the recent-deploy view. I treat the link as mandatory because the person reading a page is frequently the least-context member of the rotation at the worst possible hour; if the response only exists in the head of whoever wrote the rule, the response time depends on who happens to be on call. So the review rule is simple: no new paging alert ships without a runbook link in its annotations, and a link that just restates the alert or points at a wiki search page fails that review the same as an empty one.
go deeper
Know the fields a page must carry — user-visible impact, how bad and since when, the owning service and team, and a runbook link — and be able to criticise a page that carries only a rule name.
Explain the runbook link as a forcing function: if nobody can write down the response at authoring time, the rule has not proved a human can act and probably belongs at a lower tier.
Describe how you enforce it — a required annotation validated in review or in the pipeline — and how you keep linked runbooks from going stale by capturing failures during real incidents.
Set the standard fleet-wide: which annotation fields are mandatory, who grants exceptions and for how long, and how you measure whether the payload is actually shortening time to first useful action.
## What the responder needs in the first sixty seconds A page is read half-awake, on a phone, by whoever is on the rotation — which may be the newest person on the team, or someone who has never touched this service. The payload has one job: get that person from "my phone buzzed" to "I know what is happening and what to do first" without them having to reconstruct anything. That means four things. **Impact in user terms.** "HighErrorRate" is the name of a rule, not a statement about the world. "About 7% of checkout attempts are failing" is a statement about the world, and it lets the responder decide in seconds whether this is a real incident. **Magnitude and duration.** How bad, and since when. A metric that has been degraded for forty minutes and one that broke ninety seconds ago call for different first moves, and the difference between them is also the first clue about cause. **Ownership.** The service and the team that owns it, carried on the alert itself. Even when routing is correct, the responder often needs to know who else to pull in, and a page that does not name its service is unreadable in a channel where several services alert. **Where to go next.** The runbook link, a dashboard link scoped to the affected service and time window, and a link to recent deploys or config changes. These are the three places a responder goes anyway; putting them in the payload removes several minutes of navigation from every single incident. ## Why the runbook link is mandatory The argument against it is always the same: "anyone on this team knows what to do". That is usually true of the person saying it and false of the rotation as a whole. The runbook link converts response quality from a property of the individual on call into a property of the team. There is a second, sharper reason. Requiring a link is a forcing function at authoring time. If you cannot write down what a responder should do when this fires, you have not established that a human can do anything about it — which is the actionability test, applied before the alert ever reaches production. A rule whose author cannot fill in the runbook field is a rule that probably belongs in a lower tier. Rejecting it at review is far cheaper than discovering it at 03:00. Mechanically, this is enforced as a required annotation on the alert definition — most alerting systems carry an arbitrary set of annotation fields alongside a rule, and a `runbook_url`-style field is the conventional home for it. Whether the check runs in code review or as a validation step in the pipeline matters less than that it exists and blocks. ## What makes a link worthless Not every link satisfies the requirement, and the review has to be willing to say so. - **A link that restates the alert.** A page saying error rate is high, linking to a document saying error rate being high means many requests are failing, has added a click and nothing else. - **A link to a search page or a wiki homepage.** The responder now has a research task on top of an incident. Link the specific section for this specific alert. - **A runbook with no mitigations.** Diagnosis without action leaves the responder better informed and equally stuck. The document must name what levers exist, even if the answer is "escalate to the database team, here is how". - **A stale runbook.** Steps referencing a service that was decommissioned actively mislead, and a misleading runbook is worse than none because it consumes the responder's trust as well as their time. The last one is the hard one, and the practical defence is to keep the runbook next to the alert definition so that changing one prompts a review of the other, and to note during incidents when the runbook was wrong — that note is one of the cheapest and most reliably useful postmortem outputs. ## The payload is not a place for everything There is an opposite failure: a page carrying twenty labels, the full query expression, and a paragraph of templated prose. On a phone screen at 03:00 that is as unreadable as an empty body. The rule of thumb is one screen: impact, magnitude, owner, and the two or three links. The detail belongs behind those links, where the responder can reach it once they are at a keyboard. ## What good looks like "[SEV-2] checkout-service: 7.2% of checkout attempts failing for the last 4 minutes (normal < 0.3%). Owner: payments-team. Runbook: <link to the checkout-error-rate section>. Dashboard: <link, scoped to the last hour>. Deploys: <link to the last 24 hours>." A responder who has never seen this service can start working from that alone, which is the entire point.
- How do you stop runbooks linked from alerts going stale?Keep the runbook next to the alert definition so changing the rule prompts a review of the document, and make "the runbook was wrong" an explicitly captured output of every incident that used one. Both are cheap. What does not work is a periodic review of all runbooks in the abstract: nobody can evaluate a document they have not just used under pressure.
- Is there any case where a paging alert legitimately has no runbook?Rarely, and it should be visible rather than quiet. A brand-new service may page before its runbook exists, but that should be a recorded exception with an owner and a date, not the default. The usual case labelled "no runbook possible" is really an alert with no defined response, which means it should not be at paging tier yet.
- Why not put the full diagnostic detail into the page body instead of behind a link?Because the page is read on a phone. A body carrying every label and the full rule expression is as unusable as an empty one, and the important fields get lost in it. One screen — impact, magnitude, owner, two or three links — is the target; depth belongs behind the links, where the responder arrives once they are at a keyboard.
saying these in an interview costs you the question
- Alert body is only the rule name and service
- Assuming everyone on the rotation already knows the response
- Linking a wiki homepage instead of the exact section
- Treating the runbook link as documentation polish
- Cramming every label into the page body