How do you measure whether an on-call rotation is healthy, and what pager-load numbers would tell you it is not?
answer
- count per shift, not per month
- actionable rate, not raw volume
- distribution hides the worst shift
- two incidents per 12-hour shift
- rank alerts by pages generated
basics
~20 sMeasure pager load per shift, not per month: pages per shift, the share that arrive outside business hours, the share that turned out to be actionable, and hours lost to interrupts. Google's SRE book suggests at most two incidents per 12-hour shift.
solid answer
~50 sThe unit of measurement is the **shift**, because that is the unit a human actually suffers. I track four things: pages per shift, what fraction arrived outside working hours (and how many woke someone), what fraction was actionable — meaning the responder did something a machine could not have done — and total interrupt hours consumed. Google's SRE book uses a guideline of no more than two incidents per 12-hour shift, on the argument that anything more leaves no time to investigate properly and write up what happened. Then I look at the distribution rather than the mean: an average of 1.5 pages per shift can hide one weekend that took 20. Finally I rank alerts by how many pages each one generated, because in practice a handful of rules usually produce most of the volume, and that ranked list is the cleanup backlog.
go deeper
Know that on-call health is measured in pages per shift and that a page which required no action is a defect in the alert, not bad luck. Be able to say what you would record after your own shift.
Explain why the shift is the right denominator, name the four core metrics — volume, out-of-hours share, actionable rate, interrupt hours — and quote a reference figure such as two incidents per 12-hour shift with the reasoning behind it.
Show that you read distributions rather than means, that you rank alerts by pages generated before proposing fixes, and that you pair pager load with a detection-quality metric so the number cannot be gamed by silencing.
Own the argument that pager load is a staffing and funding input, not a dashboard: be ready to say what headcount, ownership or roadmap decision changes when the trend climbs, and how you would set the cap for a team of six rather than copying Google's numbers.
## What pager load actually measures Pager load is the volume of interrupt work a single on-call shift absorbs. It is deliberately a *human* metric, not a system metric: it does not say how reliable the service is, it says how much of a person's attention and sleep the service consumes. A service can be perfectly within its reliability targets and still have a brutal pager, and vice versa. That is why pager load is measured and governed separately from availability. The key modelling decision is the denominator. Pages per month is nearly useless — it averages a catastrophic weekend into three quiet weeks and reports a comfortable number. Pages per shift is the honest unit, because the shift is what one person carried. ## The metrics worth collecting - **Pages per shift.** Reported as a distribution — median, p90, and the worst shift — not a mean. - **Out-of-hours share.** What fraction fired outside working hours, and separately how many actually woke someone. A page at 03:00 is not the same event as a page at 14:00 and should never be counted as though it were. - **Actionable rate.** What fraction of pages required a human to do something that could not have been automated or deferred. This is the single strongest signal of alert quality: a rotation with 6 pages a shift at 90% actionable has a reliability problem; the same volume at 10% actionable has an alerting problem, and the fixes are completely different. - **Interrupt hours.** Wall-clock time consumed from page to stand-down, including write-up. Two pages that each take three hours are not a light shift. - **Repeat rate.** How many pages were the same condition firing again. Repeats mean the last responder mitigated without fixing, or the alert has no owner. - **Detection quality, as a counterweight.** How many incidents were reported by a customer or another team before the pager fired. Without this, a team can "improve" pager load simply by turning alerts off. ## The reference numbers Google's SRE book offers two figures that get cited constantly in interviews. First, a target of **at most two incidents per 12-hour on-call shift** — the reasoning is not comfort, it is quality: each incident needs investigation, mitigation and a proper write-up, and beyond roughly two the responder starts merely acknowledging and moving on. Second, SREs are expected to spend **no more than 25% of their time on-call**, with a further ceiling on other operational work, so that engineering capacity survives. These are Google's guidelines from their own staffing model, not industry law; the useful part in an interview is that you can state a number, justify why it exists, and say what you would do differently at a smaller company where six people carry everything. ## Reading the numbers Always look at the shape, not the average. Three patterns recur: - **Concentration in one shift.** The mean is fine; one rotation slot (often the one covering a batch job or a weekly peak) is unlivable. Fix the schedule or the job, not the average. - **Concentration in one alert.** Ranked by page count, the top few rules usually dominate the volume. This is the cheapest possible win and the reason the ranked list is worth producing before any other analysis. - **Low volume, high interrupt hours.** Few pages, each an all-night incident. That is a reliability and runbook problem, not a noise problem. ## Closing the loop Metrics that nobody reads change nothing. The working pattern is a short standing review — commonly at shift handoff — where the outgoing on-call walks the shift's pages and each one gets a disposition: keep as a page, demote to a ticket, delete, or file a reliability fix with a named owner. The pager-load numbers then feed the team's planning: if the trend is climbing, cleanup takes priority over roadmap work rather than being done in the gaps. ## Anti-patterns Counting only mean time to acknowledge tells you the pager is being answered, not whether it should have fired. Counting alert *rules* rather than pages fired confuses configuration with experience. And treating every page as legitimate because someone once wrote the rule is how rotations get to 50 pages a week: the default assumption should be that an alert must justify continuing to wake people, not that it is entitled to.
- Your average is 1.5 pages per shift, yet everyone says on-call is brutal. What are you missing?The distribution and the clock. Averages hide a p90 shift that took twenty pages, a slot that always covers the weekly batch window, and the fact that overnight pages cost far more than daytime ones. Look at the worst shift, the out-of-hours share, and total interrupt hours rather than the mean, and rank pages by which alert produced them.
- If you could publish only one on-call number to the whole team, which would it be?The percentage of shifts that exceeded the agreed cap. It is a rate people feel — "one shift in three was over the line" — it cannot be flattered by averaging, and it maps directly onto an action, because every breach is a specific shift whose pages can be reviewed. I would pair it with the actionable rate so nobody hits the cap by silencing alerts.
- How would you stop a team from improving pager load simply by deleting alerts?Track detection quality alongside it: how many incidents were noticed by a customer, another team, or an SLO violation before the pager fired. Pager load falling while customer-reported incidents rise is a regression, not a win. Every alert removed should also be recorded with the decision and who accepted the detection risk.
saying these in an interview costs you the question
- Reporting pages per month, which hides the worst shift
- Assuming every alert that fires was legitimate
- Tracking only acknowledge time, never page volume
- Counting a 3am page the same as a midday one
- Judging the rotation by the average rather than the tail