skip to content

Why do experienced teams track on-call interrupt hours rather than page counts alone, and what does a single overnight page actually cost the next day?

level: seniorimportance: should knowfreq 45%

answer

  1. events are not equal in cost
  2. wall-clock, not incident count
  3. re-focus time after every interruption
  4. overnight pages budgeted separately
  5. cost model settles page-versus-ticket

basics

~20 s

Page counts treat a two-minute acknowledgement and a four-hour incident as the same event. Interrupt hours capture triage, recovery and the lost re-focus time around each page, and an overnight page costs the incident plus a degraded next day, which is why some conditions should be tickets rather than pages.

solid answer

~50 s

A page count is a count of *events*, and events are wildly unequal. Two pages that each consume four hours and end at 04:00 are not a lighter shift than six that auto-resolved in a minute. Interrupt hours measure the thing that actually competes with engineering work: time from page to stand-down, plus write-up, plus the re-focus cost afterwards — commonly cited research on interrupted work puts task resumption in the tens of minutes, so even a trivial page is never trivially cheap. An overnight page costs its own duration and then a degraded following day, which is why healthy teams treat comp time as the norm rather than a favour. The practical payoff is a routing rule: if the condition can wait until morning without user-visible harm, it is a ticket. Costing the interrupt honestly is what makes that argument win against "but I want to know about it".

go deeper

for a junior

Know that a page costs more than the minutes spent handling it, and that the honest thing to record after a shift is the total time it took you, including writing it up.

for a middle

Explain the components of interrupt cost — response, follow-up, re-focus — and why counting events alone lets a rotation look healthy while the team ships nothing.

for a senior

Use the cost model to settle real routing arguments: bring an alert's fire rate, overnight share and action rate to the discussion rather than debating preferences, and defend comp time for overnight pages as policy.

for a principal

Own the capacity link: show how interrupt hours spilling past one shift's absorption silently taxes the whole team, and how you would cap operational time by policy so engineering throughput is protected structurally.

## The problem with counting pages Pager load is usually first measured as pages per shift, and that is the right starting point. But a page count is a count of discrete events, and it silently assumes every event is equivalent. Real shifts violate that assumption badly: - A self-clearing alert acknowledged from a phone in ninety seconds. - A three-hour incident with a war room, a mitigation, and a write-up. - A page at 03:10 that lasted twenty minutes but ended the night's sleep. Counting these as "three pages" produces a number that is technically true and operationally misleading. It is also gameable in the wrong direction: a team can look busy while doing nothing, or look quiet while being destroyed. ## What interrupt hours capture Interrupt hours measure the wall-clock time on-call actually consumed. The components worth including: - **Response time** — page to stand-down, including investigation that led nowhere. - **Follow-up** — the write-up, the ticket, the handoff note. Real work, often invisible. - **Re-focus cost** — the period after an interruption before deep work resumes. This is not folklore; frequently cited research on interrupted work measures task resumption in the region of twenty-plus minutes, and engineering work is close to the worst case for it. The consequence is arithmetic that surprises people. Five "cheap" daytime pages, each five minutes of actual handling, is not twenty-five minutes gone; with re-focus it can consume the productive part of a day. That is why a rotation can report a comfortable page count and still deliver nothing. ## The overnight multiplier A page that wakes someone has a cost profile of its own. You pay the incident itself, then a next day at reduced capability, and then a compounding effect if it happens on consecutive nights. There is no precise universal multiplier and you should not invent one in an interview; what you should say is that overnight pages are tracked and budgeted *separately* from daytime pages, and that the norm which follows is comp time — a late start or a day off after a bad night — treated as automatic rather than as something an engineer has to ask a manager for. The second-order effect matters more than the first: a team that pays for overnight pages properly has a real incentive to stop generating them. A team that treats them as free is quietly financing its noisy pager with unpaid human cost, which shows up later as attrition and as people declining to join the rotation. ## The decision this metric drives All of this exists to settle one recurring argument: should this condition page a human now, or become a ticket for the morning? The instinct of the person who wrote the alert is almost always "page me, I want to know". Interrupt cost is the counter-argument, and it is quantitative: this alert fires about four times a week, roughly half of them overnight, and it has never required action before business hours. That framing converts a preference contest into a cost comparison. The rule that falls out is simple: **page only when a human must act within minutes to prevent or limit user-visible harm.** Everything else — capacity trends, degraded redundancy that still has headroom, batch failures with a retry window — is a ticket. The cost model is what gives you the standing to enforce it. ## Where this fits the wider budget Interrupt hours also connect on-call to team capacity planning. If the on-call engineer's interrupt hours routinely exceed what one shift can absorb, the overflow does not vanish — it spills onto teammates, who are then also interrupted, and the team's project throughput collapses in a way that looks mysterious on a burndown chart. Google's SRE model handles this by capping the proportion of time spent on operational work outright, so that engineering capacity is protected by policy rather than by hope. Whatever cap you choose, the value of measuring interrupt hours is that you can tell when you have blown through it. ## Anti-patterns Counting only pages that were acknowledged misses the ones that auto-resolved but still woke someone. Excluding the write-up understates the real cost by a large margin. And averaging interrupt hours across the whole team rather than per shift returns to the original sin: hiding one person's terrible week inside everyone else's quiet one.

  • Five daytime pages a day, each handled in under five minutes — is that a healthy rotation?
    No. The handling time is trivial but the interruption is not: with re-focus overhead, five scattered interruptions can consume the deep-work portion of a day. The engineer will report that on-call "wasn't bad" while delivering nothing that week. Judge it on interrupt hours and on what project work actually shipped, not on how quickly each page was closed.
  • How would you use interrupt cost to argue against an alert whose owner insists on being paged?
    Bring its numbers: how often it fires, what share is overnight, and how many times a responder took action before business hours. If that last figure is near zero, the alert is buying nothing and costing measurable hours. Offer the ticket path plus a user-impact alert as the paging route, and record who accepts the slower detection.

saying these in an interview costs you the question

  • Treating a two-minute page and a four-hour incident identically
  • Counting only handling time, ignoring re-focus cost
  • Assuming a quick overnight page costs nothing next day
  • Averaging interrupt hours across the team, not per shift
  • Excluding write-up time from the cost of a page

context