Your team's on-call rotation is taking roughly 50 pages a week and two engineers have asked to come off it. What do you do in the first two weeks, and what do you change structurally?
answer
- data before opinions
- rank alerts by pages produced
- every alert gets a disposition
- silences need owner and expiry
- cap needs a pre-agreed consequence
basics
~20 sStart with data, not sympathy: rank every alert by how many pages it produced last month, then give each of the top offenders a disposition — delete, demote to a ticket, or fund the reliability fix — with a named owner and a date. Structurally, agree a pager-load cap with a pre-committed consequence.
solid answer
~50 sFirst two weeks are triage. Pull 30 days of pages, group them by the alert that fired, and sort by count — in most unhealthy rotations a handful of rules account for the bulk of the volume. For each of the top ones I force a decision with a cost attached: delete it, demote it to a ticket that gets looked at in working hours, or keep it paging and fund the fix that stops it firing. Silences are allowed but only time-boxed with a named owner, otherwise they become permanent blindness. I also make the current on-call's *job* during their shift the cleanup itself, not project work, so the effort is staffed rather than volunteered. Structurally, three things change: a pager-load cap with an agreed consequence when it is breached, a standing page review at handoff, and a rule that any new paging alert ships with an owner and a runbook. And I record who accepted the detection risk for anything we stopped paging on.
go deeper
Be ready to say that the first move is counting which alerts fired and how often, and that a page requiring no action is a bug to be filed rather than a normal part of the job.
Explain the disposition set — delete, demote to ticket, fix the cause, change the condition — and why a silence without an expiry and an owner is just a hidden deletion.
Demonstrate that you can staff the cleanup: say explicitly what the on-call engineer stops delivering, what slips on the roadmap, and who accepted the detection risk for each alert you removed.
Own the pre-commitment. Show how you would get a pager-load cap with an automatic consequence agreed across teams before a crisis, and how you prevent the metric from being gamed by pairing it with customer-reported detection.
## Why this question is asked Fifty pages a week is roughly seven a day, every day, and it is the standard interview scenario for on-call health because almost everyone has an opinion and most of those opinions are useless. "Improve the culture", "reduce noise" and "tune the thresholds" are not answers. The interviewer is listening for a sequence with decisions and costs in it, and for evidence that you know who pays for the cleanup. ## Week one: get the ranked list Nothing is decided before the data exists. Export the last 30 days of pages and group them by the alert that fired. The near-universal finding is a steep head: a small number of rules produce most of the volume. That ranking is the entire prioritisation exercise, and it is why "we'll review all our alerts" is a worse plan than "we'll review these six". For each alert on the head of the list, record three facts: how many pages, how many were actionable, and what the responder actually did. "Acknowledged and it recovered on its own" appearing repeatedly is a self-clearing condition that should never have paged. ## The disposition, which is a forced choice Every alert on the list gets exactly one of four outcomes, and each has a cost you must be willing to name: - **Delete.** Cheapest, and it means accepting that this condition will now be detected some other way or not at all. Record who accepted that risk. - **Demote to a ticket.** The condition still matters but does not need a human within minutes. Cost: detection-to-response time grows to hours, and the ticket queue must actually be worked or you have simply hidden the problem. - **Keep paging, fix the cause.** The alert is right and the service is wrong. Cost: real engineering time, which must come out of the roadmap. - **Keep paging, change the condition.** Same symptom, better trigger — usually meaning it should fire on user-visible impact rather than on an internal cause. Cost: engineering time plus a window where you may under-detect. A silence is not a disposition. Silences are legitimate as a stopgap, but only with an expiry and a named owner, because an indefinite silence is an alert you have deleted without admitting it. ## Staffing the cleanup This is where most answers fall apart. Alert cleanup done "when we get a moment" never happens, because the team is already interrupted 50 times a week. Two mechanisms work: - **Make it the on-call's job.** During their shift the on-call engineer is not expected to deliver project work; their deliverable is the previous shift's cleanup. This is honest about capacity — that person was never going to finish a feature anyway. - **Fund it as a project.** If the backlog is bigger than shift capacity, it becomes a sprint with an owner, not background work. That means something on the roadmap slips, and saying so explicitly is the point. ## Structural changes so it does not recur 1. **A pager-load cap with a pre-agreed consequence.** A number alone changes nothing; the cap must specify what automatically happens on breach — cleanup takes precedence over feature work, or the pager routes back to the team that shipped the change. Agree it while things are calm, because arguing for it mid-quarter always loses. 2. **A page review at every handoff.** Ten minutes, every page from the shift walked, each one given a disposition. This is what keeps the list from regrowing. 3. **Admission criteria for new paging alerts.** A new alert that pages a human must name an owner, link a runbook, and state what the responder is expected to do. If nobody can write that sentence, it is a ticket or a dashboard. 4. **A counterweight metric.** Track incidents first reported by customers or by SLO violation. Pager load falling while customer-reported incidents rise means you optimised the wrong number. ## What to say about the two engineers Do not skip the human part, but keep it concrete: take them off the rotation temporarily if the schedule allows, honour comp time already owed, and be explicit that the cleanup is funded rather than promised. The credibility of everything above depends on the first visible win landing within days — which is exactly why you start with the top of the ranked list rather than a comprehensive audit.
- Which do you tackle first: the alert that pages 40 times a week and is never actionable, or the one that pages twice and is always real?The noisy one, immediately. It is consuming most of the attention while contributing nothing, and removing it costs almost no engineering time. The always-real alert is doing its job — its pages point at a reliability defect, which is a separate and slower piece of work. Fixing noise first also restores enough capacity to do the reliability work at all.
- That noisy alert did catch one genuine incident six months ago. Do you still remove it?Usually yes, but explicitly. Convert it to a ticket or a dashboard panel, and rely on a user-impact alert as the paging path for that failure mode instead. Then record the trade: detection for this case may now be minutes slower, and name who accepted that. The alternative — keeping it — costs roughly a thousand interruptions per genuine catch.
- How do you stop the alert list from regrowing six months later?Admission criteria plus a standing review. Any new paging alert must name an owner, link a runbook and state the action expected of the responder; anything failing that becomes a ticket. Then every handoff includes a short walk of the shift's pages where each gets a disposition. Regrowth happens when adding an alert is free and removing one needs a meeting.
saying these in an interview costs you the question
- Proposing to add people to the rotation instead of removing pages
- Silencing noisy alerts with no owner or expiry
- Planning a full alert audit before any quick win
- Treating cleanup as background work nobody is funded for
- Calling it a culture problem and stopping there