skip to content

You join a team whose pager delivers roughly 50 alerts a week, and most are acknowledged with no action taken. Describe the process you would run over the next month to reduce that, and how you would decide the fate of each rule.

level: seniorimportance: must knowfreq 66%

answer

  1. measure before you tune
  2. volume is concentrated in a few rules
  3. did a human do anything?
  4. demote before you delete
  5. hang the review on the handoff

basics

~20 s

Measure before changing: per rule, count firings over the last quarter and what share led to a human action. Then apply a disposition ladder — retire, demote off the paging path, retune thresholds and durations, group or suppress the fan-out, or fix the underlying instability — and make the review recurring at every handoff.

solid answer

~1 min

I would start with data, not opinions. Export every alert firing for the last 60 to 90 days and build a per-rule table: how many times it fired, how many pages that produced, median time to acknowledge, and — the column that decides everything — how often a human did something as a result. That last figure usually has to be captured going forward, by asking the on-call to tag each page as actionable or not at acknowledgement, which takes seconds and produces the evidence within one rotation. Then work the ranked list from the top with a disposition ladder: **retire** rules that have never been actionable or that duplicate another rule; **demote** ones that carry information but do not need a human now; **retune** thresholds and pending durations where the rule is right but its timing is wrong; **collapse** fan-outs with grouping or dependency suppression; and best of all, **fix the system** so the condition stops happening. Set a target — Google's SRE guidance is that an on-call shift should handle at most about two incidents per 12 hours — make the review a standing item at every handoff, and require any new paging rule to state what a human will do at 3 a.m.

go deeper

for a junior

Know that a pager full of alerts nobody acts on is a problem to be fixed rather than endured, and that the first step is counting which rules actually produce the volume.

for a middle

Be able to lay out the disposition options for a single rule — retune the threshold or pending duration, collapse a fan-out by grouping, demote off the paging path, retire — and say what evidence would pick between them.

for a senior

Run the whole process: gather per-rule volume and actionability, attack the concentrated leaders, prefer fixing the system to tuning and tuning to deleting, demote before retiring, and watch the alert-detected share of incidents so you can prove you did not trade noise for blindness.

for a principal

Own the durable version — a stated pager-load target, a bar that every new paging rule must clear, the counter-metric that stops noise reduction becoming blindness, and a review anchored to an existing ritual so the improvement survives the people who made it.

## Why this question is asked Every SRE interview reaches some version of it, because it separates people who have carried a pager from people who have read about one. The weak answer is a list of techniques. The strong answer is a *process* with an order, evidence at the front, and an explicit account of what could go wrong. 50 pages a week across a rotation is roughly one every three hours around the clock. No human sustains that, and the damage is not only fatigue: a pager that is mostly wrong trains its recipients to acknowledge without reading, so the genuine page gets the same reflex as the other forty-nine. ## Step 1: measure before you touch anything Do not start by deleting the rule that annoyed you last night. Build the table: | per rule | why it matters | |---|---| | firings in 60–90 days | finds the volume leaders — noise is almost always concentrated in a handful of rules | | pages vs non-paging notifications | separates 'loud' from 'merely present' | | median time to acknowledge | very fast acks with no follow-up work is the signature of reflex dismissal | | share that led to a human action | the column that decides disposition | | clustering in time | distinguishes one fan-out event from a chronically noisy rule | The action column rarely exists already. Capture it going forward: at acknowledgement, the on-call marks the page actionable or not, optionally with one word of reason. It costs seconds, and after one full rotation you have real evidence rather than a debate about whose intuition is right. Distribution matters as much as totals — 50 pages a week is often three rules producing 40 of them, which means the first week of work is disproportionately valuable. ## Step 2: the disposition ladder Work the ranked list top-down. For each rule, in this order of preference: 1. **Fix the underlying system.** If a rule fires twice a week because a job legitimately runs out of memory twice a week, the honest fix is the memory, not the alert. This is the only disposition that reduces noise *and* improves reliability, and it is the one teams skip. 2. **Collapse the fan-out.** If one fault produces forty notifications, grouping on the right label subset or suppressing the downstream alerts behind their cause removes the volume without removing any signal at all. Cheapest real win available. 3. **Retune.** The rule is right but the timing or the number is wrong: a threshold set at normal peak behaviour, or a missing pending duration so it fires on transients. Retune and re-measure next rotation. 4. **Demote off the paging path.** The condition is worth recording and worth working during business hours, but nothing useful happens at 3 a.m. Move it to a queue that a human reads in daylight. 5. **Retire.** It has never produced an action, or another rule always fires alongside it and is the one people actually work. ## Step 3: the risk, and how to bound it The way this goes wrong is deleting the only detector for a rare but severe failure. It fired twice in three years, both times legitimately, and it looks like noise in a 90-day window. Bound it: **demote before you delete.** Move the candidate off the paging path for a quarter while keeping it firing and recorded, then check whether any incident in that quarter went undetected or was found by a customer instead of by an alert. If nothing was missed, retire it. Keep the pruning decisions themselves in version control with a reason attached, so a later postmortem can ask 'did we remove the alert that would have caught this?' and get an answer. Also track the counter-metric the whole time: **what fraction of incidents were detected by an alert rather than reported by a user?** Noise reduction that quietly drives that number down has made things worse, not better, and it is the single check that keeps this work honest. ## Step 4: make it recurring, or it decays A one-off cleanup regrows within two quarters. Anchor it to something that already happens: - **At every handoff**, the outgoing on-call presents the shift's pages and proposes a disposition for the noisiest one. This is the highest-leverage habit in the whole practice: it puts the change proposal in the hands of the person with the freshest evidence, at a meeting that already exists. - **A standing target.** Google's SRE book argues an on-call shift should handle at most roughly two incidents per 12-hour shift, on the reasoning that properly handling one — response, mitigation, write-up — consumes about six hours. Whatever number you adopt, having one turns 'the pager is bad' into a measurable, arguable state. - **A bar for new rules.** Any new page-severity rule states what a human is expected to do when it fires. If nobody can write that sentence, it is not a page. ## What to say in the interview Compress it: measure per-rule volume and actionability, attack the concentrated leaders first, prefer fixing the system over tuning the alert and tuning over deleting, demote for a quarter before retiring anything, watch the alert-detected share of incidents so you can prove you did not go blind, and hang the review on the on-call handoff so it does not decay.

  • How do you get the 'was this page actionable?' data if nobody has been recording it?
    Capture it forward rather than reconstructing it. Add a one-touch prompt at acknowledgement — actionable yes/no, plus an optional word — and after a single full rotation you have per-rule evidence. Reconstructing from chat history and ticket links is possible but slow and biased toward whatever was memorable, which is exactly the bias you are trying to escape.
  • You want to retire a rule that fired twice in three years, both times for a genuine severe failure. Would you?
    No — low frequency is not the same as low value, and that rule is not what is hurting the team. It is not in the volume leaders, so retiring it buys nothing while removing the only detector for a rare severe failure. Pruning should be aimed at the concentrated noise, and rare-but-true rules are the ones you deliberately keep.
  • After a month the pager is quiet. How do you prove you haven't just gone blind?
    Track the share of incidents detected by an alert rather than reported by a user, over the same period. If page volume fell and that ratio held or improved, the noise was genuinely noise. If customer-reported incidents rose, you removed signal — and you can go back to the versioned pruning decisions and identify which change to reverse.
  • Why is the on-call handoff the right place to anchor this review?
    Because it already happens, and because the person with the freshest, most specific evidence is in the room and highly motivated. Asking the outgoing on-call to propose a disposition for the shift's noisiest page turns a vague backlog item into one concrete change per week, which compounds faster than a quarterly cleanup that everyone dreads.

saying these in an interview costs you the question

  • Raise every threshold until the pager goes quiet
  • Delete any rule that fired more than ten times
  • We just need the team to be more disciplined about acknowledging
  • Silence the noisy rules and revisit next quarter
  • Fewer pages is always better, no counter-metric needed

context