A single database failover pages the same on-call engineer five separate times, from five alert rules written by three different teams. Suppressing the extra notifications only hides the problem — what do you change about the paging rules themselves, and how do you decide which one survives?
answer
- five rules, one underlying event
- suppression is cross-team and lossy
- group rules by symptom, not by metric
- keep the one closest to user impact
- pages per incident, target about one
basics
~20 sConsolidate ownership: one user-visible symptom should have exactly one paging rule, owned by one team. The other four become ticket- or log-tier diagnostic signals. Duplicate paging is a rule-inventory and ownership problem, not a notification-filter problem.
solid answer
~50 sFive rules firing on one event usually means four teams each instrumented their own layer of the same failure — client errors, dependency timeouts, database saturation, a synthetic probe — and nobody ever deleted anything. Notification-side suppression can quieten a storm within one rule set, but it cannot resolve rules owned by different teams, and it can hide a genuinely separate second failure. So I run the inventory: for the last ten incidents, list which paging rules fired, and group them by the symptom they detect. For each symptom, keep as the page the rule closest to user impact, give it a single owning team, and demote the rest to ticket tier or to incident-channel context where they remain valuable for diagnosis. Then track pages-per-incident as a standing metric, with a target of about one.
go deeper
Recognise that receiving five pages for one outage is a defect in the alerting, not a sign of thorough monitoring, and that the extra pages did not add information.
Explain how duplicates accumulate across teams and why grouping rules by the symptom they detect — rather than by the metric they read — is what reveals the true count.
Argue why suppression is the wrong instrument across team boundaries, choose the surviving rule as the one closest to user impact, and preserve the demoted ones as incident context.
Own the governance: a generated registry of paging rules, a symptom check at promotion time, pages-per-incident as a tracked metric, and making it politically safe for a team to give up a page.
## Why the duplicates exist Nobody sets out to build five pagers for one event. It accretes. The client team adds an error-rate page after being blamed for a silent outage. The service team adds a dependency-timeout page. The database team adds a saturation page. Someone adds a synthetic probe after a postmortem action item. Each addition is locally rational, each was approved by a different reviewer, and none of them is aware of the others because there is no place where paging rules are listed as a set. The result is a system where a single underlying failure produces five interrupts, and where reducing the count requires a conversation across three teams — which is why it does not happen spontaneously and why it is genuinely a leadership problem rather than a tuning task. ## Why suppression is not the fix here Grouping, dependency-based inhibition and deduplication are real and useful, and inside one team's rule set they are the right instrument. They do not solve this case for two reasons. First, they are configured by whoever owns the notification path, and the five rules here are owned by three teams with three different views of what depends on what. Encoding "suppress the database saturation page when the checkout error page is firing" requires a dependency model that spans team boundaries and must be maintained as the architecture changes. That is a large ongoing commitment to preserve rules that should not exist. Second, suppression is lossy in a dangerous direction. A rule configured to swallow the database page during a checkout incident will also swallow it during a *different*, coincident database problem. You have bought quiet by accepting a class of blind spot. Suppression is the right tool for a storm from one rule across many instances. It is the wrong tool for four teams alerting on the same symptom. ## The inventory The exercise is concrete and takes an afternoon. 1. Take the last ten incidents. For each, list every paging rule that fired. 2. Group the rules by the *symptom* they detect, not by the metric they read. Client 5xx, dependency timeout and connection-pool exhaustion during a failover are three measurements of one symptom: the service cannot serve. 3. Count, per incident, how many distinct rules paged. This is your pages-per-incident figure, and it is usually a shock the first time it is calculated. The grouping step is where the judgment lives, and the useful question is: *if this rule fired alone, would the response be different?* If two rules always fire together and lead to the same first action, they are one alert wearing two names. ## Choosing the survivor For each symptom group, keep as the paging rule the one closest to user impact. It is the one that best justifies waking a human, it is the one whose absence would mean a silent outage, and it does not go quiet when the cause changes — a saturation-based rule detects one specific failure mode of the database, while a "checkout is failing" rule detects all of them. The demoted rules are not deleted. They move to ticket tier, or become signals published into the incident channel and the service dashboard, where they remain genuinely valuable: during the incident, the demoted saturation alert is the fastest available clue about *why*. Preserving them as diagnostic context, rather than deleting them, is usually what makes the owning team willing to give up their page. Every surviving paging rule then gets one named owning team. The three-teams-one-symptom situation is exactly the thing that regenerates duplicates if the owner is left ambiguous. ## The residual legitimate case Sometimes two pages for one cause are correct: when one root cause produces two independent user-visible symptoms that different teams must act on in parallel — say, checkout failing and a data pipeline stalling, requiring simultaneous mitigation and backlog handling. That is real, and the test is whether the two responses are genuinely different work happening at the same time. If the second team's action is "watch and wait for the first team", it is not a page, it is an incident-channel notification. ## Governance that keeps it from regrowing The inventory decays without something holding it. Three mechanisms, in increasing cost: - **A visible registry of paging-tier rules per service**, generated from the definitions rather than maintained by hand, so that adding a page is an act that others can see. - **A review at promotion time**: adding a rule at paging tier requires naming the symptom it detects and confirming no existing page covers that symptom. This is cheap and catches most regrowth. - **Pages-per-incident as a tracked metric**, reviewed alongside pager load. A target near one is the honest goal; anything consistently above two says the inventory has drifted again. ## Why this is worth leadership attention The engineering fix is small. The reason it stays broken is that each duplicate is a rational local decision by a team protecting itself from blame for a missed outage, and no individual team can safely remove a page that someone else's postmortem created. Making it safe to give up a page — by keeping the signal alive at a lower tier, by naming a symptom owner, and by ensuring the team that keeps the page is accountable for detecting the symptom — is the actual work.
- When is notification-layer suppression the right instrument rather than rule consolidation?When one rule produces many notifications for one event — a fleet-wide condition firing per instance, or a storm across many series. That is genuinely a delivery-shaping problem. It is the wrong instrument when several distinct rules, especially rules owned by different teams, detect the same symptom, because the dependency model then spans team boundaries and the suppression also blinds you to coincident unrelated failures.
- How do you get a team to give up a page they added after being blamed for a missed outage?Keep their signal alive at a lower tier and in the incident channel, so detection is not lost, and make explicit who is accountable for detecting that symptom — normally the owner of the surviving page. The objection is almost never about the notification; it is about being exposed again. Removing that exposure is what makes consolidation possible.
- Is there a case where one incident legitimately pages two teams?Yes, when a single cause produces two independent user-visible symptoms requiring genuinely different work in parallel — for example checkout failing while a data pipeline stalls, needing mitigation and backlog handling at once. The test is whether both responses are real, simultaneous work. If the second team's role is to watch the first, that is an incident-channel notification rather than a page.
saying these in an interview costs you the question
- Fixing cross-team duplicates with notification suppression
- Assuming every existing paging rule earned its place
- Deleting demoted alerts instead of keeping diagnostic value
- Grouping alert rules by metric rather than by symptom
- Leaving a symptom with no single accountable owner