skip to content

You are responsible for reliability standards across 60 teams holding roughly 4,000 alert rules, and you have no authority to edit another team's rules. How would you drive pager noise down across the whole organisation?

level: principalimportance: nice to knowfreq 30%

answer

  1. you cannot review 4,000 rules
  2. instrument what you cannot control
  3. defaults beat mandates
  4. budget with a pre-agreed consequence
  5. pages down, detection holding

basics

~20 s

Make noise measurable per team and visible, ship good behaviour as platform defaults rather than mandates, set a page-load budget teams own themselves, and pair it with a detection counter-metric so nobody hits the budget by going blind.

solid answer

~60 s

Central rule review does not scale to 4,000 rules and mandates get ignored by teams you cannot compel. Three levers actually work. First, **measurement and visibility**: publish pages per shift per team, actionability rate, and the share of incidents detected by an alert rather than by a customer — teams change behaviour when their number sits next to their peers'. Second, **defaults through the platform**: the shared routing pipeline supplies grouping keys, dependency suppression for shared infrastructure, and expiring silences; rule templates make the good shape the easy shape; and a lint in the delivery pipeline rejects a paging rule with no owner and no runbook link. Third, **a budget rather than an audit**: each team owns a pager-load target, and exceeding it for a period commits them to spend the next cycle on noise rather than features. The trap is that a page budget alone rewards deleting alerts, so it is only safe alongside the detection metric. And accept the residue: some teams will stay noisy, and central deletion of their rules is a fight not worth having.

go deeper

for a junior

Know that alerting standards across many teams are set through shared defaults and visible measurement rather than by one team editing everyone's rules.

for a middle

Be able to name the concrete central levers — shared grouping and routing, inhibition for shared infrastructure, expiring silences, rule templates, a pipeline lint for owner and runbook — and say why each works without needing authority.

for a senior

Show how you would measure it: per-team pages per shift, noise concentration in the top rules, and the alert-detected share of incidents, plus how you would pick the few teams worth hands-on help.

for a principal

Own the tradeoff explicitly. Argue why a budget with a pre-agreed consequence beats an audit, name the under-alerting incentive a page target creates and the counter-metric that neutralises it, and be honest about the residue you choose not to fight.

## Why the per-rule process does not scale here A single team can sit down with 90 days of firing data and work a ranked list. Sixty teams cannot be worked that way by one central group: 4,000 rules is more context than any platform team can hold, and the person who knows whether a rule is worth keeping is on the team that wrote it. The central role is not to make the decisions — it is to change the conditions under which sixty teams make them. The honest constraint in the question is the important one: **no authority to edit another team's rules.** Any answer that assumes you can mass-delete is answering a different question. ## Lever 1: measurement and visibility You can almost always instrument what you cannot control. Central paging infrastructure sees every notification, so you can publish, per team: - **pages per on-call shift**, especially out-of-hours pages; - **actionability rate**, if acknowledgement captures it; - **the noise concentration** — how much of a team's volume comes from its top three rules, which is usually most of it and makes the fix look tractable; - **the detection ratio** — incidents found by an alert versus reported by a customer. Publishing these side by side does most of the work, because a team seeing itself at 60 pages a shift next to peers at 3 does not need to be told anything further. Present it as a health signal, not a league table to punish with, or teams will optimise the number instead of the pager. ## Lever 2: defaults through the platform The highest-leverage central work is making the good configuration the one teams get for free: - **shared routing with sane grouping** so nobody has to discover grouping keys themselves, and a per-instance fan-out cannot reach a pager as forty notifications; - **dependency suppression for shared infrastructure** — the datastore, cluster control plane and network alerts you own centrally can inhibit downstream symptoms across all sixty teams from one rule set; - **silences that expire by policy**, with a maximum duration and automatic reporting of long-lived ones, so blind spots cannot accumulate quietly; - **rule templates** that carry a pending duration, an owner label and a runbook annotation by default; - **a lint in the delivery pipeline** — a paging-severity rule with no owner and no runbook link fails the check. This is enforcement without authority over content: you are not saying what may be alerted on, only that a page must name who and what. Each of these reduces noise for teams that never asked, which is the only kind of central intervention that scales. ## Lever 3: a budget, not an audit A reliability organisation already knows this shape from budgeted objectives: set the target, let the owning team choose how to meet it, and attach a commitment to what happens when it is missed. Applied to the pager, a team owns a load target — say, no more than a small handful of pages per 12-hour shift, and out-of-hours pages counted more heavily — and sustained breach commits them to spend the following cycle on alert quality rather than roadmap work. The commitment is what makes it real. A target with no consequence is a dashboard. The consequence does not need to be punitive; it needs to be pre-agreed with the team's own leadership, so that when it triggers it is a policy taking effect rather than an argument starting. ## The trap, and the counter-metric **A pager-load budget on its own rewards under-alerting.** The fastest way to hit any page target is to delete the rules, and a team under delivery pressure will find that path. This is the single most important thing to say about org-wide noise governance. The pairing that prevents it is the detection ratio: what share of a team's incidents were detected by an alert rather than reported by a customer or discovered by another team. A team is healthy only when both move in the right direction — pages down and detection holding. One number alone is gameable in an obvious direction, and both together are not. Additional guards: require that pruning decisions are recorded with a reason, and ask postmortems the direct question of whether a removed or suppressed alert would have caught the incident earlier. ## What you accept Some teams will remain noisy. Central deletion of their rules is a fight that costs more political capital than it returns and that ends with a team who no longer talks to you about reliability at all. The workable posture is: make the good path effortless, make the numbers visible, make the budget a pre-agreed commitment, help directly with the two or three worst offenders where you have a relationship — and let the residue be visible rather than hidden. Visible noise in three teams is a manageable state; hidden blindness across sixty is not.

  • Why can a pager-load target be actively harmful without a second metric?
    Because the cheapest way to meet any page target is to delete alerts, and a team under delivery pressure will take that route. Pairing it with the share of incidents detected by an alert rather than reported by a customer makes the shortcut visible: pages fall, detection falls, and the pairing exposes what a single number would have rewarded.
  • What can you enforce centrally without any authority over what teams alert on?
    Structural requirements rather than content ones. A pipeline lint can reject a paging-severity rule with no owner label and no runbook annotation, and platform policy can cap silence duration and expire silences automatically. You are not deciding what deserves a page — only that every page names who owns it, what to do, and that suppression cannot become permanent.
  • You have capacity to work directly with only three of the sixty teams. How do you choose?
    By out-of-hours page volume weighted against the team's willingness to engage. Pick from the top of the volume list, because noise is heavily concentrated and a small number of teams generate most of it, but spend the effort where there is an invitation — a team that wants the help produces a case study others copy, while a resistant one consumes the whole budget and produces nothing reusable.

saying these in an interview costs you the question

  • Mandate a rule review and delete what teams don't justify
  • Set a page budget and let each team hit it however they like
  • One central team should own all alert rules
  • Publish a league table and let shame do the work
  • If pages went down, the programme worked

context