A paging alert on your service fired 20 times last month, and only 3 of those pages led a responder to take any action. Express that as the alert's precision, explain what you give up when you tighten the rule to improve it, and say how you would decide where to sit on that trade.
answer
- three actionable out of twenty pages
- the two errors do not cost the same
- low precision destroys human recall
- misses are invisible; use discovery source
- fast and strict, plus slow and broad
basics
~20 sPrecision is 3/20, or 15% — roughly six of every seven pages were noise. Tightening the rule raises precision but lowers recall, missing real events. Paging tier should favour precision; lower tiers can absorb lower precision.
solid answer
~40 sPrecision is the share of pages that were genuinely actionable — here 3 of 20, about 15%. Recall is the share of real incidents the rule actually caught, and the two move against each other: any change that makes the rule fire less often raises precision and risks dropping a real event. Where I sit depends on the cost asymmetry and on whether anything else covers a miss. At paging tier I want precision high, because a 15% alert quietly destroys human recall — responders learn to assume it is noise, and the one page that mattered gets acknowledged and set aside with the rest. I would rather run one high-precision fast page backed by a slower, higher-recall detector at ticket tier than one rule trying to do both jobs badly.
go deeper
Know that a page which usually turns out to be nothing is a defect, and that the fix is not simply to ignore it. Be able to say what fraction of pages led to action.
Define precision and recall for an alert and show why they move against each other, and give the arithmetic on a concrete case rather than describing the trade in the abstract.
Demonstrate that you measure the actionable rate, use incident discovery source as the recall proxy, and layer a strict paging rule with a broader lower-tier detector instead of tuning one rule to do both.
Own the standard: a published precision floor for paging-tier rules, who enforces it, and how you fund the layered detection that lets teams keep recall while the pager stays quiet.
## The two numbers Borrowing the classification terms is worth doing because they force honesty about what you are trading. - **Precision** = actionable pages ÷ all pages. In the example, 3 ÷ 20 = 15%. Six out of seven interruptions bought nothing. - **Recall** = real events detected ÷ all real events. This is the number that says whether the rule is doing its job at all. Every knob you can turn — the threshold, how long the condition must hold, how narrowly the rule is scoped — moves both, in opposite directions. Make the rule stricter and fewer false pages get through, but some real event that was just under the line now goes undetected. There is no setting that improves both; there is only a choice about which error you would rather make. ## The asymmetry that decides the answer The two errors do not cost the same, and the ratio differs by service. A false positive costs a human interrupt: call it half an hour of attention, plus the residue of a broken night. That sounds survivable, and per-page it is. What makes it dangerous is that the cost compounds into a *detection* failure. At 15% precision the rational response of an experienced responder is to assume the page is noise until proven otherwise. The rule's measured recall is unchanged, but the recall of the human plus the rule — which is the only thing that matters — has collapsed. This is the concrete mechanism behind alert fatigue, and it is why precision is not a comfort metric. A false negative costs undetected user impact, running until something else surfaces it — usually a customer. The incident clock effectively starts at the complaint, and the postmortem timeline shows the data was there all along. So the decision rule is: at paging tier, favour precision, because the failure mode of low precision is itself a loss of recall. At ticket tier you can afford far lower precision, because the cost of a wrong ticket is a few minutes of someone's working day rather than a night. ## You can measure precision; you cannot directly measure recall This asymmetry catches people out. Precision is straightforward: after each page the responder records whether they took an action, and the actionable rate over a month falls out. Many teams set an explicit floor for paging-tier rules — something like "at least seven in ten pages must have been actionable" — and treat a rule below it as a defect with an owner. Recall is harder, because misses are invisible by construction. The standard proxy is how incidents were discovered: the share detected by an alert versus by a customer report, a support ticket, or an engineer noticing something. If most of your incidents arrive by a route other than the pager, your alerting has a recall problem, whatever the individual rules look like. That single metric is often the most informative thing on an alerting review. ## Defence in depth instead of one perfect rule The most useful practical move is to stop trying to make one rule both fast and complete. Layer instead: - A **fast, strict, high-precision** rule at paging tier, deliberately tuned to fire only on conditions that are unambiguous. It will miss slow or subtle degradations, and that is accepted. - A **slower, broader, high-recall** detector that catches what the strict rule missed, delivered at a lower tier — a ticket, or a review in the working day. It will be wrong more often, and that is affordable at that tier. The cumulative recall of the pair beats either alone, and the pager only carries the high-precision half. Note that budget-consumption-based detection is a common backstop for exactly this role, and its windowing math belongs to the error-budget discipline rather than here. ## How to actually move a bad rule When you have a 15% rule in front of you, resist reaching straight for the threshold. Look first at *why* the seventeen non-actionable pages fired, because they are usually not one population: - Some are genuine transients that resolve before anyone can log in — those are a duration-of-condition problem. - Some are a single misbehaving instance or tenant dragging an aggregate — a scoping problem. - Some are real degradations that simply were not worth a human — those are a tier problem, not a threshold problem. Each class has a different fix, and lumping them together into one threshold bump is how teams end up with a rule that is both quiet and blind. Whatever you change, change one thing, keep the previous behaviour visible at a lower tier for a few weeks, and re-measure the actionable rate before declaring victory. ## Stating it in an interview The answer that lands is the numeric one: name the precision, name what tightening costs, name the cost asymmetry that decides the direction, and name the metric you would watch afterwards. "We'd tune it" is what everyone says; "it was at 15%, our paging floor is 70%, and here is how I'd find which of the three noise populations dominates" is what a lived answer sounds like.
- If misses are invisible, how do you know your alerting has a recall problem?Track how incidents were discovered. The share detected by an alert versus by a customer report, a support ticket or an engineer noticing is the closest thing to a measurable recall figure. If most incidents arrive by a route other than the pager, alerting has a recall problem regardless of how healthy the individual rules look on their own dashboards.
- Why is lowering precision more dangerous at paging tier than at ticket tier?Because the cost of a false page compounds into lost detection. Responders learn to treat a noisy rule as noise, so the rule's measured recall stays flat while the recall of the human-plus-rule system collapses. At ticket tier a wrong item costs a few minutes of working time and nobody's trust, so a much lower precision is affordable.
- A team proposes fixing a 15%-precision alert by raising its threshold by 50%. What would you push back on?That the seventeen non-actionable pages are probably several different populations — brief transients, one instance skewing an aggregate, and real-but-minor degradations — and each needs a different fix. A single threshold bump treats them as one and typically buys quiet at the cost of blindness. Change one thing, keep the old behaviour visible at a lower tier, and re-measure.
saying these in an interview costs you the question
- Judging an alert only by whether it fires
- Assuming a quieter alert is automatically a better one
- Treating alert fatigue as morale rather than lost detection
- Claiming recall can be measured directly
- Fixing every noisy alert by raising the threshold