skip to content

How do you set an anomaly-detection threshold that trades false alarms against detection delay for on-call?

level: principalimportance: should knowfreq 42%

answer

  1. you cannot improve both sides
  2. quiet and fast are the same knob
  3. mean gap between alarms, mean delay to detect
  4. multiply the rate by the fleet size
  5. budget pages per week, then derive

basics

~20 s

Set the threshold from an alarm budget, not from a default number of sigmas. Measure false alarms per week and median detection delay across the whole fleet of series, then pick the point on that curve on-call can sustain.

solid answer

~50 s

There is no threshold that improves both sides — every notch of quiet buys delay, and the honest deliverable is a curve, not a number. I frame it with average run length: the in-control ARL is the mean gap between false alarms when nothing is wrong, and the out-of-control ARL is the detection delay under the shift I care about. Then I convert both into the units the business argues in: alarms per on-call week, and minutes of undetected incident. The step most people miss is **fleet multiplication** — a per-series false-alarm rate of one in a thousand points looks excellent until five hundred series run it, and that is an alarm every two points. So the budget is set at the fleet level and pushed down. I also decouple sensitivity from paging: cheap signals go to a dashboard or a ticket, and only a persistent, corroborated breach pages someone. And I backtest on labelled history so the curve is measured, not asserted.

go deeper

for a junior

Be ready to say that a tighter threshold catches problems sooner but raises false alarms, and a looser one does the reverse, so the setting depends on which mistake costs more.

for a middle

Explain the two quantities precisely: the mean gap between false alarms when nothing is wrong, and the mean delay to detect a shift of a given size, and how moving one threshold changes both in opposite directions.

for a senior

Show you have tuned a real pager: converting run lengths into alarms per week, backtesting on labelled incidents, using persistence rules and severity tiers, and suppressing known deploy windows instead of loosening thresholds.

for a principal

Own the budget. Set a fleet-wide alarm allowance, account for multiplication across hundreds of series, price a false page against a minute of undetected incident, and present the delay-versus-alarms curve so the pager owner picks the operating point.

## The tradeoff is a curve, and the curve is the deliverable Every detector has one knob that trades the same two quantities in opposite directions: - **False-alarm rate** — how often it cries wolf when nothing is wrong. - **Detection delay** — how long a genuine change runs before it is caught. Tighten the threshold and you find things faster and wake people for nothing more often. Loosen it and the pager goes quiet while incidents run longer. No amount of cleverness in the detector removes the tradeoff; a better detector moves the whole curve inward, but you still choose a point on it. The standard vocabulary is **average run length (ARL)**. The in-control ARL is the expected number of observations between false alarms while the process is fine — bigger is quieter. The out-of-control ARL is the expected number of observations to detect a shift of a given size — smaller is faster. Quoting both, for the specific shift size worth catching, is what separates an engineering answer from "we used three sigma". ## Convert statistics into operational units ARL in "observations" persuades nobody. Convert it. If the series is sampled every minute and the in-control ARL is 1,000 points, that is a false alarm roughly every 17 hours — about ten a week from **one** series. Detection delay of 10 points is 10 minutes of undetected incident. Now the conversation is about things people can price: an on-call engineer's tolerance is a handful of actionable pages per week, and the cost of an undetected outage is revenue or trust per minute. A rough decision frame: choose the threshold that minimises expected total cost, `alarms_per_week * cost_per_false_page + incidents_per_week * cost_per_minute * detection_delay`. The numbers will be soft, but writing them down forces the real conversation — how much is a 3am page worth versus five more minutes of downtime — and that conversation is the actual deliverable at a senior level. ## Fleet multiplication is where thresholds die The most common production failure is tuning per series and deploying across a fleet. A false-alarm rate of one per thousand points per series is fine alone. Run it on 500 series and the expected rate is `500/1000 = 0.5` alarms per time step — one every two points, forever. So set the budget globally first — "no more than N actionable pages per on-call week across all monitors" — and derive per-series thresholds from it. Practical levers: - **Tier the series.** A handful of top-line metrics get sensitive thresholds and page; the long tail gets loose thresholds and a dashboard. - **Aggregate before detecting.** Monitoring one well-chosen aggregate beats monitoring a thousand slices, each of which alarms independently. - **Correct for multiplicity.** If you insist on per-slice detection, the per-series threshold must tighten as the fleet grows, otherwise the global rate grows linearly with it. ## Levers other than the threshold The threshold is not the only knob, and the others often buy a better trade: - **Persistence (k-of-n).** Require, say, three consecutive breaches before alarming. Random breaches rarely repeat, so false alarms fall steeply while delay grows only by the extra points — usually the best value per unit of delay available. - **Accumulate rather than threshold per point.** Cumulative methods catch a small sustained shift that a per-point rule at the same false-alarm rate never sees, buying detection power at no false-alarm cost for the shift sizes they are tuned to. - **Severity tiers.** Page, ticket, dashboard. Most detections do not need a human awake; routing by severity lets you run a sensitive detector without a punishing page rate. - **Suppression windows.** Mute known deploys, migrations and scheduled events. An alarm you already know the cause of is pure noise and trains people to ignore the channel. - **Corroboration.** Require two independent signals to agree before paging. This cuts false alarms far more than it costs in delay, because independent false alarms rarely coincide. ## Measure the curve; do not assert it Backtest on labelled history: take incidents you know about with their start times, replay the detector across a grid of thresholds, and plot alarms-per-week against median detection delay. Present that curve to whoever owns the pager and let them choose the operating point — the choice is theirs, the measurement is yours. Two cautions. Labels are incomplete: you know the incidents that were noticed, which biases delay estimates optimistically for loud failures and hides the quiet ones entirely. And the curve drifts as the system changes, so it needs re-measuring, along with an ongoing count of what fraction of pages were actionable. A precision rate below about a third destroys trust in the channel, and once on-call stops reading the alerts your effective detection delay is infinite regardless of what the threshold says. ## The one-line version No threshold is correct in isolation; it is correct relative to an alarm budget, a fleet size, a shift size worth catching, and the cost asymmetry between a false page and a slow detection. Say that, show the curve, and name fleet multiplication — that is the principal-level answer.

  • What does an in-control average run length of 500 actually tell an on-call team?
    That while nothing is wrong, roughly 500 observations pass between false alarms on that series. Its operational meaning depends entirely on sampling rate and fleet size: at one point per minute it is a false alarm about every eight hours from a single series, and across a hundred such series it is several per hour. Always convert it to alarms per on-call week before quoting it.
  • Why does requiring three consecutive breaches before paging usually pay for itself?
    Because false breaches are largely independent while genuine ones persist. Needing a run of them cuts the false-alarm rate steeply, while the cost is only the extra observations spent waiting — a small, bounded delay. It is the best exchange rate on the curve for most metrics. The exception is a fast, catastrophic failure mode, where even a few points of delay is unaffordable and a separate fast path is warranted.
  • How do you set thresholds when you have almost no labelled incidents to backtest against?
    Calibrate the false-alarm side from in-control history, which you have plenty of: replay the detector over known-quiet periods and measure the alarm rate directly. Then derive detection delay analytically or by simulating injected shifts of the sizes worth catching. That gives the whole curve without incident labels, and the few labels you do have become a sanity check rather than the foundation.
  • What do you monitor about the alerting system itself?
    The actionable fraction of pages, the fleet-wide alarm rate against its budget, the share of incidents found by a human before the detector, and the time from onset to acknowledgement. Falling precision predicts that alerts will start being ignored, at which point effective detection delay is unbounded no matter what the threshold says. Detector quality is an operational metric, not a one-time tuning exercise.

A smoke alarm sensitive enough to catch smouldering wiring also goes off when you make toast. Households solve it the same way teams should: not one setting, but different responses for different signals.

saying these in an interview costs you the question

  • Quotes three sigma as a universal threshold
  • Tunes on one series and ships to a whole fleet
  • Talks about false alarms with no mention of detection delay
  • Reports run lengths in points, never in pages per week
  • Assumes a better model removes the tradeoff entirely
  • Ignores that ignored alerts make detection delay effectively infinite

context