skip to content

You inherit an on-call rotation whose forty-odd paging rules are almost all host-level thresholds — CPU, memory, disk, process restarts, queue depth. How do you move that alert set to symptom-based paging without going blind during the transition?

level: seniorimportance: should knowfreq 52%

answer

  1. write the promise before touching the rules
  2. add first, remove second
  3. run both sets in parallel
  4. the interesting quadrant is impact with no cause page
  5. demote to ticket and dashboard, never delete

basics

~20 s

Define the service's user-visible signals first, add symptom pages for them, then run both sets side by side for a few weeks. Demote every cause rule that fired without user impact; keep the few that caught real incidents nothing else saw.

solid answer

~50 s

I would not delete anything on day one. First I write down what the service promises and measure it at the boundary — error ratio and tail latency per user journey — and add a small number of symptom pages, typically one to three per journey. Then I run both alert sets in parallel for two to four weeks and build a two-by-two: cause alerts that fired with no user impact get demoted to tickets or dashboards; incidents that only a cause alert caught are the real coverage gaps and I either keep that alert or, better, find the symptom that was missing. The output is a short paging list plus an explicitly kept set of cause pages for the certain-and-imminent cases. The parallel period is what makes this safe, and the demotion — rather than deletion — is what makes the team accept it.

go deeper

for a junior

Be able to say the order of operations — define the user-visible signals, add those alerts, only then demote the old ones — and that nothing gets deleted outright.

for a middle

Explain the parallel-run comparison and what each outcome implies: no-impact firings get demoted, and an incident caught only by a cause alert marks a coverage gap to investigate.

for a senior

Show the traps you would check for in a real estate: low-traffic ratio noise, asynchronous impact, internal-only callers, and host rules that quietly covered several services.

for a principal

Own the measurement that justifies the change — pages per week, the share of pages that produced action, and the share of incidents found by alert rather than by a customer — and be honest about where detection got slower.

## Start from the promise, not from the alerts The instinct is to open the alert file and start deleting. That is backwards: you cannot tell what the forty rules were protecting until you know what the service is supposed to do. So the first step is a page of writing, not configuration. Enumerate the user journeys the service owns — for a payments service perhaps *authorise a card*, *refund a charge*, *fetch transaction history*. For each, decide where the boundary is (the edge proxy, the API gateway, the caller's client) and what "served correctly" means, including the failures that return a 200. That gives you a handful of measurable symptom signals. ## Add before you remove Write the symptom alerts and turn them on **alongside** the existing forty. One to three per journey is normally enough: an error-ratio condition, a tail-latency condition, and where relevant a traffic-collapse condition that catches the case where requests stop arriving entirely. At this moment the pager is noisier, not quieter. That is the price of not going blind, and it is temporary. Set the parallel window in advance — two to four weeks, long enough to include a deploy cycle, a peak-traffic day and ideally a real incident. ## Classify what fires During the parallel window, record every firing in one table with two columns: *did a cause alert fire?* and *was there user-visible impact?* ``` impact no impact cause alert fired keep or DEMOTE replace (the bulk) cause alert silent GAP - fine investigate ``` - **Fired, no impact** — the dominant quadrant, and the whole reason the rotation is miserable. These become tickets or dashboard panels. Do not delete the metric; change its destination. - **Fired, with impact, and the symptom alert also fired** — the cause alert added nothing to detection. Demote it too; it stays valuable as diagnostic context, and it will be on the incident dashboard. - **Fired, with impact, symptom alert silent** — a genuine coverage gap and the most valuable finding of the exercise. Ask *why* the symptom was invisible. Usually one of: the impact was on a journey you did not instrument; the service is low-traffic so the ratio was too noisy to trip; or the harm was asynchronous and arrives at the user hours later. Fix the symptom coverage if you can, and keep the cause page if you cannot. - **Silent, with impact** — the symptom alert caught something the old set never would have. This is the evidence you present when the change is questioned. ## Demote, do not delete Every cause metric stays collected, and most of them end up on the service's incident dashboard, laid out so that a responder who has just been paged on the error ratio can see CPU, memory, pool occupancy, restarts and queue depth in one glance. This matters politically as well as technically: the engineer who wrote the memory alert three years ago is much easier to convince when the answer is "it moves to the dashboard and the ticket queue" than "we are deleting it". A reasonable end state for the forty rules is something like: three to six symptom pages, three to five kept cause pages for the certain-and-imminent conditions (disk projected to fill, certificate expiry, redundancy loss on a stateful tier), and the remaining thirty as ticket rules and dashboard panels. ## The traps **Low-traffic services.** An error-ratio alert on a service handling two requests a minute is a coin flip; one failure is 50%. For these, alert on absolute counts over a longer window, on synthetic probes that generate steady traffic, or accept slower detection deliberately. **Asynchronous and batch work.** If the user impact of a stuck consumer only appears in tomorrow's report, the symptom is too late to be the only page. The right symptom is usually freshness or the age of the oldest unprocessed item, which is user-visible in the sense that matters, and which is a symptom rather than a cause. **Internal-only services.** Your "user" is the calling service. Measure at the interface you expose to it. Do not conclude that a service with no human users must be alerted on host metrics. **Multiple services per host.** Host-level rules often protected several services at once. When you demote them, check that each co-located service actually acquired its own symptom coverage. ## Prove it afterwards Measure before and after: pages per week, the fraction of pages that led to a change or a mitigation, and how many incidents were detected by an alert rather than by a customer report. Those three numbers turn the migration from a matter of taste into a result you can defend, and they tell you honestly if detection got slower.

  • How do you handle a service whose traffic is too low for an error-ratio alert to be meaningful?
    Ratios are unstable at low volume, so use absolute failure counts over a longer window, or run synthetic probes so there is a steady stream of requests to measure. If neither is worth the effort, decide deliberately that detection will be slower and record that choice, rather than quietly reverting to host thresholds.
  • During the parallel run, a cause alert catches an incident that your symptom alerts missed. What do you do with it?
    Treat it as a coverage gap first and a keeper second. Ask why the symptom was invisible — an uninstrumented journey, too little traffic, or impact that lands asynchronously — and fix that if you can, because a symptom alert generalises to causes you have not met. Keep the cause page only if no symptom can be measured.
  • How do you get the team to accept losing alerts they wrote themselves?
    Frame it as demotion, not deletion: every metric is still collected, most land on the incident dashboard, and the noisy ones become tickets. Then show the parallel-run data — which of their alerts fired without impact, and which incidents the symptom alerts caught. Numbers from their own rotation settle it far better than principle does.

saying these in an interview costs you the question

  • Deleting the old rules before symptom coverage is proven
  • Assuming an error-ratio alert works on a two-request-per-minute service
  • Treating the migration as done when the alert count drops
  • Forgetting host rules that were covering several co-located services
  • Leaving asynchronous pipelines with no freshness signal

context