A service pages the on-call engineer whenever a host's CPU stays above 85% for five minutes, and users have never noticed anything during those pages. In alerting design, what is the difference between a symptom-based and a cause-based alert, and which of the two belongs on the pager?
answer
- what users feel versus what machines feel
- one open set, one closed set
- pages, tickets, logs — three destinations
- CPU high, latency fine: nobody is harmed
- causes are for diagnosis, not for waking people
basics
~20 sSymptom alerts fire on what users experience — failed requests, slow responses, a service level burning down. Cause alerts fire on machine state such as CPU or disk. Page on symptoms; send cause signals to tickets and dashboards.
solid answer
~50 sA symptom alert describes something a user could feel: requests failing, requests getting slow, a queue of work that will not drain, a service level objective burning down. A cause alert describes machine state that may or may not reach a user: CPU, memory, disk, thread pools, restart counts. The pager should be reserved for symptoms, because waking a human is only justified when someone is being harmed and a human can stop it. CPU at 85% with healthy latency and error rate is not harm — it is capacity information, so it belongs on a dashboard or in a ticket. The practical consequence is coverage: one alert on the error rate of a user-facing endpoint catches every cause of failure, including the dozen you never thought to threshold, while a hundred host thresholds still miss the one that mattered.
go deeper
Be able to sort example conditions into two piles — things a user would notice, and things only a machine notices — and say plainly that the pager is for the first pile.
Explain why symptom alerts give better coverage: causes are an open-ended set, while every cause that matters surfaces in the same few user-visible measurements. Name the page/ticket/log destinations.
Show the demotion in practice — what you did with the old host thresholds, how you verified nothing went uncovered, and how the cause metrics still shorten diagnosis once a symptom page fires.
Own the tradeoff at fleet scale: symptom paging is a policy that trades a small chance of late detection for a large gain in responder trust, and you should be able to say how you would measure whether that trade paid off.
## The two kinds of alert Every monitoring condition you write sits somewhere on a line between two poles. A **symptom-based** alert fires on evidence that the service is failing at the point where a user or a calling system touches it. Typical forms: the fraction of requests returning server errors over the last few minutes, the 99th-percentile response time of a public endpoint, the age of the oldest unprocessed item in a work queue, the fraction of a service level objective's budget consumed so far. These conditions say *something is wrong for someone*. A **cause-based** alert fires on internal machine state that is one or more steps removed from that experience: CPU utilisation, memory used, disk free, connection-pool occupancy, garbage-collection pause time, a process restart counter, replica count. These conditions say *something is unusual inside the box*. The distinction is not about which metric is more useful — cause metrics are indispensable for diagnosis. It is about which one earns the right to wake somebody up. ## Why paging on symptoms wins Three arguments, in the order interviewers usually want to hear them. **Coverage.** The set of causes of an outage is open-ended and you cannot enumerate it. A bad deploy, a poisoned cache, a dependency's timeout, a certificate, a leap-second bug, a noisy neighbour on shared hardware — each has its own cause signal, and you will not have thought of all of them. But every one of them, if it matters, shows up in the same small set of user-visible measurements. One alert on the error ratio of the checkout endpoint covers causes nobody has invented yet. **Actionability.** A page is a claim that a human must act now. CPU at 92% with normal latency and no errors is not a claim of harm; it is a claim about headroom, and headroom is a planning problem with a calendar, not a pager. When a page arrives that requires no action, the responder learns something dangerous: that pages can be ignored. Precision loss is cumulative and hard to undo. **Correctness under redundancy.** Modern systems are built so that individual components can be sick without users noticing. A single node pegged at 100% CPU behind a load balancer with nine healthy peers is exactly the design working. A cause-based pager punishes you for building redundancy, because it fires on every degraded component whether or not the redundancy absorbed it. ## Where the golden signals fit The common taxonomy for symptom signals is latency, traffic, errors and saturation. Three of the four — latency, errors and, when it collapses to zero, traffic — describe what a request experienced, and are natural paging material. Saturation is the odd one out: it measures how full a constrained resource is, which is a cause. It is the leading indicator you want on a dashboard and in capacity reviews, not the thing you page on by default. ``` page : error ratio, tail latency, request rate collapsed to zero, objective budget burning down ticket : CPU, memory, disk trend, pool occupancy, restart counts, single-replica failures absorbed by redundancy log only: everything you want during a postmortem and nothing else ``` The three-way split — page, ticket, log — is the practical form of the rule. Nothing gets deleted; things get demoted. ## Working the CPU example Apply the test to the alert in the question. Is any user harmed? No — latency and error rate are normal, so the answer is no by measurement, not by assumption. Is there an action only a human can take right now? No — the action is to decide whether to add capacity, and that decision has days of runway. Therefore: demote. Keep the CPU series, put it on the capacity dashboard next to the request rate that drives it, and file a ticket when the trend crosses a planning threshold. Then ask the question the CPU alert was standing in for. Probably: *will this host run out of headroom and start dropping requests?* That is answerable directly by measuring the dropped requests and the latency, which is a symptom, and which will fire whether the cause turns out to be CPU, a slow dependency or a lock. ## The honest exception The rule is not absolute. A cause deserves a page when the impact is certain and the lead time is short relative to the time it takes to fix — a disk that will be full within the hour, a certificate expiring tonight. That is a narrow, justified list, not a licence to keep the old thresholds. ## What the interviewer is listening for Candidates who have only read about this recite "alert on symptoms". Candidates who have carried a pager say what they did with the cause signals afterwards — that they were kept, that they were routed to tickets and dashboards, and that during an incident the symptom page tells you *that* you are broken while the cause metrics tell you *where*.
- If you only page on symptoms, how do you still find the cause quickly during an incident?You keep every cause metric — you just stop paging on it. The symptom page tells you the service is failing; the cause signals, on a dashboard beside it, tell you where. Alerting and diagnosis are different jobs done by the same telemetry, and demoting an alert to a ticket or a dashboard removes none of its diagnostic value.
- Does the symptom rule change for a backend service whose only callers are other internal services?No, but the definition of "user" moves. The symptom is what the calling service experiences at your boundary: the error ratio and latency you return to it. Measure at that interface rather than inventing internal proxies, and page on it exactly as a public service would.
- Where should latency be measured for a symptom alert — client side or server side?As close to the user as you can measure reliably. Server-side latency misses connection setup, network problems and requests that never arrived; a load balancer or edge measurement catches those. Server-side is a reasonable default because it is cheap and complete, but be explicit that it under-reports failures upstream of your process.
saying these in an interview costs you the question
- Treating high CPU as an outage even when latency and errors are normal
- Believing more alerts means better coverage
- Thinking cause metrics must be deleted once you page on symptoms
- Assuming a single degraded replica behind a load balancer deserves a page
- Claiming symptom alerts cannot detect a problem until users complain