If the rule is "page on symptoms, not on causes", when is it nevertheless correct to page on a cause signal such as a filling disk or an expiring TLS certificate?
answer
- certain, not merely possible
- compare runway against repair time
- some failures are cheaper before than after
- N+1 down to N: nobody feels it yet
- a deadline is a better condition than a level
basics
~20 sPage a cause when impact is near-certain and the remaining lead time is close to the time needed to fix it — a disk filling within the hour, a certificate expiring tonight. Otherwise the runway is long enough for a ticket.
solid answer
~50 sThe exception has a test with three parts. First, **certainty**: the condition will become user-visible if nobody acts — a disk projected to fill, a certificate with a fixed expiry, a licence or quota with a deadline. Second, **lead time versus fix time**: if the remaining runway is comparable to how long the fix takes, waiting for the symptom means the outage is already unavoidable. Third, **cost asymmetry**: some failures are far more expensive after impact than before — a full disk can corrupt state or block the very writes you need to recover, and a lost quorum cannot be restored by adding capacity later. Loss of redundancy fits too: with N+1 gone to N, users feel nothing yet, but the next failure is an outage. I keep that list short and written down, and I require every cause page to name what the responder does in the next few minutes.
go deeper
Know that the symptom rule has exceptions, and be able to name two — a disk about to fill and a certificate about to expire — and say why they cannot wait for user impact.
Explain the test rather than the examples: certainty of impact, lead time compared with fix time, and whether the failure gets more expensive once it lands.
Do the arithmetic out loud for a real case, and show how you keep the exception list small — written justifications, projections instead of levels, and review of what actually fired.
Own the policy boundary: decide who may add a cause page, what evidence they must supply, and how you audit the list so the exception does not quietly reconstruct the old cause-based pager.
## Why the rule needs an exception at all Symptom-based paging is right because it maximises coverage and protects the responder's attention. But it has one structural weakness: it is inherently *reactive*. A symptom alert fires when harm has already begun. For most failures that is fine, because the mitigation is fast — roll back, fail over, shed load — and the difference between detecting at minute zero and minute two is small. It is not fine when the harm, once begun, is either irreversible or takes far longer to undo than to prevent. That is the entire justification for the exception, and it is worth stating in that form rather than as a list of blessed metrics. ## The three-part test **1. Is impact certain, not merely possible?** A certificate expiring at 02:00 will break TLS handshakes at 02:00. A disk filling at a steady 2 GB/hour with 6 GB left will be full in three hours. Contrast with CPU at 85%, which may stay at 85% indefinitely. Certainty is what turns a cause into a symptom you happen to be seeing early. Note that a projection is only as good as its assumptions: a fill-rate projection assumes the rate holds, so make the projection window short enough that the assumption is defensible and prefer a trend over a static threshold, since 10% free means very different things on a 100 GB volume and a 10 TB one. **2. Is the lead time comparable to the fix time?** This is the decisive part and the one candidates skip. Compute both sides. If the certificate renewal requires a change approval that takes two hours and you find out six hours before expiry, you are inside the danger zone and that is a page. If automated renewal runs daily and failure is detected thirty days out, that is a ticket — and if it is still failing at seven days, *then* it becomes a page. Escalating a condition from ticket to page as the runway shrinks is the right shape for these alerts. **3. Is the failure worse after impact than before?** Some conditions are cheap to fix while healthy and brutal afterwards. - A full disk on a database can stop the write-ahead log, corrupt state, and block the log rotation or compaction you would use to free space. You may not be able to log the incident on the box that is having it. - Losing quorum in a consensus system is not recoverable by adding a node once the cluster is already down. - A queue with no consumer may cross a retention boundary and drop data permanently, which no amount of later capacity restores. Where the cost curve has that cliff, paying the price of an early page is rational. ## Loss of redundancy The redundancy case deserves separate mention because it is the most commonly missed one. A service designed to survive one failure (N+1) has, after one failure, no margin left. Users see nothing. Every symptom signal is green. But the system is now one event away from an outage and it will stay that way until someone restores the spare. Whether that pages depends on the same arithmetic: how long does restoring the redundancy take, how likely is a second failure in that window, and how bad is the outage if it lands? Losing one of nine stateless replicas that auto-heal in ninety seconds is a ticket. Losing one of three zones, or one of three members of a consensus group, is a page. ## Discipline that keeps the exception from eating the rule The risk with any exception is that every team decides its favourite threshold qualifies. Three practices keep it honest: - **Write the justification into the alert.** The alert's own documentation should state the certainty argument, the lead time, and the action the responder takes in the next few minutes. If nobody can write the action, the condition is not a page. - **Prefer a deadline to a threshold.** "Disk will be full in under one hour at the current rate" is a better paging condition than "disk above 90%", because the first encodes lead time and the second encodes only a level. Same for certificates: page on days remaining, not on existence. - **Review the list.** Cause pages should be a small, named set — typically a handful per service — reviewed periodically against what actually fired and what the responder actually did. A cause page that fired five times and produced no action in any of them has failed its own test. ## What to say in an interview Give the rule, then the test, then a worked example with real arithmetic: "disk fills in 25 minutes, the fix is to expand the volume or purge logs, and that takes ten to fifteen — I want a human on it now, not after the writes start failing." The arithmetic is what shows you have made this call under pressure rather than read the principle.
- Why is "disk above 90%" a worse paging condition than "disk will be full within an hour"?A level says nothing about time. Ninety percent of a 10 TB volume with a slow fill rate has days of runway; the same percentage on a small, fast-filling volume has minutes. A projection encodes the lead time directly, so the alert fires when action is genuinely required, and the same condition works across volumes of very different sizes.
- How would you treat the loss of one node in a three-node consensus cluster, given users are unaffected?Page it. With one member gone the cluster tolerates no further failure, and quorum, once lost, cannot be recovered by adding capacity during the outage. That is the cost-asymmetry case: restoring the third member while healthy is routine, restoring service after quorum loss is not.
- How do you stop the cause-page exception from quietly growing back into a cause-based pager?Keep the exception list explicit and small, require each entry to state its certainty argument, lead time and immediate action, and review what actually fired. Any cause page whose last several firings produced no human action within minutes has failed its own test and should be demoted to a ticket.
saying these in an interview costs you the question
- Treating any high-utilisation threshold as an emergency
- Paging on a static percentage rather than projected time to failure
- Ignoring that redundancy loss removes all margin without user impact
- Waiting for the symptom when the fix takes longer than the runway
- Letting each team add its own cause pages with no written justification