skip to content

Where does a strictly symptom-based paging policy break down across a large service estate, and what would you mandate instead of a flat "page only on user-visible symptoms" rule?

level: principalimportance: nice to knowfreq 32%

answer

  1. reactive by construction
  2. late impact, thin traffic, silent success, lost margin
  3. redefine the symptom before abandoning the rule
  4. mandate a floor, not a prohibition
  5. audit what fired, do not gatekeep what exists

basics

~20 s

Symptom paging is reactive, so it fails where impact arrives late, where traffic is too thin to measure, and where redundancy hides a fault until the next one lands. Mandate symptom coverage as a floor, and permit cause pages by exception with written justification.

solid answer

~50 s

The rule holds for high-traffic request paths and breaks in four places: asynchronous work where the user feels the failure hours later, low-traffic or internal services where ratios are statistically meaningless, silent failures such as data corruption or a stale cache that return success, and degraded redundancy where nothing is user-visible until the next fault. So I would not mandate the rule as an absolute. I would mandate a *floor* — every user-facing journey must have symptom coverage measured at its boundary — and then allow cause pages by exception, where the owner writes down the certainty argument, the lead time versus the fix time, and the action the responder takes. The exception list gets reviewed against what actually fired. That gives me a consistent minimum without pretending every system is a synchronous HTTP service.

go deeper

for a junior

Know that symptom-based paging is a default rather than a law, and that asynchronous or very low-traffic services need their symptoms defined differently.

for a middle

Name the four breakdown cases — delayed impact, thin traffic, silent success, lost redundancy — and give the redefined symptom for at least the first, such as data freshness.

for a senior

Show how you would cover each gap concretely: freshness signals, synthetic probes, correctness checks, and a short justified list of cause pages for imminent conditions.

for a principal

Own the policy design and its cost: what you centralise (the coverage floor and the audit), what you leave to teams, and which two numbers tell you the policy is drifting in either direction.

## The rule's blind spots "Page on symptoms" is the right default because it maximises coverage per alert and protects responder attention. As an absolute org-wide mandate it fails in four recognisable places, and a principal-level answer names them before defending the policy. **Delayed impact.** In an asynchronous pipeline, the gap between the fault and the user's experience can be hours. A consumer stops at 22:00; the report is wrong at 07:00; the customer notices at 09:00. Paging on the user-visible symptom means paging eleven hours after the point where the fix was cheap — and possibly after retention has discarded the data. The fix is not to abandon symptom alerting but to *redefine the symptom*: the age of the oldest unprocessed item, or end-to-end data freshness, is genuinely user-facing and available immediately. **Thin traffic.** Error ratios need volume. On a service taking a few requests a minute, one failure is a large fraction and one quiet hour is indistinguishable from an outage. Options are synthetic probes to create measurable traffic, absolute counts over longer windows, or an explicit decision that this service tolerates slower detection. **Silent failure.** Some faults return success. A cache serving stale data, a partial index, a feature returning empty results, a replication stream that has quietly stopped. No latency or error signal moves. Detecting these needs correctness or freshness probes — a different class of signal, still symptom-shaped, but one that has to be deliberately built rather than derived from request telemetry. **Degraded redundancy.** A system designed to survive one failure looks perfectly healthy after its first. Every symptom is green and the margin is gone. Whether that pages depends on how long restoration takes and what the second failure costs. ## What I would actually mandate The policy I would write has three parts, deliberately asymmetric. **A floor that is not negotiable.** Every user-facing journey has symptom coverage measured at its boundary — a failure signal and a latency or freshness signal at minimum — and that coverage is a launch requirement, not a follow-up ticket. This is the part worth centralising, because it is what makes services comparable and what lets an incident responder from another team read a service they have never seen. **An exception path with a cost.** Cause pages are allowed, and the cost of adding one is a written justification: why impact is certain, what the lead time is compared with the repair time, and what the responder does in the first five minutes. The requirement is not bureaucracy for its own sake — most proposed cause pages fail on the third item, and having to write it is what surfaces that. **An audit, not an approval board.** Rather than gatekeeping additions, review outcomes: which cause pages fired, and what the responder did. A cause page whose last several firings produced no action is demoted automatically. This scales far better than a central team approving alert rules, and it is much harder to argue with. ## What the mandate deliberately does not fix Two honest limits worth stating. First, symptom-only paging can make diagnosis *slower* if teams read "do not page on causes" as "do not instrument causes". The policy must say explicitly that cause telemetry is required and simply routed elsewhere. Detection and diagnosis are separate jobs. Second, a uniform rule imposed on a heterogeneous estate produces malicious compliance — teams write a symptom alert that technically satisfies the floor and keep their old thresholds beside it. That is why the floor is stated as an outcome (a journey is covered) rather than as a rule count, and why the audit looks at what fired rather than at what exists. ## The organisational tradeoff There is a real cost on both sides, and naming it is the point of the question. Centralising the floor buys consistency, cross-team readability and a defensible on-call load; it costs teams autonomy and will occasionally be wrong for an unusual system. Leaving it entirely to teams buys local fit; it costs you an estate where no two rotations are comparable and where the worst pager quietly stays the worst. My resolution is to centralise the *minimum* and the *review*, and to decentralise everything above it. Teams may always add more; they may not go below the floor without an exception on record. And the thing I would measure to know whether the policy is working is not the alert count — it is the share of incidents detected by an alert rather than by a customer, and the share of pages that led to an action. If the first falls, the floor is too thin. If the second falls, the exception list has grown back into a cause-based pager.

  • A team says symptom paging does not fit their nightly batch job. How do you respond?
    Push back on the definition, not the principle. The symptom for batch work is freshness or the age of the oldest unprocessed item, both measurable while the job is still running and both genuinely user-facing. If the job's only observable outcome truly is tomorrow's report, then a completion-deadline page is the justified exception — documented with its lead time and its first action.
  • How would you detect a fault that returns successful responses, such as a cache serving stale data?
    Request telemetry will not show it, so build a correctness or freshness signal deliberately: probe a known value and assert the answer, compare a sampled result against the source of truth, or track the age of the data being served. These are still symptom signals — they measure what the user gets — but they have to be designed rather than derived.
  • What single number tells you the policy has failed?
    The share of incidents first reported by a customer rather than by an alert. If that rises, the symptom floor is too thin somewhere and the exception path is not compensating. A second number guards the other direction: the share of pages that produced an action within minutes, which falls when cause pages have quietly accumulated again.

saying these in an interview costs you the question

  • Applying request-shaped symptom rules unchanged to batch and streaming systems
  • Reading 'do not page on causes' as 'do not instrument causes'
  • Assuming an error ratio is meaningful at very low request volume
  • Believing every failure eventually shows up in latency or error rate
  • Approving cause pages centrally instead of auditing what they produced

context