skip to content

Alerting Philosophy

What deserves to wake a human: symptom-based paging, actionable alerts, and keeping signal-to-noise high. Interviewers probe this to separate engineers who have carried a pager from those who have only added alerts.

on this pageshow

questions

16

Name the four golden signals from Google's SRE monitoring guidance, and say which of them make good paging alerts for a request-serving service and which one does not.

level: juniorimportance: must knowfreq 68%

answer

  1. four words, one of them is not like the others
  2. three describe requests, one describes resources
  3. the odd one out is a leading indicator
  4. fast failures flatter your average
  5. zero traffic is itself a symptom

basics

~20 s

Latency, traffic, errors and saturation. Latency and errors are what users feel and are the natural paging signals; a traffic collapse is also user-visible. Saturation measures resource fullness — a cause, so it belongs on dashboards and tickets.

solid answer

~50 s

The four golden signals, as set out in Google's Site Reliability Engineering book, are **latency** (how long requests take, read as a distribution and not a mean), **traffic** (demand on the service, such as requests per second), **errors** (the rate of requests that failed, including those that succeeded too slowly to be useful), and **saturation** (how full the most constrained resource is). Latency and errors describe what a caller actually experienced, so they are the two I would build paging alerts on. Traffic is mostly a denominator and a context signal, but it becomes a symptom when it collapses to zero — that usually means callers cannot reach you at all. Saturation is the odd one out: it is a cause, a leading indicator of trouble rather than trouble itself, so I put it on capacity dashboards and use it for tickets, not pages.

go deeper

for a junior

Memorise latency, traffic, errors, saturation, and be able to give one concrete metric for each on an HTTP service. Say which of them you would page on.

for a middle

Explain why saturation is a cause rather than a symptom, why the errors definition needs care beyond status codes, and how the signals translate to a queue consumer or batch pipeline.

for a senior

Show the instrumentation decisions: where latency is measured, what counts as an error for this service's contract, and why a zero-traffic condition needs its own alert alongside the error ratio.

for a principal

Be ready to argue for a minimum signal set every service in the estate must publish, and to say what you gain from that consistency and what you give up by not letting each team invent its own.

## The four signals The golden signals are a checklist for instrumenting any request-serving system, published in Google's *Site Reliability Engineering* (2016) in the monitoring chapter. They are useful here as the **symptom taxonomy** — the vocabulary you use to decide what deserves a page. **Latency** — how long a request takes. The important discipline is to read it as a distribution, not an average, and to separate successful from failed requests. A fast 500 will drag your average latency *down* and make an outage look like a performance improvement, so error responses must be excluded or measured separately. **Traffic** — the demand placed on the system, in whatever unit is natural: HTTP requests per second, sessions, messages consumed per second, bytes served. It is the denominator for the other signals and the context that explains them. **Errors** — the rate of requests that failed. Failure is broader than a 5xx status. It includes responses that were technically successful but wrong, responses that arrived after the caller had already timed out, and policy failures such as a request served without the personalised content it was supposed to carry. Deciding what counts as an error for a given service is most of the work. **Saturation** — how full the service's most constrained resource is. For one service that is memory, for another it is IOPS, for another a fixed worker pool or a database connection limit. Saturation matters because most systems degrade sharply once a resource approaches full, well before it hits 100%. ## Which of these page Apply the symptom test — *would a user feel this?* — to each signal. - **Errors: page.** A failed request is user harm by definition. An alert on the ratio of failed to total requests over a short window is the single most valuable alert most services have. - **Latency: page.** Slow enough is indistinguishable from broken. Alert on a tail percentile — 95th or 99th — against a threshold derived from what the caller can tolerate, not from what the service currently does. - **Traffic: usually not, with one exception.** Rising traffic is not a fault; it is the reason you have capacity planning. But traffic falling to *zero*, or far below the value the same weekday and hour normally shows, is a symptom: it usually means a load balancer, DNS entry, upstream gateway or client release has stopped requests from arriving. A no-traffic alert also protects you from the failure mode where an error-ratio alert goes quiet simply because nothing is being served. - **Saturation: no — ticket it.** Saturation is machine state. A thread pool at 100% for ten minutes with healthy latency and no errors means the service is sized close to its demand, which is a planning conversation. The same pool at 100% *with* rising latency is already showing up in the latency signal, and that is the signal that should carry the page. ## Two related checklists Interviewers sometimes ask how the golden signals relate to other mnemonics. The RED method — rate, errors, duration — is the request-oriented subset, essentially traffic, errors and latency: it is what you apply to a *service*. The USE method — utilisation, saturation, errors — is resource-oriented and is what you apply to a *resource* such as a disk, a NIC or a CPU. The mapping is clean: RED is your paging vocabulary, USE is your diagnosis vocabulary, and the golden signals span both because they include saturation. ## Defining them for a real service The checklist is worth little until it is instantiated. For a checkout API the four might be: latency = 99th percentile of successful POST /checkout calls measured at the edge; traffic = checkout attempts per second; errors = the fraction of attempts returning 5xx or failing payment authorisation for infrastructure reasons; saturation = payment-provider connection pool occupancy plus database CPU. Notice that three of those definitions required a decision about what counts. That is the part that separates a candidate reciting four words from one who has instrumented a service. A common trap is to define errors as "5xx responses" and miss the outage where the service returned a cheerful 200 with an empty cart. ## Non-request systems The signals need translation for a pipeline or a consumer. Latency becomes end-to-end freshness or the age of the oldest unprocessed item. Traffic becomes items ingested per second. Errors becomes the fraction of items that failed or were dead-lettered. Saturation becomes consumer lag against the partition's write rate. The symptom test survives translation intact: freshness and failed items are what a downstream consumer feels; lag on its own is not, until it stops draining.

  • Why should a latency alert use a tail percentile rather than the mean?
    The mean hides the minority of requests that are badly served, and it moves the wrong way during an outage because fast error responses pull it down. A 95th or 99th percentile tells you what the worst-served slice of real users experienced, which is what they will complain about and what a service level objective is normally written against.
  • How would you apply the golden signals to an asynchronous queue consumer rather than an HTTP service?
    Translate each: latency becomes the age of the oldest unprocessed message or end-to-end freshness, traffic becomes messages consumed per second, errors becomes the fraction failed or dead-lettered, saturation becomes consumer lag relative to the producer's rate. Page on freshness and failure rate; keep lag as the diagnostic and capacity signal.
  • What is the most common mistake in defining the errors signal?
    Equating errors with 5xx status codes. Services fail while returning 200 — an empty result set, a stale cache, a page missing its personalised section, or a response that arrived after the caller timed out. Define errors from the caller's contract, not from the status code, or the alert will stay silent during a real outage.

saying these in an interview costs you the question

  • Listing saturation as a primary paging signal
  • Alerting on mean latency instead of a tail percentile
  • Treating rising traffic as a fault condition
  • Defining errors as 5xx only, missing successful-but-wrong responses
  • Reciting the four names without defining them for a real service

context

open as a page

Your team is adding a new alert for a production service and must decide whether it pages a human immediately, files a ticket, or is only recorded for later investigation. What test decides the tier, and what does each of the three tiers commit the team to?

level: middleimportance: must knowfreq 78%

basics

~20 s

Page only when a human must act within minutes and can actually change the outcome. Ticket when the work is real but can wait days. Log when nobody will act. Each tier is a response-time promise.

open as a page

A single database failover causes 200 alert notifications — one per application instance, differing only in the instance label — to hit the pager inside a minute. Explain how alert deduplication and grouping would collapse that into one notification, and what you give up as you group more aggressively.

level: middleimportance: must knowfreq 62%

basics

~20 s

Deduplication collapses byte-identical alerts from redundant senders into one. Grouping batches alerts that share a chosen set of labels into a single notification after a short wait window. Grouping harder costs detection delay and can bury an unrelated failure inside an already-acknowledged page.

open as a page

A service pages the on-call engineer whenever a host's CPU stays above 85% for five minutes, and users have never noticed anything during those pages. In alerting design, what is the difference between a symptom-based and a cause-based alert, and which of the two belongs on the pager?

level: middleimportance: must knowfreq 78%

basics

~20 s

Symptom alerts fire on what users experience — failed requests, slow responses, a service level burning down. Cause alerts fire on machine state such as CPU or disk. Page on symptoms; send cause signals to tickets and dashboards.

open as a page

You join a team whose pager delivers roughly 50 alerts a week, and most are acknowledged with no action taken. Describe the process you would run over the next month to reduce that, and how you would decide the fate of each rule.

level: seniorimportance: must knowfreq 66%

basics

~20 s

Measure before changing: per rule, count firings over the last quarter and what share led to a human action. Then apply a disposition ladder — retire, demote off the paging path, retune thresholds and durations, group or suppress the fan-out, or fix the underlying instability — and make the review recurring at every handoff.

open as a page

Your team will fail over a database at 02:00 and expects a burst of alerts. What is a maintenance silence, what does it actually stop, and why should every silence carry an expiry, a narrow scope and an owner?

level: juniorimportance: should knowfreq 48%

basics

~20 s

A maintenance silence suppresses notifications for alerts matching a set of criteria during a bounded window. It stops the notification only — rules keep evaluating and alerts stay visible. It needs an expiry so it cannot outlive the work, a narrow scope so unrelated failures still page, and an owner so someone can explain it.

open as a page

A paging rule that fires when a service's p99 latency exceeds 500 ms re-fires and resolves five times during every traffic peak. Which rule-level knobs would you reach for to damp that, and what does each one cost?

level: middleimportance: should knowfreq 52%

basics

~20 s

Require the condition to hold continuously before firing — a pending or 'for' duration — and keep the alert firing for a defined period after it clears, so brief recoveries do not close and reopen it. The first adds exactly that much detection delay; the second delays the resolved notification.

open as a page

A checkout latency alert always pages the observability platform team, because the rule lives in the dashboard folder that team maintains. Why is that routing wrong, who should receive the first page instead, and how do you keep routing following ownership as services change hands?

level: seniorimportance: should knowfreq 52%

basics

~20 s

The first page belongs to the team that can change the failing service, not the team that authored the rule. Route on an ownership attribute carried by the service and sourced from a service catalogue, so routing changes automatically when ownership does.

open as a page

A shared database primary fails. Its own alert fires, and so do the alerts of thirty dependent services, paging thirty teams at once. How would you suppress the downstream pages, and what goes wrong with dependency-based suppression?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Use inhibition: a firing source alert suppresses matching downstream alerts, scoped by labels that must be equal on both — same cluster or region — so suppression cannot leak across environments. The risk is that the encoded dependency graph drifts from reality and starts hiding genuinely independent failures.

open as a page

If the rule is "page on symptoms, not on causes", when is it nevertheless correct to page on a cause signal such as a filling disk or an expiring TLS certificate?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Page a cause when impact is near-certain and the remaining lead time is close to the time needed to fix it — a disk filling within the hour, a certificate expiring tonight. Otherwise the runway is long enough for a ticket.

open as a page

You inherit an on-call rotation whose forty-odd paging rules are almost all host-level thresholds — CPU, memory, disk, process restarts, queue depth. How do you move that alert set to symptom-based paging without going blind during the transition?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Define the service's user-visible signals first, add symptom pages for them, then run both sets side by side for a few weeks. Demote every cause rule that fired without user impact; keep the few that caught real incidents nothing else saw.

open as a page

A single database failover pages the same on-call engineer five separate times, from five alert rules written by three different teams. Suppressing the extra notifications only hides the problem — what do you change about the paging rules themselves, and how do you decide which one survives?

level: principalimportance: should knowfreq 35%

basics

~20 s

Consolidate ownership: one user-visible symptom should have exactly one paging rule, owned by one team. The other four become ticket- or log-tier diagnostic signals. Duplicate paging is a rule-inventory and ownership problem, not a notification-filter problem.

open as a page

A paging alert on your service fired 20 times last month, and only 3 of those pages led a responder to take any action. Express that as the alert's precision, explain what you give up when you tighten the rule to improve it, and say how you would decide where to sit on that trade.

level: seniorimportance: nice to knowfreq 40%

basics

~20 s

Precision is 3/20, or 15% — roughly six of every seven pages were noise. Tightening the rule raises precision but lowers recall, missing real events. Paging tier should favour precision; lower tiers can absorb lower precision.

open as a page

You are responsible for reliability standards across 60 teams holding roughly 4,000 alert rules, and you have no authority to edit another team's rules. How would you drive pager noise down across the whole organisation?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Make noise measurable per team and visible, ship good behaviour as platform defaults rather than mandates, set a page-load budget teams own themselves, and pair it with a detection counter-metric so nobody hits the budget by going blind.

open as a page

Where does a strictly symptom-based paging policy break down across a large service estate, and what would you mandate instead of a flat "page only on user-visible symptoms" rule?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Symptom paging is reactive, so it fails where impact arrives late, where traffic is too thin to measure, and where redundancy hides a fault until the next one lands. Mandate symptom coverage as a floor, and permit cause pages by exception with written justification.

open as a page