skip to content

Name the four golden signals from Google's SRE monitoring guidance, and say which of them make good paging alerts for a request-serving service and which one does not.

level: juniorimportance: must knowfreq 68%

answer

  1. four words, one of them is not like the others
  2. three describe requests, one describes resources
  3. the odd one out is a leading indicator
  4. fast failures flatter your average
  5. zero traffic is itself a symptom

basics

~20 s

Latency, traffic, errors and saturation. Latency and errors are what users feel and are the natural paging signals; a traffic collapse is also user-visible. Saturation measures resource fullness — a cause, so it belongs on dashboards and tickets.

solid answer

~50 s

The four golden signals, as set out in Google's Site Reliability Engineering book, are **latency** (how long requests take, read as a distribution and not a mean), **traffic** (demand on the service, such as requests per second), **errors** (the rate of requests that failed, including those that succeeded too slowly to be useful), and **saturation** (how full the most constrained resource is). Latency and errors describe what a caller actually experienced, so they are the two I would build paging alerts on. Traffic is mostly a denominator and a context signal, but it becomes a symptom when it collapses to zero — that usually means callers cannot reach you at all. Saturation is the odd one out: it is a cause, a leading indicator of trouble rather than trouble itself, so I put it on capacity dashboards and use it for tickets, not pages.

go deeper

for a junior

Memorise latency, traffic, errors, saturation, and be able to give one concrete metric for each on an HTTP service. Say which of them you would page on.

for a middle

Explain why saturation is a cause rather than a symptom, why the errors definition needs care beyond status codes, and how the signals translate to a queue consumer or batch pipeline.

for a senior

Show the instrumentation decisions: where latency is measured, what counts as an error for this service's contract, and why a zero-traffic condition needs its own alert alongside the error ratio.

for a principal

Be ready to argue for a minimum signal set every service in the estate must publish, and to say what you gain from that consistency and what you give up by not letting each team invent its own.

## The four signals The golden signals are a checklist for instrumenting any request-serving system, published in Google's *Site Reliability Engineering* (2016) in the monitoring chapter. They are useful here as the **symptom taxonomy** — the vocabulary you use to decide what deserves a page. **Latency** — how long a request takes. The important discipline is to read it as a distribution, not an average, and to separate successful from failed requests. A fast 500 will drag your average latency *down* and make an outage look like a performance improvement, so error responses must be excluded or measured separately. **Traffic** — the demand placed on the system, in whatever unit is natural: HTTP requests per second, sessions, messages consumed per second, bytes served. It is the denominator for the other signals and the context that explains them. **Errors** — the rate of requests that failed. Failure is broader than a 5xx status. It includes responses that were technically successful but wrong, responses that arrived after the caller had already timed out, and policy failures such as a request served without the personalised content it was supposed to carry. Deciding what counts as an error for a given service is most of the work. **Saturation** — how full the service's most constrained resource is. For one service that is memory, for another it is IOPS, for another a fixed worker pool or a database connection limit. Saturation matters because most systems degrade sharply once a resource approaches full, well before it hits 100%. ## Which of these page Apply the symptom test — *would a user feel this?* — to each signal. - **Errors: page.** A failed request is user harm by definition. An alert on the ratio of failed to total requests over a short window is the single most valuable alert most services have. - **Latency: page.** Slow enough is indistinguishable from broken. Alert on a tail percentile — 95th or 99th — against a threshold derived from what the caller can tolerate, not from what the service currently does. - **Traffic: usually not, with one exception.** Rising traffic is not a fault; it is the reason you have capacity planning. But traffic falling to *zero*, or far below the value the same weekday and hour normally shows, is a symptom: it usually means a load balancer, DNS entry, upstream gateway or client release has stopped requests from arriving. A no-traffic alert also protects you from the failure mode where an error-ratio alert goes quiet simply because nothing is being served. - **Saturation: no — ticket it.** Saturation is machine state. A thread pool at 100% for ten minutes with healthy latency and no errors means the service is sized close to its demand, which is a planning conversation. The same pool at 100% *with* rising latency is already showing up in the latency signal, and that is the signal that should carry the page. ## Two related checklists Interviewers sometimes ask how the golden signals relate to other mnemonics. The RED method — rate, errors, duration — is the request-oriented subset, essentially traffic, errors and latency: it is what you apply to a *service*. The USE method — utilisation, saturation, errors — is resource-oriented and is what you apply to a *resource* such as a disk, a NIC or a CPU. The mapping is clean: RED is your paging vocabulary, USE is your diagnosis vocabulary, and the golden signals span both because they include saturation. ## Defining them for a real service The checklist is worth little until it is instantiated. For a checkout API the four might be: latency = 99th percentile of successful POST /checkout calls measured at the edge; traffic = checkout attempts per second; errors = the fraction of attempts returning 5xx or failing payment authorisation for infrastructure reasons; saturation = payment-provider connection pool occupancy plus database CPU. Notice that three of those definitions required a decision about what counts. That is the part that separates a candidate reciting four words from one who has instrumented a service. A common trap is to define errors as "5xx responses" and miss the outage where the service returned a cheerful 200 with an empty cart. ## Non-request systems The signals need translation for a pipeline or a consumer. Latency becomes end-to-end freshness or the age of the oldest unprocessed item. Traffic becomes items ingested per second. Errors becomes the fraction of items that failed or were dead-lettered. Saturation becomes consumer lag against the partition's write rate. The symptom test survives translation intact: freshness and failed items are what a downstream consumer feels; lag on its own is not, until it stops draining.

  • Why should a latency alert use a tail percentile rather than the mean?
    The mean hides the minority of requests that are badly served, and it moves the wrong way during an outage because fast error responses pull it down. A 95th or 99th percentile tells you what the worst-served slice of real users experienced, which is what they will complain about and what a service level objective is normally written against.
  • How would you apply the golden signals to an asynchronous queue consumer rather than an HTTP service?
    Translate each: latency becomes the age of the oldest unprocessed message or end-to-end freshness, traffic becomes messages consumed per second, errors becomes the fraction failed or dead-lettered, saturation becomes consumer lag relative to the producer's rate. Page on freshness and failure rate; keep lag as the diagnostic and capacity signal.
  • What is the most common mistake in defining the errors signal?
    Equating errors with 5xx status codes. Services fail while returning 200 — an empty result set, a stale cache, a page missing its personalised section, or a response that arrived after the caller timed out. Define errors from the caller's contract, not from the status code, or the alert will stay silent during a real outage.

saying these in an interview costs you the question

  • Listing saturation as a primary paging signal
  • Alerting on mean latency instead of a tail percentile
  • Treating rising traffic as a fault condition
  • Defining errors as 5xx only, missing successful-but-wrong responses
  • Reciting the four names without defining them for a real service

context