skip to content

RED and USE Checklists

The instrumentation checklists that tell you which metrics a service or resource needs before you ship it. Interviewers ask 'what would you instrument first?' and expect RED, USE, or the golden signals by name.

on this pageshow

questions

3

Under the RED checklist, how do you instrument a request-serving service, and what counts as an error?

level: middleimportance: must knowfreq 64%

answer

  1. Three measurements, one request-serving component
  2. Rate and errors share one counting event
  3. Duration recorded as a distribution
  4. The verdict belongs in a dimension

basics

~20 s

RED means emitting three things for a request-serving component: request rate, failed requests, and a duration distribution. Rate and errors come from one counter carrying an outcome dimension, and "error" must be defined in writing before the count means anything.

solid answer

~50 s

**RED** is a per-component instrumentation checklist for anything that answers requests: emit the **rate** of requests, the **errors** among them, and each one's **duration**. In practice rate and errors are one instrument, not two — a single counter incremented once per completed request, carrying an outcome dimension — because a ratio built from two independently maintained counters drifts the first time a code path increments one and forgets the other. Duration is recorded as a distribution rather than a mean. The hard part is *errors*, which must be defined before it is counted. A 5xx is obvious; a client fault usually is not the component's failure; a timed-out request was often never counted at all because the handler never returned; a `200` carrying a failure payload counts for nothing unless the handler says so. Emit the classification as a dimension so consumers can redraw the line later.

code

text · 4 lines
text
turbine_schedule_requests_total{route="/plans/{planId}",outcome="ok"} 1842317
turbine_schedule_requests_total{route="/plans/{planId}",outcome="client_error"} 41208
turbine_schedule_requests_total{route="/plans/{planId}",outcome="server_error"} 7431
turbine_schedule_requests_total{route="/plans/{planId}",outcome="shed"} 2196

go deeper

for a junior

Be ready to name the three RED measurements — rate, errors, duration — and to say they apply to anything that serves requests. Knowing that "error" is a decision somebody makes, rather than a fact the framework hands you, already puts you ahead.

for a middle

Explain that rate and errors come from a single counting event carrying an outcome dimension, that duration is recorded as a distribution rather than a mean, and what happens to a request whose handler never returns.

for a senior

Show that you have argued about the definition in production. Talk about where the clock starts, how shed and rate-limited requests are classified, and how you keep components across a fleet measuring at the same boundary.

for a principal

Own the estate-level decision: one outcome vocabulary that every service uses, delivered by a shared instrumentation library rather than by convention, so two teams' numbers are comparable and changing collection tooling cannot silently redefine the error rate.

## What the checklist actually asks for **RED** is an instrumentation checklist, not a dashboard layout. It applies to any component that receives a request and returns a response — an HTTP handler, a gRPC method, an internal service call — and says that before the component ships it must emit three things: the **rate** at which requests complete, the **errors** among them, and the **duration** each took. The value of a checklist is that it is mechanical: you apply the same three to every request-serving component in the estate, so that a query, a dashboard and a new engineer's mental model transfer unchanged between services. That transferability is also why the definitions must be nailed down — three services that each interpret "error" their own way produce three numbers that share a name and mean different things. ## Rate and errors are one instrument, not two Rate is completed requests per unit of time. Errors is the subset of those requests that failed. That subset relationship decides how you emit them, because the number anyone actually reads is the **ratio** of the second to the first, and a ratio is only trustworthy when its numerator and its denominator come from the same counting event. So: one counter, incremented exactly once per completed request, carrying a dimension that records how the request ended. - **Rate** is that counter summed across every value of the outcome dimension. - **Errors** is the same counter filtered to the failing values. - Nothing can drift, because there is only one increment. Two separately maintained counters — one named for requests, one for errors — diverge the first time a code path increments one and forgets the other, and the divergence is undetectable from outside because both numbers still look plausible. | Signal | What you emit | What it is attached to | The usual mistake | |---|---|---|---| | Rate | one counter per completed request | the route template the framework matched | counting arrivals rather than completions, so shed load vanishes | | Errors | the same counter, filtered on its outcome dimension | the same route template | a second counter maintained by hand, which drifts | | Duration | a distribution recorded per request | the same route template, plus outcome | recording a mean, from which no tail can be recovered | ## "Error" is a decision, not a fact The framework hands you a status code. It does not hand you a verdict. At least five different things get called an error, and they behave differently: 1. **Server faults.** The handler threw, a dependency was unreachable, the response was a 5xx. Nobody argues about these. 2. **Client faults.** A malformed body, a missing field, an unauthenticated call. These are usually the caller's defect rather than the component's, and folding them into the same number makes your error rate track your callers' bugs instead of your own. 3. **Requests that never returned.** The caller timed out, the connection dropped, the process was replaced mid-request. The handler never reached the increment, so from inside the component the request does not exist at all — the caller counted a failure and the server counted nothing. 4. **Successful-looking failures.** A `200` whose body says the plan could not be computed. The transport layer cannot know this failed; only the handler can, and only if it says so. 5. **Deliberate rejections.** Load shedding, rate limiting, an open circuit breaker. These are failures for the caller and correct behaviour for the component, and which of those you mean decides where they are counted. The instrumentation answer to all five is the same: **do not bake the verdict into the metric's name.** A series whose name ends in something like `_errors_total` has permanently decided which of the five count, for every consumer, forever. Emit the outcome as a dimension with enough resolution to keep the five separable, and let each reader draw the line where their question needs it. Redrawing a line is a query change; re-deciding a metric name is a redeploy of every service. ## Where the clock starts and stops Duration hides a definition too. Measured at the framework's outermost server-side filter it includes routing, deserialization, authentication and any wait for a worker; measured inside the business method it includes none of that. Neither boundary is wrong, but a fleet whose components disagree about it produces a fleet-wide latency figure that means nothing. Pick one boundary and write it down beside the outcome vocabulary. ## What it looks like when the definition moves A wind-farm maintenance planner serves roughly 1,842 scheduling requests per second across 2,311 turbines. Mid-quarter the team migrated off a hosted monitoring vendor and re-emitted RED from inside the application rather than from the vendor's proxy-level integration. The reported error ratio fell from 2.7% to 0.4% overnight without one line of application code changing: the proxy had counted every 4xx as an error and had counted requests the application never completed, while the new in-process counter incremented only on handler return and classified client faults under their own outcome value. Nothing improved; one definition replaced another. And because the old metric had baked its verdict into its name, the old reading could not be reconstructed from the new data — the practical argument for the outcome dimension, made by an estate that lacked one.

  • Your request counter increments when the handler returns. Which requests does that miss?
    Everything that never reaches the return statement: a caller that disconnected, a request killed when the process was replaced, a handler still blocked past the caller's timeout. The caller counted those as failures and the component counted nothing at all. Incrementing from a completion hook that also runs on abnormal termination, with its own outcome value, closes most of the gap; the remainder is only visible from the caller's own RED metrics.
  • Does a retried request count as one failure or several in a component's RED metrics?
    Several. Each attempt is a real request the component served, so its counter increments per attempt — three retries of one logical operation are three requests and up to three errors. That is correct for the component and misleading about the user's experience, so the caller instruments the logical operation separately, emitting one outcome per operation rather than per attempt. Reporting only the component's view makes a retry storm look like a traffic increase.
  • Where in the request path should RED instrumentation sit so a whole fleet agrees?
    At the outermost server-side filter or interceptor you control — outside routing, authentication and deserialization — so that malformed, rejected and shed requests are counted rather than invisible. Instrumenting inside the business method makes each component's numbers depend on how much of its own plumbing it happens to wrap, and a fleet-wide latency figure assembled from components measuring at different boundaries means nothing. Comparability comes from choosing one boundary and writing it down.

Counting errors without defining one is like a factory reporting defects with no written spec: the number moves when the inspector changes, not when the product does.

saying these in an interview costs you the question

  • Counts only 5xx responses and calls everything else a success
  • Keeps a separate errors counter that drifts from the request counter
  • Reports duration as an average response time
  • Counts requests on arrival, so shed load disappears from the rate
  • Bakes the error verdict into the metric name instead of a dimension
  • Assumes a timed-out request was counted by the component that served it
open as a page

In the USE method, what is saturation for a thread pool or a work queue, and why is utilization alone a weak overload signal?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Saturation is work a resource has accepted but cannot serve yet: queued tasks, queue depth, wait time before starting, rejections. Utilization is capped at one hundred percent, so it stops resolving exactly when pressure starts growing.

open as a page

Neither RED nor USE fits a batch job or an event consumer. What do you instrument on those components instead?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Instrument what the component owes someone: for a batch job, the time of its last successful completion; for an event consumer, the age of the oldest unprocessed item; for both, items processed, failed and skipped.

open as a page