skip to content

A team monitoring a fleet of stateless API services builds dashboards using the RED method, while the team monitoring the underlying VMs and disks builds dashboards using the USE method. What does each method measure, and why is one better suited to request-driven services and the other to resources?

level: middleimportance: must knowfreq 75%

answer

  1. RED = Rate, Errors, Duration (per service)
  2. USE = Utilization, Saturation, Errors (per resource)
  3. Duration as percentiles, not averages
  4. Saturation catches trouble before 100% utilization
  5. RED for user-facing, USE for infrastructure underneath

basics

~20 s

RED tracks Rate, Errors, and Duration of requests a service handles, good for things that answer requests, like an API. USE tracks Utilization, Saturation, and Errors of a resource, like a CPU or a disk, good for things that don't handle requests themselves but can still run out of capacity.

solid answer

~50 s

RED (Rate, Errors, Duration) is a service-level methodology: for every request-driven component, track the rate of requests per second, the rate of failed requests, and the distribution of request duration, typically as percentiles. It maps directly onto what a caller experiences and what an SLO is usually defined against. USE (Utilization, Saturation, Errors) is a resource-level methodology: for every finite resource, CPU, memory, disk I/O, network, track how busy it is, how much extra work is queued waiting for it, and its error count. USE is better suited to infrastructure because resources like a disk don't handle requests with a duration in the RED sense, but they do have a hard capacity ceiling that RED doesn't capture directly, a service can look fine on RED metrics while its underlying CPU is saturated and about to cause queuing latency. Mature systems run both, RED for user-facing services, USE for the resources underneath them, because each catches problems the other misses.

go deeper

for a junior

Should know RED tracks traffic, errors, and speed of a service and USE tracks how busy a resource is, even without precise definitions.

for a middle

Should give the precise three signals for each method and say why Duration should be a percentile.

for a senior

Should explain why Saturation is often the earlier warning sign than Utilization, and design dashboards that pair RED, service layer, with USE, resource layer, for full-stack root-causing.

for a principal

Should set org-wide dashboard conventions so every team's RED and USE dashboards are structurally comparable, enabling fast cross-team root-causing during a multi-service incident.

## Two ways to pick what to measure RED and USE are two complementary methodologies for choosing which metrics matter, both created to answer a simpler underlying question: out of the effectively infinite number of things you could measure about a system, which handful will actually tell you it's unhealthy? | | RED | USE | |---|---|---| | Target | anything that serves requests: a service, an endpoint, a queue consumer | a finite resource with a hard capacity ceiling, such as a CPU core, a disk, a network interface, or a memory pool | | Signals | Rate, Errors, Duration | Utilization, Saturation, Errors | ## RED, at the service layer RED, popularized around monitoring Kubernetes microservices with Prometheus, targets anything that serves requests: a service, an endpoint, a queue consumer. It defines exactly three signals. - **Rate** is the number of requests the component handles per second, which establishes the baseline traffic context, a latency spike means something different at 10 requests per second than at 10,000. - **Errors** is the rate of requests that failed, typically defined by a protocol-level signal such as an HTTP 5xx status or a gRPC error code, expressed as a fraction or rate rather than a raw count so it's comparable across different traffic levels. - **Duration** is the distribution of how long requests took, almost always reported as percentiles such as p50, p95, and p99 rather than an average, because averages hide the long tail of slow requests that actually hurts users. Together, these three answer whether the service is serving traffic successfully and quickly, which is precisely the question a caller or an SLO cares about. ## USE, at the resource layer USE targets a different kind of thing: a finite resource with a hard capacity ceiling, such as a CPU core, a disk, a network interface, or a memory pool. It also defines exactly three signals, but with a different meaning. - **Utilization** is the percentage of time the resource was busy doing work, for example the percent of time a CPU was not idle. - **Saturation** is the amount of extra work the resource couldn't get to immediately, typically measured as a queue length such as the CPU run queue or the disk I/O queue, and it is the more important of the two for catching trouble early, because a resource can be at high utilization with no queuing, which is fine, or hit saturation well before 100 percent utilization on components with variable service times. - **Errors**, in this framing, means literal hardware or driver-level error counts, like disk I/O errors, distinct from RED's application-level request errors. ## Why you need both The reason both exist, rather than one method covering everything, is that they answer questions at different layers of the stack, and a problem at one layer often doesn't show up cleanly at the other, at least not immediately. A service's RED dashboard can look completely green, low error rate, healthy p99, right up until the moment the underlying VM's CPU saturates and a queue starts building, at which point RED's duration metric will start climbing, but by then the USE dashboard for that VM would already have shown rising saturation minutes earlier as an early warning. Conversely, a disk can sit at 40 percent utilization, which sounds fine, while still causing real user-facing latency if that 40 percent happens to include long, bursty writes that saturate its queue; USE metrics alone won't tell you that this is actually hurting checkout latency, only the RED dashboard for the checkout service will show that. ## The trade-off The trade-off is mostly about where you point which method rather than one method being strictly better. RED requires the component to be something that meaningfully handles requests, so it doesn't apply cleanly to a raw resource like a disk. USE requires the thing being measured to be a finite, quantifiable resource, so it doesn't map cleanly onto an abstract service's business logic, since utilization of a stateless API process isn't a well-defined single number the way CPU utilization is. Both methods deliberately keep the metric count small, three signals each, as a trade-off against completeness: this is by design, to keep dashboards actionable during an incident rather than an overwhelming wall of graphs, at the cost of not directly capturing things neither method was designed for, like business-level correctness or data quality. ## Where dashboards go wrong Common failure modes: - Using RED's Duration as an average instead of percentiles hides exactly the slow-tail requests SLOs are meant to protect. - Skipping the Rate signal on a RED dashboard means an error-rate spike caused by a traffic collapse, fewer total requests but a higher fraction failing, looks identical to a genuine failure spike, when the underlying cause is completely different. - On the USE side, watching Utilization alone and ignoring Saturation is the single most common miss, since a resource can look not that busy by utilization while already queuing work and hurting latency. ## How the pair is used in practice A concrete, real-world pairing: a team running services on Kubernetes typically builds RED dashboards per HTTP endpoint from Prometheus histograms generated by request middleware, and USE dashboards per node from node exporter metrics covering CPU, memory, and disk I/O, with the RED dashboards used to judge whether an SLO is at risk and the USE dashboards used to diagnose why.

  • Why does RED recommend percentiles for Duration instead of an average?
    An average can look healthy even when 5% of requests are taking ten times as long, because the mass of fast requests drags the mean down; percentiles like p95 or p99 directly expose that slow tail, which is usually what's actually violating an SLO and hurting the users who experience it.
  • Why is Saturation often considered the earliest warning signal in the USE method?
    A resource's queue can start growing well before utilization hits 100%, especially for resources with variable or bursty service times, so saturation trending upward is often the first visible sign of trouble, arriving before utilization looks alarming and well before user-facing latency, a RED signal, degrades.
  • Could you apply RED metrics to a database, and USE metrics to an API service?
    Yes to both, with adjustment: a database is also a request-handling component, queries in, results out, so RED metrics apply directly to it. An API service also runs on finite resources, its process's CPU and memory, so USE-style metrics can describe the process itself, even though the service's business logic isn't a resource in the USE sense.

RED is like watching a restaurant from the dining room: how many tables are served per hour, how many orders come back wrong, how long people wait for their food. USE is like watching the kitchen's equipment: how busy each burner is, how long the ticket rail is backing up, whether any appliance is throwing error codes. A dining room can look calm right up until the kitchen's ticket queue backs up and food starts arriving late.

saying these in an interview costs you the question

  • Uses Duration as a plain average rather than percentiles
  • Thinks RED and USE are interchangeable or redundant with each other
  • Applies USE's utilization/saturation framing to describe a service's business errors
  • Doesn't know Saturation is distinct from Utilization
  • Can't say which method targets services versus resources

context