skip to content

Your team proposes CPU utilization and cache hit rate as the SLIs for a user-facing API. Why are those poor choices for a service level indicator, and what test does a candidate SLI have to pass before you adopt it?

level: juniorimportance: must knowfreq 72%

answer

  1. does a user feel this number?
  2. proportional to user happiness
  3. saturation signals are diagnostics
  4. start from journeys, not endpoints
  5. high CPU, everyone happy

basics

~20 s

A good SLI tracks what a user experiences — success, latency, freshness — and moves with user happiness. CPU and cache hit rate are internal resource metrics: they spike while users are fine, and stay flat while users suffer.

solid answer

~50 s

An SLI has to stand in for the question "are users getting what they came for right now?", so the test I apply is proportionality: as the number degrades, real user experience must degrade with it, and when the number is healthy users must actually be fine. CPU utilization fails in both directions — a service tuned to run hot sits at 90% CPU with everyone happy, and a deadlocked thread pool sits at 5% CPU while serving nothing. Cache hit rate is the same: it is a cause, not an experience. Both are genuinely useful as saturation and debugging signals, they just cannot carry a service level, because no target on them promises a user anything. Instead I start from the critical user journeys — browse, search, checkout — and for each pick from the small menu of user-visible indicator types: availability, latency, correctness, freshness, durability. Usually one availability and one latency indicator per journey is enough.

go deeper

for a junior

Be ready to say plainly that an SLI measures what the user experiences — success, speed, freshness — and that CPU or cache hit rate is an internal metric. Name at least two real indicator types rather than describing the idea abstractly.

for a middle

Explain the proportionality test in both directions: the indicator must degrade when users suffer and stay healthy when they are fine. Give a concrete case where CPU is high and users are happy, and one where CPU is low and the service is dead.

for a senior

Show that you would start from critical user journeys and separate low-volume, high-value paths like checkout from high-volume browse traffic, so one cannot hide inside the other. Say what you would do about a journey whose failure is invisible to status codes.

for a principal

Own the tradeoff between an indicator that is cheap and uniform across the estate and one that genuinely reflects a particular product's experience. Be able to say when you would fund building a correctness or freshness indicator that does not exist yet, and when the generic one is good enough.

## What an SLI is actually for A service level indicator is a single measured number that stands in for the question an operator can never ask directly: *are the people using this service getting what they came for right now?* Everything else about service levels — the target you set on it, the budget you compute from it, the alerts you build on it — inherits the quality of that stand-in. Choose the indicator badly and every layer above it is measuring the wrong thing very precisely. ## The proportionality test The test to apply to any candidate indicator is proportionality: **as the number gets worse, user happiness must get worse with it, and while the number is healthy, users must be fine.** Both directions matter, and each has a distinct failure mode. - *False alarm* — the number degrades while nobody is affected. A batch-heavy service is often designed to run at 90–95% CPU; that is the machine doing its job, not an outage. Paging on it teaches the team that the indicator lies. - *False comfort* — the number looks fine while users are locked out. A service whose thread pool is deadlocked, whose downstream dependency is timing out, or whose process has stopped accepting connections can show *low* CPU and a *high* cache hit rate precisely because it is serving nothing. CPU utilization, memory, cache hit rate, queue depth, GC pause time and connection-pool usage all fail this test. That does not make them worthless — they are exactly what you want when diagnosing an incident or planning capacity — but they belong to the diagnostic and saturation layer, not to the service level. The giveaway is that you cannot state a target on them that promises a user anything: "CPU below 70%" is not a commitment any customer can understand or care about. ## The small menu of user-visible indicator types In practice almost every SLI is one of a handful of shapes: - **Availability** — the proportion of requests the service handled successfully rather than failing. - **Latency** — the proportion of requests served faster than a stated threshold. - **Correctness / quality** — the proportion of responses that were right, not merely returned: search results that are not empty, a recommendation payload that is not the degraded fallback, a computed total that reconciles. - **Freshness** — for anything derived or cached, the proportion of data served that was updated recently enough to be useful. - **Durability** — for storage, the proportion of stored objects still retrievable and intact. - **Coverage / throughput** — the proportion of the offered work the system actually got through. Picking from this menu is most of the job. The remaining judgment is which journey each indicator belongs to and where you measure it. ## Start from journeys, not from endpoints The common mistake after abandoning CPU is to swing to one indicator per HTTP endpoint, which produces forty numbers nobody looks at. Instead enumerate the **critical user journeys** — the handful of things a user came to do — and give each one or two indicators. For a storefront that might be: - *Browse catalogue*: availability and latency of the catalogue read path. - *Checkout*: availability and latency of the order-submission path, plus a correctness indicator that the order was actually persisted. - *Order status*: freshness of the status data, because a stale-but-fast answer is a failure to this user even though the request succeeded. Note that checkout and browse deserve **separate** indicators even though they run in the same binary. Checkout is lower volume and higher value; folding it into a single service-wide ratio lets a total checkout failure hide inside a healthy browse number. The rule of thumb is one indicator per journey per *dimension that can fail independently*. ## A worked selection For a payments API, the candidate list would be: - Availability: proportion of payment-submission requests that did not return a server error. - Latency: proportion of those requests completed under an explicit threshold, chosen from what the checkout page can tolerate before users abandon. - Correctness: proportion of accepted payments that reconcile against the ledger within the settlement window. CPU, cache hit rate and queue depth stay on the dashboards as capacity and diagnostic signals — and note that when the cache hit rate collapses, the latency indicator moves on its own, which is the point. The user-visible indicator catches the *class* of problems the resource metric only catches one instance of. ## What interviewers are checking The substance of this question is whether you reach for what is easy to measure or for what a user feels. Every service already emits CPU; almost none emit a checkout-journey success ratio until someone deliberately builds it. Naming that gap, and naming the journeys rather than the endpoints, is the answer.

  • If CPU makes a bad SLI, is there any place for it in how you run the service?
    Yes — as a saturation and capacity signal. It tells you how much headroom is left and it is often the first thing you look at during an incident. The distinction is that it informs provisioning and diagnosis, while the service level is stated in terms a user would recognise. Keep it graphed; just do not set a service commitment on it.
  • How many SLIs should one service have?
    Few enough that a human actually watches them — typically one or two per critical journey, so a handful per service. Every extra indicator dilutes attention and adds a definition somebody must maintain. If a journey can fail independently and a user would notice, it earns its own indicator; if it fails only when another journey already has, fold it in.
  • A service returns HTTP 200 with an empty result set whenever its search backend is down. Which indicator type catches that?
    A correctness or quality indicator, because availability by status code will read 100%. You need something that inspects the response — proportion of searches returning a non-empty, non-fallback result — or a business signal such as the conversion rate on the search page. This is the classic case where status codes alone are not a service level.

saying these in an interview costs you the question

  • Anything we already graph can serve as an SLI
  • More SLIs mean better coverage of the service
  • High CPU causes slowness, so CPU is a fine proxy
  • Host uptime is the same thing as user availability
  • One service-wide success ratio covers every journey

context