skip to content

Service Levels & Error Budgets

The measurement backbone of SRE: SLIs, SLOs, SLAs, and the error budgets that turn reliability into a negotiable engineering resource. A staple of SRE interviews because it tests whether you can define 'reliable enough' quantitatively instead of chasing 100%.

on this pageshow

questions

22

Your team proposes CPU utilization and cache hit rate as the SLIs for a user-facing API. Why are those poor choices for a service level indicator, and what test does a candidate SLI have to pass before you adopt it?

level: juniorimportance: must knowfreq 72%

answer

  1. does a user feel this number?
  2. proportional to user happiness
  3. saturation signals are diagnostics
  4. start from journeys, not endpoints
  5. high CPU, everyone happy

basics

~20 s

A good SLI tracks what a user experiences — success, latency, freshness — and moves with user happiness. CPU and cache hit rate are internal resource metrics: they spike while users are fine, and stay flat while users suffer.

solid answer

~50 s

An SLI has to stand in for the question "are users getting what they came for right now?", so the test I apply is proportionality: as the number degrades, real user experience must degrade with it, and when the number is healthy users must actually be fine. CPU utilization fails in both directions — a service tuned to run hot sits at 90% CPU with everyone happy, and a deadlocked thread pool sits at 5% CPU while serving nothing. Cache hit rate is the same: it is a cause, not an experience. Both are genuinely useful as saturation and debugging signals, they just cannot carry a service level, because no target on them promises a user anything. Instead I start from the critical user journeys — browse, search, checkout — and for each pick from the small menu of user-visible indicator types: availability, latency, correctness, freshness, durability. Usually one availability and one latency indicator per journey is enough.

go deeper

for a junior

Be ready to say plainly that an SLI measures what the user experiences — success, speed, freshness — and that CPU or cache hit rate is an internal metric. Name at least two real indicator types rather than describing the idea abstractly.

for a middle

Explain the proportionality test in both directions: the indicator must degrade when users suffer and stay healthy when they are fine. Give a concrete case where CPU is high and users are happy, and one where CPU is low and the service is dead.

for a senior

Show that you would start from critical user journeys and separate low-volume, high-value paths like checkout from high-volume browse traffic, so one cannot hide inside the other. Say what you would do about a journey whose failure is invisible to status codes.

for a principal

Own the tradeoff between an indicator that is cheap and uniform across the estate and one that genuinely reflects a particular product's experience. Be able to say when you would fund building a correctness or freshness indicator that does not exist yet, and when the generic one is good enough.

## What an SLI is actually for A service level indicator is a single measured number that stands in for the question an operator can never ask directly: *are the people using this service getting what they came for right now?* Everything else about service levels — the target you set on it, the budget you compute from it, the alerts you build on it — inherits the quality of that stand-in. Choose the indicator badly and every layer above it is measuring the wrong thing very precisely. ## The proportionality test The test to apply to any candidate indicator is proportionality: **as the number gets worse, user happiness must get worse with it, and while the number is healthy, users must be fine.** Both directions matter, and each has a distinct failure mode. - *False alarm* — the number degrades while nobody is affected. A batch-heavy service is often designed to run at 90–95% CPU; that is the machine doing its job, not an outage. Paging on it teaches the team that the indicator lies. - *False comfort* — the number looks fine while users are locked out. A service whose thread pool is deadlocked, whose downstream dependency is timing out, or whose process has stopped accepting connections can show *low* CPU and a *high* cache hit rate precisely because it is serving nothing. CPU utilization, memory, cache hit rate, queue depth, GC pause time and connection-pool usage all fail this test. That does not make them worthless — they are exactly what you want when diagnosing an incident or planning capacity — but they belong to the diagnostic and saturation layer, not to the service level. The giveaway is that you cannot state a target on them that promises a user anything: "CPU below 70%" is not a commitment any customer can understand or care about. ## The small menu of user-visible indicator types In practice almost every SLI is one of a handful of shapes: - **Availability** — the proportion of requests the service handled successfully rather than failing. - **Latency** — the proportion of requests served faster than a stated threshold. - **Correctness / quality** — the proportion of responses that were right, not merely returned: search results that are not empty, a recommendation payload that is not the degraded fallback, a computed total that reconciles. - **Freshness** — for anything derived or cached, the proportion of data served that was updated recently enough to be useful. - **Durability** — for storage, the proportion of stored objects still retrievable and intact. - **Coverage / throughput** — the proportion of the offered work the system actually got through. Picking from this menu is most of the job. The remaining judgment is which journey each indicator belongs to and where you measure it. ## Start from journeys, not from endpoints The common mistake after abandoning CPU is to swing to one indicator per HTTP endpoint, which produces forty numbers nobody looks at. Instead enumerate the **critical user journeys** — the handful of things a user came to do — and give each one or two indicators. For a storefront that might be: - *Browse catalogue*: availability and latency of the catalogue read path. - *Checkout*: availability and latency of the order-submission path, plus a correctness indicator that the order was actually persisted. - *Order status*: freshness of the status data, because a stale-but-fast answer is a failure to this user even though the request succeeded. Note that checkout and browse deserve **separate** indicators even though they run in the same binary. Checkout is lower volume and higher value; folding it into a single service-wide ratio lets a total checkout failure hide inside a healthy browse number. The rule of thumb is one indicator per journey per *dimension that can fail independently*. ## A worked selection For a payments API, the candidate list would be: - Availability: proportion of payment-submission requests that did not return a server error. - Latency: proportion of those requests completed under an explicit threshold, chosen from what the checkout page can tolerate before users abandon. - Correctness: proportion of accepted payments that reconcile against the ledger within the settlement window. CPU, cache hit rate and queue depth stay on the dashboards as capacity and diagnostic signals — and note that when the cache hit rate collapses, the latency indicator moves on its own, which is the point. The user-visible indicator catches the *class* of problems the resource metric only catches one instance of. ## What interviewers are checking The substance of this question is whether you reach for what is easy to measure or for what a user feels. Every service already emits CPU; almost none emit a checkout-journey success ratio until someone deliberately builds it. Naming that gap, and naming the journeys rather than the endpoints, is the answer.

  • If CPU makes a bad SLI, is there any place for it in how you run the service?
    Yes — as a saturation and capacity signal. It tells you how much headroom is left and it is often the first thing you look at during an incident. The distinction is that it informs provisioning and diagnosis, while the service level is stated in terms a user would recognise. Keep it graphed; just do not set a service commitment on it.
  • How many SLIs should one service have?
    Few enough that a human actually watches them — typically one or two per critical journey, so a handful per service. Every extra indicator dilutes attention and adds a definition somebody must maintain. If a journey can fail independently and a user would notice, it earns its own indicator; if it fails only when another journey already has, fold it in.
  • A service returns HTTP 200 with an empty result set whenever its search backend is down. Which indicator type catches that?
    A correctness or quality indicator, because availability by status code will read 100%. You need something that inspects the response — proportion of searches returning a non-empty, non-fallback result — or a business signal such as the conversion rate on the search page. This is the classic case where status codes alone are not a service level.

saying these in an interview costs you the question

  • Anything we already graph can serve as an SLI
  • More SLIs mean better coverage of the service
  • High CPU causes slowness, so CPU is a fine proxy
  • Host uptime is the same thing as user availability
  • One service-wide success ratio covers every journey

context

open as a page

A service has a 99.9% availability SLO measured over a 30-day rolling window. How much error budget does that give you, and what does it mean for a team to "spend" it?

level: juniorimportance: must knowfreq 78%

basics

~20 s

An error budget is one minus the SLO target over the window. A 99.9% target across 30 days allows about 43 minutes of unavailability, or 0.1% of requests, before the objective is missed. Spending it means shipping change against that allowance.

open as a page

In service reliability practice, define SLI, SLO and SLA, explain how the three relate to one another, and say which one a financial penalty attaches to.

level: juniorimportance: must knowfreq 88%

basics

~20 s

An SLI is a measured number, usually good events divided by valid events. An SLO is the internal target that number must meet over a defined window. An SLA is the customer-facing contract, set looser than the SLO, and it is the only one carrying penalties.

open as a page

A service has a 99.9% availability SLO measured over a rolling 30-day window. Explain what error-budget burn rate means, what a burn rate of exactly 1 signifies, and how you compute it from an observed error ratio.

level: middleimportance: must knowfreq 65%

basics

~20 s

Burn rate expresses error-budget consumption as a multiple of the sustainable pace: rate 1 spends the whole budget exactly at the window's end. Compute it as the observed bad-event ratio divided by (1 minus the SLO target).

open as a page

How would you define a latency SLI for an HTTP API, and why is "average response time under 300 ms" the wrong shape for one?

level: middleimportance: must knowfreq 66%

basics

~20 s

Define latency as the proportion of valid requests served faster than an explicit threshold — for example 99% of checkout requests under 400 ms. An average hides the slow tail entirely: a handful of thirty-second requests barely move it while those users are the ones leaving.

open as a page

Why do reliability teams deliberately set an availability SLO below 100%, and how would you choose between 99.9% and 99.99% for a specific service?

level: middleimportance: must knowfreq 70%

basics

~20 s

A 100% target is unachievable and forbids all change, since every deploy and dependency can fail. Choose between 99.9% and 99.99% by what users can actually perceive and what the extra nine costs: 99.9% allows about 43 minutes of downtime a month, 99.99% about 4.

open as a page

You are designing SLO-based paging for a service with a 99.9% availability SLO over 30 days, using multi-window multi-burn-rate alerts. Which burn-rate and window pairs would you choose, and what job does each of the two windows in a pair do?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Use two paging tiers: 14.4x averaged over 1 hour for fast burns and 6x over 6 hours for slower ones, each ANDed with a short window (5 minutes and 30 minutes) of the same threshold. The long window sets sensitivity; the short one proves the burn is still happening.

open as a page

Your service exhausts its error budget in week two of the quarter. What should a written error budget policy already commit the team to at that moment, and what stops that commitment from being quietly ignored?

level: seniorimportance: must knowfreq 66%

basics

~20 s

A budget policy, agreed in advance by engineering, SRE and product, states what happens at zero: feature releases stop while reliability and security fixes continue, reliability work is funded, and a single named owner may override — with every override logged and counted.

open as a page

Many teams page on a static rule such as 'error ratio above 2% for five minutes'. Against a 99.9% availability SLO over 30 days, explain the two opposite ways that static threshold fails, and what burn-rate alerting changes.

level: middleimportance: should knowfreq 45%

basics

~20 s

A static error-ratio threshold ignores both the SLO and how long the condition lasts. It pages too early on brief or low-volume blips that cost almost no budget, and stays silent through a sustained sub-threshold error rate that quietly drains the entire budget in days.

open as a page

When defining an availability SLI for an HTTP API, which responses should count as bad — and do 4xx errors, rate-limit 429s, health checks and bot traffic belong in the denominator?

level: middleimportance: should knowfreq 48%

basics

~20 s

Default to 5xx as bad and 2xx/3xx as good, then decide each grey case deliberately: health checks and scanner traffic come out of the denominator, 429s count as bad when you shed because you were overloaded, and 4xx stays out unless your own change caused it.

open as a page

Two teams both report a 99.9% availability SLO, but one tracks its error budget in bad minutes and the other in failed requests. How does that choice change what the budget shows during a partial outage, and how would you pick?

level: middleimportance: should knowfreq 45%

basics

~20 s

Time-based budgets count bad minutes and weight every minute equally; request-based budgets count failed requests and weight an incident by the traffic it hit. Partial outages and off-peak incidents cost very differently under the two.

open as a page

An SLO target must be evaluated over a window. Compare a rolling 28-day window with a calendar-quarter window for an availability SLO, and say what each choice changes for the team.

level: middleimportance: should knowfreq 45%

basics

~20 s

A rolling window recomputes continuously, so an outage ages out gradually and there is never a reset to wait for. A calendar window aligns with billing and planning cycles but resets abruptly, so an early outage poisons the whole period and a late one barely registers.

open as a page

An SLI is conventionally specified as the ratio of good events to valid events. Why is defining "valid" as consequential as defining "good", and what do teams commonly get wrong in that denominator?

level: middleimportance: should knowfreq 42%

basics

~20 s

The denominator decides what the service is accountable for. Health checks and bot traffic inflate it with easy successes that mask real user failures, and quietly excluding awkward traffic makes the indicator look better without the service improving.

open as a page

A page configured as 'average error ratio over the last hour exceeds 14.4 times the budget rate' fires during a ten-minute total outage of a 99.9% service. The outage is mitigated, but the page keeps firing for most of the following hour. Explain why, and how the standard fix works.

level: seniorimportance: should knowfreq 32%

basics

~20 s

The outage stays inside the trailing one-hour average until it rolls out of the window, so the condition remains true long after the burn stops. The fix is to AND the long window with a short one, typically a twelfth of its length, so the alert clears roughly five minutes after recovery.

open as a page

An availability SLI for a web API can be computed from load-balancer logs, from counters inside the application, from client-side telemetry, or from synthetic probes. What does each vantage point miss, and how do you choose between them?

level: seniorimportance: should knowfreq 54%

basics

~20 s

Each vantage point sees a different slice of the request path. Application counters miss everything that never reached the process, load-balancer logs miss DNS and network failures, client telemetry only reports from clients healthy enough to report, and probes measure a synthetic journey at low resolution.

open as a page

A nightly ETL job and a streaming ingest feed have no request-and-response pattern, so availability and latency SLIs do not fit. What SLIs would you define for them, and why is "the job exited successfully" not one?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Define indicators on the data, not the job: freshness (how recently the output was updated), coverage or completeness (share of input records processed), and correctness (share of outputs that validate). A zero-exit job can emit nothing at all, so exit status says nothing about what a consumer received.

open as a page

A service ends every quarter with more than 90% of its error budget unspent. What does that tell you, and what would you actually change?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A permanently unspent error budget is not a win. It usually means the target is looser than what users notice, or the team is buying reliability nobody asked for. Tighten the SLO, or deliberately spend the budget on velocity.

open as a page

Your company publishes a 99.9% availability SLA to customers. What availability SLO should the team run to internally, and why should the two numbers not be the same?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Run a tighter internal target than the contract — commonly 99.95% behind a 99.9% SLA. The gap is the team's reaction margin: internal alarms fire and remediation starts while the customer is still well inside the promised level and owed nothing.

open as a page

You own the alerting standard for a platform of several hundred services, some with 30-day SLO windows and some with 7-day windows. Would you mandate a single burn-rate alert configuration for all of them? Explain what must vary per service and what you would hold fixed.

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Hold the budget-spend policy fixed, not the multipliers. Burn rate already normalises across different SLO targets, but the multipliers are derived from the window length, so a 7-day SLO needs 3.36x where a 30-day SLO needs 14.4x, and low-traffic services need an extra event-count guard.

open as a page

You own reliability standards for an organisation where every team has invented its own SLIs and none of them are comparable. How would you standardise SLI definitions across teams, and where would you deliberately let them diverge?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Standardise the shape and the measurement point, not the targets: compute one default availability and latency indicator centrally from a shared ingress for every service, and let teams add journey-specific indicators — freshness, correctness — that only they can define.

open as a page

A cloud zone failure took your service down for an afternoon and burned 80% of its quarterly error budget, while your own code behaved perfectly. Should that spend be exempted from the budget, and what does your answer commit you to?

level: principalimportance: nice to knowfreq 33%

basics

~10 s

Count it. An error budget measures user experience, not team fault, and exempting dependency failures hides the case for removing the dependency. Adjust the policy response to the cause instead of adjusting the measurement.

open as a page

A large prospective customer asks for a 99.99% availability SLA with financial penalties. As the engineering owner in that negotiation, how do you decide what to commit to?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Start from the worst measured month, not the average, then compose the dependency chain: serial dependencies multiply, so three components at 99.99% cap you near 99.97%. Four nines allows 4.32 minutes a month, which is an automated-failover architecture, not a promise you can talk your way into.

open as a page