An SLI is conventionally specified as the ratio of good events to valid events. Why is defining "valid" as consequential as defining "good", and what do teams commonly get wrong in that denominator?
answer
- the denominator picks who you answer to
- probes succeed even when users cannot
- every exclusion improves the number
- a 400 storm can hide inside "client error"
- one bad minute weighs the same as another
basics
~20 sThe denominator decides what the service is accountable for. Health checks and bot traffic inflate it with easy successes that mask real user failures, and quietly excluding awkward traffic makes the indicator look better without the service improving.
solid answer
~50 sThe numerator answers "what counts as success", but the denominator answers the harder question: which events the service is answerable for at all. Get it wrong and the ratio flatters you. The classic case is health-check and synthetic probe traffic sharing the counter with real requests — if probes outnumber users ten to one, a complete outage of user traffic barely moves the indicator, because the probes keep succeeding. The same applies to crawler traffic and to requests the service never had a chance to serve. The other direction is equally dangerous: excluding 4xx responses as "the client's fault" is defensible right up to the release that turns valid requests into 400s, which then becomes invisible. And because every exclusion improves the number, the definition of "valid" is a lever someone can pull under pressure. Treat changes to it as contract changes: written down, reviewed, and dated.
code
python · 9 linesdef sli(events):
valid = [e for e in events
if not e["is_probe"] # health checks are not user traffic
and not e["is_synthetic"] # neither are uptime checkers
and e["status"] not in (404,)] # absent resource: not a service failure
if not valid:
return None # no traffic: no measurement, not 100%
good = [e for e in valid if e["status"] < 500 and e["latency_ms"] <= 300]
return len(good) / len(valid)go deeper
Know the shape of the ratio — good events over valid events — and be able to say that health-check traffic should not be counted alongside real user requests. Recognising that the denominator is a choice at all is the main step here.
Work through the dilution arithmetic out loud: with probes at ten times user volume, a total user-facing outage still reads as roughly 91% success. Be ready to argue both sides of excluding 4xx responses.
Show the governance instinct: exclusions only ever improve the number, so the definition of valid needs the same review as the target and must change going forward, never retroactively. Be able to pick between a request-based and a time-window form and justify it.
Own the comparability problem across an organisation — if every team defines valid differently, a fleet-wide reliability number means nothing. Be ready to describe the minimum shared definition you would mandate and where you would let teams diverge.
## The ratio and its two halves The standard shape of an SLI is ``` SLI = good events / valid events ``` Most discussion goes into "good" — which status codes count, which latency threshold. The denominator gets less attention and decides more, because it defines the population the service is being judged on. Two services with identical behaviour and different denominators report different reliability. ## The dilution trap The most common defect is a denominator that includes traffic which is trivially successful. Health checks fire every few seconds from every load balancer and every orchestrator, synthetic probes run on a schedule, and uptime checkers hit a cheap endpoint that touches nothing. All of them succeed almost always, and all of them are cheap enough to run at a rate far above real user traffic. Suppose probes generate ten times the request volume of users. If every real user request fails for an hour, the measured success rate over that hour is about 91% rather than 0% — the service looks degraded rather than down. Over a 28-day window that outage may not even breach the target. The fix is to filter the denominator to requests that represent real usage, and to keep probe results as their own separate signal rather than mixed into the user-facing indicator. Crawler and scanner traffic causes a subtler version of the same problem, and can distort in either direction: bots that hammer a cached endpoint dilute the denominator with easy wins, while a scanner probing nonexistent paths can generate a wall of errors that has nothing to do with user experience. ## The exclusion trap The opposite mistake is excluding too much. Client errors are the usual argument: a 400 or a 404 is the caller's fault, so it should not count against the service. Often true — and it creates a blind spot the moment a deployment starts rejecting valid requests as malformed. A change that turns 2% of real traffic into 400s is a serious regression that an indicator excluding all 4xx will never see. Workable positions on this exist and should be chosen deliberately rather than by default: - Exclude 4xx from the numerator's fault but keep a separate indicator or alert on client-error *rate*, so a sudden rise is still visible. - Treat specific codes differently — 429 (rate-limited by you) and 401/403 arguably reflect your behaviour, while 404 for a genuinely absent resource does not. The same reasoning applies to declared maintenance windows, requests from internal test accounts, and traffic from a customer running a load test against production. Each exclusion may be individually reasonable; the pattern to watch is that every one of them makes the number better. ## Governance: the denominator is a lever Because exclusions only ever improve the ratio, the definition of "valid" is where pressure lands when a target is being missed. "Should we really be counting those requests?" is a question that always has a plausible answer. The defence is procedural rather than technical: write the definition down alongside the target, require the same review to change it as to change the threshold, apply changes going forward rather than retroactively, and record when each change took effect so historical comparisons stay honest. ## Request-based versus time-window indicators Not everything can be expressed as a count of requests. A second common form is the time-window ratio: ``` SLI = good time slices / valid time slices ``` A minute is "good" if some condition held throughout it — success rate above a bar, a pipeline's data fresher than a threshold, a queue below a depth. This suits things with no natural per-request event: batch pipelines, data freshness, background workers. Its trade-off is important and often missed. A time-window indicator weights every bad minute equally: a minute in which one request out of a million failed and a minute in which everything failed both count as one bad minute. That makes low-traffic incidents look as severe as total outages and vice versa. Request-based indicators weight by volume, which usually tracks user impact better — but they need enough events to be meaningful, which is exactly what low-traffic services lack. Choose the form to match what the service actually produces, and say which one you chose when quoting a number, because the two are not comparable.
- During a one-hour window the service received no valid requests at all. What does the indicator report?Nothing — the correct answer is no measurement, not 100%. A zero denominator means the window carries no evidence either way, and reporting perfection for it lets quiet periods paper over the noisy ones. Most implementations either skip empty windows or carry the last known state forward explicitly; either way it should be a stated decision, not an accident of the division.
- A customer runs a load test against production and generates a wall of errors. Should those count?It depends on whether the service was supposed to survive it. If they exceeded documented limits and were correctly rate-limited, the 429s are the service working as designed and can be excluded with the exclusion recorded. If the service collapsed under a load it was meant to handle, that is a genuine failure and excluding it hides a real capacity problem — the traffic being deliberate does not make the outage imaginary.
- How would you keep the definition of "valid" from being quietly loosened over time?Version it with the SLO itself, in the same reviewed document, and make every change dated and forward-looking rather than retroactive. Keeping the raw unfiltered counts alongside the filtered ones makes drift visible, since anyone can compare the two series. The signal to watch for is a run of individually reasonable exclusions that all move the number the same way.
saying these in an interview costs you the question
- Counting health-check and probe traffic in the user-facing indicator
- Excluding all 4xx without any separate watch on client-error rate
- Changing what counts as valid while a target is being missed
- Treating a time-window ratio and a request ratio as interchangeable
- Reporting 100% for a window in which the service received no traffic