skip to content

When defining an availability SLI for an HTTP API, which responses should count as bad — and do 4xx errors, rate-limit 429s, health checks and bot traffic belong in the denominator?

level: middleimportance: should knowfreq 48%

answer

  1. classification is the definition
  2. health checks dilute the denominator
  3. who caused the 4xx?
  4. overload shedding versus documented quota
  5. graph what you exclude

basics

~20 s

Default to 5xx as bad and 2xx/3xx as good, then decide each grey case deliberately: health checks and scanner traffic come out of the denominator, 429s count as bad when you shed because you were overloaded, and 4xx stays out unless your own change caused it.

solid answer

~50 s

The classification is the design work, not a detail. Server errors are unambiguously bad and successes are good; the arguments are all at the edges. Health checks, uptime probes and internal warmers must leave the denominator entirely, because on a low-traffic service they can outnumber real requests and hold the ratio high while every real user is failing. Bot and scanner traffic goes the same way, since you should not spend reliability on scrapers. A 4xx is the client's fault by definition, so it stays out — except when a deploy of yours starts producing 400s or a broken token service produces an auth storm, which is real user pain wearing a client-error status code. A 429 is the one I argue about: if you shed because you were overloaded, that user was failed by you and it should count; if it is a documented quota an abusive client blew through, exclude it. Whatever you exclude, write it down, because a definition with enough exclusions reports 100% forever.

go deeper

for a junior

Know that an availability indicator needs an explicit rule for which responses are good and which are bad, and that 5xx is the usual bad case while 2xx and 3xx are good. Say that health-check traffic should not be counted.

for a middle

Work through the grey cases and justify each: why probes and bots leave the denominator, why 4xx normally stays out, and what makes a 429 arguable. Give the health-check dilution effect concretely on a low-traffic service.

for a senior

Show that you would classify a 429 by the reason for shedding rather than the status code, catch the self-inflicted 4xx that hides a broken deploy, and handle responses with no status code at all. Name the discipline of graphing excluded volume.

for a principal

Own the governance angle: definitions drift, exclusions accumulate, and a number that can only read 100% is worse than no number. Be able to say who reviews classification changes and what evidence forces a redefinition.

## Why classification is the real work An availability SLI is a ratio, and the interesting decisions are all about **which events go in the denominator and which of those count as bad**. Two teams can compute "availability" for the same service and differ by a full percentage point purely through classification. Because a target is only meaningful against a definition, the classification *is* the service level. Start from the anchor points nobody argues about: a 5xx response is a bad event; a 2xx or 3xx is a good one. Everything else needs a decision with a stated reason. ## Health checks, probes and internal traffic These must come out of the denominator, and the reason is quantitative. A service receiving 20 real requests a minute may receive 200 liveness and readiness checks a minute from its scheduler, plus uptime probes and mesh health checks. If those are in the denominator and the shallow health endpoint keeps returning 200 while the real dependency is dead, the ratio during a **total** functional outage sits above 90%. The indicator has been diluted by traffic that has no user behind it. The same argument applies to internal cache warmers, load-test traffic against production, and your own synthetic probes if they share the ingress path. Filter them by path, by user agent, or better, by an explicit header or client identity so the exclusion cannot be spoofed by accident. ## Bot, scanner and abusive traffic Security scanners generate large volumes of requests to paths that do not exist, producing 404s, and occasionally trip real 5xx paths in code nobody meant to expose. Crawlers hit expensive endpoints at odd hours. Including them means your reliability number partly describes your experience of being scanned. Exclude them where you can identify them, and accept that identification is imperfect. ## The 4xx family The default is that a 4xx is not a service failure — the client sent something invalid, and a service correctly rejecting it is behaving well. Two exceptions matter in practice: - **You caused it.** A schema change, a stricter validator, or a new required header shipped in your deploy turns previously valid requests into 400s. Users experience a broken product; your availability indicator reports perfection. This is the most common way an outage hides from an SLI. - **Authentication and authorisation.** A 401 or 403 storm caused by an expired signing key, a clock skew, or a broken token service is your outage, delivered as a client-error status code. Two workable responses: either count a specific subset (401/403, and 400 above a baseline) as bad, or keep 4xx out of the availability indicator and cover it with a separate quality indicator that watches the 4xx *rate* against its own baseline. The second is cleaner, because a raw 400 count is dominated by genuinely malformed client traffic. ## The 429 argument Rate-limit responses are the sharpest case because the answer depends on **why** you shed the request. - Shed because the service was **overloaded** and protecting itself: that user asked for something legitimate and did not get it. It is a bad event. Excluding it means a service can survive any load spike with a perfect availability number simply by rejecting more traffic — an indicator that rewards the failure it is supposed to detect. - Shed because the client **exceeded a documented quota** it agreed to: that is the contract working. Exclude it, or the SLI measures your customers' behaviour instead of your reliability. The implementable version is to distinguish the two at the source: emit different reasons for quota-based and pressure-based shedding, and classify on the reason rather than on the status code. ## Requests with no status code at all Connection resets, timeouts and TLS handshake failures produce no HTTP status. They are real failures and they are invisible to anything counting status codes in the application. This is one more reason the count usually comes from the ingress, which records the aborted connection. ## Multi-status and partial responses An API that returns 200 with a body containing per-item errors, or a GraphQL endpoint that returns 200 with an `errors` array, will report perfect availability while failing. If your API has this shape, the classification must read the body or the service must emit an explicit outcome label alongside the response, otherwise the status code is measuring nothing. ## Write it down, and watch the exclusions The discipline that keeps this honest: **every exclusion is documented, and the excluded volume is graphed.** Two things then become visible — an exclusion that starts growing (a client suddenly producing 30% of traffic as 429s), and the slow accumulation of exemptions that turns an indicator into a number that can only ever read 100%. When the SLI is healthy and the support queue is not, the classification is the first place to look.

  • A deploy makes the API reject previously valid requests with 400, and availability stays at 100%. How do you catch that class of failure?
    Either classify a subset of 4xx as bad — typically 401/403 plus 400s above their normal baseline — or run a separate quality indicator that watches the client-error rate against its own baseline and treats a step change as a failure. The underlying point is that a status code records who the protocol blames, not who actually broke the product.
  • How would you implement the distinction between quota-based and overload-based 429s?
    Do not classify on the status code. Have the shedding layer emit the reason as a label or log field — quota exceeded versus pressure shed — and build the indicator on that field. It costs one attribute at the point of rejection and removes an argument that otherwise recurs in every review, since only the service itself knows why it refused.
  • What is the risk of an SLI definition with a long list of exclusions?
    It converges on always reading 100%. Each exclusion is individually defensible and collectively they remove every event that could ever be bad. The countermeasure is to graph excluded volume next to the indicator and to require that any new exclusion names the failure it removes; if the number is healthy while users complain, the exclusion list is the first suspect.

saying these in an interview costs you the question

  • Anything that is not a 5xx counts as success
  • Leave health checks in, they are requests too
  • 429 always counts as the client's fault
  • A 200 response means the request succeeded
  • Exclusions do not need to be written down anywhere

context