How would you define a latency SLI for an HTTP API, and why is "average response time under 300 ms" the wrong shape for one?
answer
- a ratio of counts, not a statistic
- the mean lives in an empty valley
- percentiles do not average across servers
- the knee of the latency curve
- a fast 500 is not fast service
basics
~20 sDefine latency as the proportion of valid requests served faster than an explicit threshold — for example 99% of checkout requests under 400 ms. An average hides the slow tail entirely: a handful of thirty-second requests barely move it while those users are the ones leaving.
solid answer
~50 sI would express it as a ratio, not as a statistic: the proportion of valid requests completed faster than a stated threshold, over a stated window. Averages are the wrong shape because latency distributions are long-tailed and bimodal — cache hits at 5 ms and cache misses at 2 s average out to a number nobody ever experienced, and one request in a thousand taking thirty seconds moves the mean by 30 ms while that user gives up. The threshold should come from what the calling experience can absorb, often the knee of the latency curve where abandonment starts climbing, not from a round number. The percentile follows the journey: a tight target at the 99th protects the worst-off users, who are frequently your largest accounts, while a looser 50th describes the typical experience. And I would exclude or separately account for errors, or a service that fails fast will report excellent latency.
go deeper
Know that latency is reported as a percentile or as a proportion under a threshold, never as an average, and be able to say why: a few very slow requests barely move a mean but ruin real sessions.
Explain the ratio form — proportion of valid requests faster than an explicit threshold — and be ready to state where the threshold comes from and why percentiles cannot be averaged across instances or time windows.
Show the operational traps: fast failures inflating a latency indicator, in-handler timers hiding queueing during overload, and tail percentiles becoming statistical noise on a low-traffic service. Say how you would pick the threshold from observed user behaviour.
Own the end-to-end budget: latency thresholds for individual services should be apportioned from what the whole user-facing chain can absorb. Be prepared to argue what tightening a tail target actually costs in architecture and spend, and when the honest answer is to leave it loose.
## State it as a ratio, not as a statistic The first correction is structural. A latency SLI is written as **the proportion of valid requests that completed faster than a threshold**, over a window — for example, *99% of checkout requests completed under 400 ms over the trailing 28 days*. It is not written as "the 99th percentile is under 400 ms", even though the two sound equivalent, and it is certainly not written as an average. Stating it as a ratio matters for two practical reasons. First, **percentiles do not aggregate.** You cannot average the p99 of ten servers to get the fleet p99, and you cannot average this hour's p99 with last hour's to get the two-hour p99. A percentile is a quantile of a distribution; combining them requires the underlying distribution, normally the histogram buckets. A ratio of counts, by contrast, adds up perfectly across servers, shards, regions and time windows — you sum the fast requests and sum the total. Any team that reports "average p99 across our instances" has produced a number with no defined meaning. Second, the ratio form composes with everything above it: the same shape of good-events-over-valid-events used for availability, so one machinery serves both. ## Why the average is the wrong statistic Latency distributions in real systems are long-tailed and usually multi-modal. A read path with a cache produces a mode at a few milliseconds (hits) and another at hundreds of milliseconds or seconds (misses, cold shards, retries, GC pauses, connection re-establishment). The mean falls in the empty valley between them — a value that describes no actual request. Worse, the mean is insensitive exactly where you need sensitivity. If one request in a thousand takes 30 seconds and the rest take 20 ms, the mean is about 50 ms and looks fine, while 0.1% of your users are staring at a spinner. Scaled to a service handling ten million requests a day, that is ten thousand abandoned sessions a day sitting invisibly inside a healthy-looking average. The tail is not noise; it is a population of users. ## Choosing the threshold The threshold is the substantive decision, and "300 ms because it sounds good" is the weak answer. Better sources, roughly in order of strength: - **What the caller can absorb.** If a page cannot render until this call returns and the product's budget for that page is one second, and two other calls sit in the same chain, your share is a few hundred milliseconds. Latency thresholds should be *derived from a budget across the chain*, not chosen per service in isolation. - **The knee of the observed curve.** Plot latency against a behavioural outcome — abandonment, retry rate, conversion. There is usually a region where the outcome is flat and then bends sharply. Setting the threshold at that knee makes the indicator track something real. - **Current achievable performance.** If the service currently serves 99.5% of requests under 250 ms, a 400 ms threshold is honest and leaves room; a 100 ms threshold means you have committed to a rewrite. A threshold that the service already meets 99.999% of the time is a vanity indicator — it will never move, and you will learn nothing from it. ## Choosing the percentile The percentile is a statement about *whose* experience you are protecting. - The **50th** describes the typical user and is nearly useless as a guarantee — half your traffic is allowed to be arbitrarily slow. - The **95th/99th** protect the tail. This is where the users on cold caches, large accounts, distant regions and unlucky retries live. In many businesses the heaviest users have the largest payloads and therefore live permanently in your tail — the 99th percentile is disproportionately your biggest customers. - The **99.9th and beyond** get expensive fast and increasingly measure your own infrastructure's noise floor rather than your code. At low traffic they are also statistically meaningless: with 1,000 requests an hour, the 99.9th percentile is a single request. A common, defensible pattern is a **two-threshold** indicator: a high proportion under a comfortable bound and a smaller proportion under a tight one — for example 99% under 400 ms *and* 90% under 100 ms. That captures both "nobody waits absurdly long" and "the normal case feels instant" without pretending one number does both. ## The fast-failure trap A latency indicator computed over *all* responses will report beautifully during a total outage, because a service returning HTTP 500 from its front door does so in single-digit milliseconds. The fix is to define the valid-event set explicitly: count only responses that were also successful, or count errors as failures of the latency indicator too. Whichever you choose, choose deliberately, because the default of "all responses, any status" is actively misleading in exactly the moments you care about. ``` good = requests where status is successful AND duration <= threshold valid = requests where status is successful (or all real user requests) SLI = good / valid ``` ## Where the number comes from Finally, latency has to be measured at the same vantage point as your other indicators, and it must include the time the request spent queued before your code ran. An in-process timer that starts when a handler is entered misses the time the request spent waiting in the accept queue during overload — which is precisely the time the user experienced. Timing at the ingress or from the client is what makes queueing visible.
- Why can't you compute a fleet-wide p99 by averaging the p99 of each instance?Because a percentile is a property of a distribution, not an additive quantity. Averaging quantiles of different distributions with different request counts yields a number that corresponds to no percentile of the combined data — it can sit well below the true fleet p99 when one instance carries most of the slow traffic. Aggregate the underlying histogram buckets, or express the indicator as a ratio of counts, which does add.
- Would you ever set a latency threshold above what the service currently achieves?Only deliberately, as a stated intent to invest. A threshold the service cannot meet means the indicator reads unhealthy from day one and the team learns to ignore it. The usual sequence is to set it near current achievable performance, confirm it correlates with real user behaviour, then tighten it as a funded piece of work rather than as an aspiration.
- How does the choice of percentile change for a low-traffic internal service?Tail percentiles stop being meaningful. At a few hundred requests an hour, the 99.9th percentile is one request, so it swings wildly on noise and cannot support a tight target. Either widen the measurement window so enough events accumulate, or set the indicator at a lower percentile with an honest threshold and accept coarser resolution.
saying these in an interview costs you the question
- Average response time is fine if the traffic is uniform
- Average the per-instance p99 to get the service p99
- Pick p99 under 100 ms because it sounds ambitious
- Timing inside the handler captures what the user waited
- Latency and errors are separate, so ignore status codes