skip to content

A release moved p95 latency while the median stayed flat, so what changed and what do you do?

level: principalimportance: should knowfreq 28%

answer

  1. a pure shift moves every quantile equally
  2. so this is shape, not location
  3. a minority of traffic degraded
  4. check the tail move against its own noise
  5. distributions compared, not individuals tracked

basics

~20 s

The distribution changed shape, not location: a pure shift moves every quantile together, so a tail-only move means part of the traffic degraded while typical traffic did not. Check the move against tail uncertainty, then segment.

solid answer

~50 s

Read it as a quantile treatment effect: the effect at `tau = 0.5` is about zero while the effect at `tau = 0.95` is positive. A pure location shift would move every quantile by the same amount, so this is a shape change - some subpopulation or code path degraded while the typical request was untouched. Two disciplines follow. First, verify the move survives its own uncertainty: a p95 rests on far fewer effective observations than a median, so put an interval on it before calling it a regression. Second, resist the individual-level reading - the effect compares the 95th percentile of two distributions, and saying the user who was at p95 got slower would need an assumption that ranks are preserved. Then act: plot the effect across a range of quantiles, segment by shard, region or cache state to find the affected slice, and decide explicitly whether the guardrail is a central or a tail metric.

go deeper

for a junior

Know that the median and the p95 describe different parts of the same distribution, so one can move while the other does not, and that this points at a subset of slow requests rather than a uniform slowdown.

for a middle

Explain the contrast between a location shift, which moves all quantiles alike, and a shape change, which moves only some, and be able to say why a tail quantile is estimated less precisely than the median.

for a senior

Demonstrate the investigation: verify the move exceeds its uncertainty, plot the effect across several quantiles, then segment by shard, region or device until the affected population shows the degradation at its own median.

for a principal

Own the policy. Decide which quantile the organisation commits to as a guardrail, require pre-registration and intervals, set traffic floors for tail alerting, and make the tail-versus-median tradeoff an explicit, owned business decision.

## Naming the pattern Compare the outcome distribution before and after the release. The quantile treatment effect at level `tau` is the difference between the treated and control quantiles at the same `tau`: `QTE(tau) = Q_tau(treated) - Q_tau(control)` A pure location shift - everything slower by a constant - has `QTE(tau)` equal to the same constant at every `tau`. What you observed is `QTE(0.5)` near zero and `QTE(0.95)` clearly positive. That is a shape change: the body of the distribution is intact and the upper tail has stretched. This is not an exotic pattern. It is the signature of a change that only bites on a subset of traffic - a slow path taken by a minority of requests, one shard or region, a cache-miss route, an older device class, a retry that only fires under contention. Because that subset is a minority, it lives in the tail and never touches the median. ## Step one: is the move real? Before anything else, put uncertainty on the tail number. A sample p95 rests on far fewer informative observations than the median, and its precision is governed by the density of the distribution where the quantile falls, which is low in the thin part of a right-skewed latency curve. A 20ms move in p95 with an interval spanning 60ms is not a finding. Teams that skip this step spend weeks chasing regressions that were resampling noise, and - worse - they learn to distrust tail metrics generally. The corollary is a reporting rule you can own as a lead: no tail metric is quoted in a decision review without an interval beside it. ## Step two: read the effect at the right altitude One quantile is a single slice. Compute the effect across a grid of quantiles - 0.5, 0.75, 0.9, 0.95, 0.99 - and look at the curve. Where does it lift off zero? A curve that is flat until 0.9 and then climbs steeply says a small, identifiable slice of traffic is affected. A curve that rises gently from 0.6 says something broader and milder. The shape of the effect curve is a much stronger lead for the investigation than any single number, and it costs nothing extra to produce. ## Step three: the interpretation trap worth knowing A quantile treatment effect is a comparison of two distributions at the same rank. It is *not* the effect on the individual who occupied that rank. Saying "the users at p95 each got 40ms slower" requires assuming that the change preserves everyone's rank in the distribution - rank invariance - which nothing in the data establishes. The same observed tail shift is consistent with a small group getting dramatically worse while others improved slightly. This distinction matters when the conclusion drives compensation, SLO negotiation, or a rollback decision, and stating it correctly is one of the cleanest senior-versus-principal separators on this topic. ## Step four: find the slice Segment. Split by the dimensions that plausibly generate a minority slow path and recompute the quantile effect within each segment. The signature you want is a segment where the effect appears at the median - that is, the affected population is degraded typically, not just in its tail. When you find that, you have converted a tail mystery into an ordinary regression with a cause. If no segmentation localises it, the change is diffuse and the story is different: a shared resource under contention, for example. ## Step five: the organisational call This is the part a lead owns rather than delegates. **Which guardrail does the team commit to?** A tail metric matches user-perceived worst-case experience but is noisy at low traffic volumes; a central metric is stable but blind to exactly this failure mode. Choosing one is a policy decision with a real cost either way, and it should be made once, in advance, not argued after a result appears. **Pre-register it.** If p95 is the guardrail, it is the guardrail before the release, not after the median came back clean. Selecting the quantile that shows the effect you want is a discipline failure that invalidates the reading. **Set an alerting volume floor.** A p99 target on a low-traffic service manufactures pages from noise. Decide the traffic level at which each quantile is meaningful and do not alert below it. **Decide the acceptable tradeoff.** Sometimes a tail regression on 2% of traffic in exchange for a large median improvement is the right business call. Making that explicit - and naming who bears the cost - is the judgment the question is really probing. ## The one-line version A flat median with a moved p95 means the distribution changed shape rather than position; confirm the move against tail uncertainty, read the whole quantile-effect curve, segment to find the affected minority, and be careful not to turn a statement about distributions into a claim about individuals.

  • What would a pure location shift have looked like in the same data?
    Every quantile would have moved by roughly the same amount - median, p75, p95 and p99 all up by a similar figure - so the quantile-effect curve would be flat and positive across the range. Seeing the effect concentrated in the upper quantiles is what rules out a uniform slowdown and points at a subpopulation.
  • Why is it wrong to say the users at p95 each got 40 milliseconds slower?
    Because a quantile treatment effect compares the 95th percentile of two distributions, not the same individuals before and after. Concluding anything about a specific user requires assuming ranks are preserved between the two conditions, which the data cannot establish. The same tail shift is consistent with a small group degrading far more than 40 milliseconds.
  • How do you stop teams from picking whichever quantile tells the story they want?
    Pre-register the guardrail quantile with the experiment or release plan, publish the whole quantile-effect curve rather than one number, and require an interval on any tail metric quoted in a decision. Post-hoc quantile selection is the tail-metric equivalent of choosing a metric after seeing results.
  • When is a tail regression an acceptable price for a median improvement?
    When the affected population is identified, small, and not systematically the people you most need to serve - and when someone has explicitly accepted the tradeoff. The failure mode is an unexamined tail regression that concentrates on one region, device class or customer tier, so segment before deciding, then record the decision and who owns it.

The average commute is unchanged but the worst days got worse - the timetable did not shift, one unreliable connection started failing.

saying these in an interview costs you the question

  • Declares a regression without an interval on the tail metric
  • Assumes a tail move means every user got slower
  • Reads a quantile treatment effect as an effect on named individuals
  • Picks the quantile that shows the desired result after the fact
  • Averages the problem away by reporting only the mean
  • Alerts on p99 for a service with almost no traffic

context