After a release, your real-user monitoring shows p75 LCP improved from 3.1s to 2.6s while p95 LCP got worse, from 6.2s to 8.0s, and support tickets about slowness went up. How do you interpret that, and how would you decide whether the release was a win?
answer
- two ranks on one shape
- the left came in, the right went out
- tickets come from the tail
- cost proportional to device weakness
- segment before you judge
basics
~20 sThe release helped typical loads and hurt the slowest ones — a classic tradeoff that shifts cost onto weak devices and networks. Decide by segmenting the distribution to find which population regressed and how many visits it covers, not by comparing two headline numbers.
solid answer
~50 sPercentiles move independently, so a distribution can compress on the left and stretch on the right at the same time. The usual cause is a change that trades work for latency in a way only capable devices can absorb — shipping more JavaScript to render earlier, inlining critical CSS at the cost of a bigger uncacheable document, preloading aggressively and starving other requests on a narrow connection. The tickets are the tell: complaints come from the tail, and the tail got 30% worse. To decide, I'd stop comparing two numbers and look at the distribution: the good/needs-improvement/poor bucket split before and after, then the same comparison per device class, connection type, country and page template, with visit counts attached. If a definable segment regressed, that is a real regression to fix or roll back, even though the number the team reports upward improved.
go deeper
Know that percentiles are separate readings of a distribution and can move in opposite directions, and that a better p75 with a worse p95 means the slowest loads got slower rather than the data being wrong.
Explain the mechanism: a change whose cost scales with how weak the client is — extra script, a bigger uncacheable document, aggressive preloading — improves capable devices and punishes the tail. Be able to name what to segment by.
Show the full diagnosis path: distribution and bucket counts before and after, segmentation by device, connection, region and template, then the LCP sub-parts to localise the cause — and a decision framed as how many real visits got worse.
Own the incentive problem. A team scored on one percentile will ship changes that move it at the tail's expense; decide what tail signal sits next to the target, what triggers a rollback, and how tail work gets funded when it never improves the reported number.
## Percentiles are independent readings of one shape It is tempting to treat p75 and p95 as two views of the same dial, so that improving one should improve the other. They are not. They are two rank positions on a distribution whose shape can change arbitrarily. A release can pull the left side in — more loads finishing quickly — while pushing the right side out, and the two summary numbers will move in opposite directions with no contradiction. If your first instinct on seeing this is "one of these numbers must be wrong", you will waste a day auditing the pipeline instead of investigating a real regression. ## What actually causes this pattern The recurring theme is a change whose benefit is unconditional but whose cost is *proportional to how weak the client is*. - **More client work for an earlier paint.** Shipping additional JavaScript that renders content sooner is nearly free on a fast CPU and brutal on a low-end phone, where parse, compile and execute dominate. Fast devices move left; slow devices move right. - **Inlining at the expense of cacheability.** Inlining critical CSS or above-the-fold data into the HTML removes a round trip, which is a clear win on a fresh visit. But the document grows and stops being cacheable, so repeat visits and slow connections pay for bytes they used to skip. - **Aggressive preloading and prioritisation.** Preloading several resources at high priority helps when there is bandwidth to spare and actively harms when there is not, because the hero image now queues behind the things you promoted. - **A new dependency on a slow third party.** If the render path now waits on an external script, the visits where that request is slow or times out land squarely in the tail. - **A new bimodal path.** A cache, a feature flag or an experiment can create a second, slow population that did not exist before — the tail gets deeper without any single request getting slower. ## Diagnosis: stop looking at two numbers The first move is to replace both scalars with the shape. Compare the before/after histogram, or at minimum the three-bucket split of good, needs-improvement and poor, **with counts**. That immediately distinguishes two very different worlds: the poor bucket holding the same share of traffic but getting deeper (existing slow visits got slower), versus the poor bucket growing (visits moved into it). The second is a much more serious regression. The second move is to segment, and to segment along the dimensions where the tail actually lives: - **Device class** — low-end Android versus flagship versus desktop. This is usually where the split falls, because CPU cost is the most common thing a release adds. - **Connection type** — this separates a bandwidth story from a CPU story. - **Country or region** — proxies for network quality and distance to your edge. - **Page template** — a regression on one heavy template can move the site-wide tail on its own. - **First versus repeat visit** — this is the one that catches the cacheability tradeoff. Compute the same percentiles per segment, with visit counts next to them. Either the regression is concentrated — one segment moved a lot — which names the cause, or it is diffuse across every segment, which points at something on the common path such as a new blocking request. The third move is to check the sub-parts. Split LCP into time to first byte, resource load delay, resource load duration and render delay. Which part grew in the regressed segment tells you whether you added server time, delayed discovery of the element, added bytes, or added main-thread work before the paint could happen. ## Making the decision Frame it as people, not percentiles. If mobile-on-slow-network is 18% of your traffic and its p75 went from 4.5s to 6s, that is hundreds of thousands of visits made materially worse — and it is the population that generates the support tickets you are already seeing. Weigh that against the size and value of the improvement on the fast side. Three honest outcomes: 1. **Roll back**, if the regression is large, affects a meaningful segment, and the gain is marginal. 2. **Make the change conditional**, which is usually the right answer — ship the heavier path only where it pays, gated on a signal such as device memory or effective connection type, so weak clients keep the old path. 3. **Accept and fix forward**, if the regression is small, well understood, and the fix is already in flight — but say so explicitly and set a date, rather than letting the improved headline close the discussion. What you should not do is declare victory because the number you are scored on improved. A metric that only ever moves in the reported direction, while tickets rise, is a metric that has stopped measuring anything. Reporting the tail alongside the target percentile is what keeps this failure mode visible; the tail's job is not to be a target, it is to catch exactly this.
- How would you tell whether the tail got deeper or simply got more crowded?Compare the bucket counts, not just the percentile values. If the share of visits in the poor bucket is unchanged but p95 and p99 moved further out, existing slow visits got slower. If the poor bucket's share grew, visits migrated into it — that means a population crossed from acceptable to bad, which is usually the more damaging outcome and easier to trace to a specific segment.
- Which single segmentation would you try first here, and why?Device class. The most common way to improve the median while wrecking the tail is to add client-side work, and CPU cost scales viciously on low-end hardware while barely registering on a flagship. If low-end Android shows the whole regression and desktop shows none, you have the shape of the answer before you have opened a single trace.
- The tail regression is real but only affects 3% of visits. Do you still act?It depends on who those visits are and what they were doing. Three percent of a large site is a lot of people, and if they cluster in a market you are investing in, or on a checkout template, the revenue impact is out of proportion to the share. If they are genuinely marginal — ancient devices on a page nobody converts from — documenting the tradeoff and moving on is a defensible call, provided you actually documented it.
- How do you keep this from being invisible next time?Report a tail percentile next to the target one on the same dashboard, and alert on segment-level regressions rather than only on the site-wide figure. The headline number is the bar you are held to; without a tail number beside it, any change that trades the slow quarter for the fast three quarters looks like an unambiguous win, and nobody sees the tradeoff until support does.
saying these in an interview costs you the question
- Claims p95 must follow p75, so one number is a bug
- Declares the release a win because the reported metric improved
- Blames tail movement on bots or noise without checking
- Compares aggregate numbers without segmenting or counting visits
- Treats the slow quarter as users not worth optimising for