Your site-wide p75 LCP improved over the last quarter, but when you split the same data by device class, p75 LCP got worse for mobile and worse for desktop. As the person who owns performance reporting for the organisation, how do you explain that, and how would you set reporting up so leadership is not misled by it again?
answer
- the blend depends on the proportions
- who was measured, not what they felt
- every subgroup down, the total up
- you cannot average two p75s
- publish the traffic share beside it
basics
~20 sThe traffic mix shifted toward the faster segment, so the blended number improved while both segments regressed — a mix-shift effect. Fix the reporting, not the arithmetic: publish segmented percentiles with their traffic shares beside the headline, and never derive an aggregate by combining per-segment percentiles.
solid answer
~50 sA site-wide percentile is a blend of populations, and its value depends on the proportions as much as on the latencies. If desktop is faster than mobile and desktop's share of traffic grew — a campaign, a seasonal shift, a market launch, an app that pulled mobile users away — the blended p75 improves even though every segment got worse. Nothing is broken; the aggregate is answering a different question from the one people think it answers. The reporting fix has three parts. Show device-class segments and their traffic shares next to the headline number, so a mix change is visible as a mix change. Never compute an overall percentile by averaging or weighting per-segment percentiles — percentiles are not additive; recompute from the underlying samples. And judge progress on segments, keeping the blended figure only for the external bar you are actually assessed against.
go deeper
Know that a site-wide percentile mixes different groups of users together, so it can move because the mix of who visited changed rather than because the site got faster or slower.
Explain the mechanism concretely — a larger share of samples from the faster segment pulls the pooled rank down — and know that you cannot rebuild an aggregate percentile by averaging or weighting per-segment percentiles.
Show how you would prove it from the data: segment percentiles with traffic shares over several periods, a mix-normalised comparison against fixed weights, and a pipeline that keeps distributions so the roll-up can be recomputed.
Own the reporting contract: which segments exist and who owns each, what the blended figure is allowed to be used for, how period-over-period claims are normalised, and how you stop a headline metric from being satisfied by traffic drift instead of engineering.
## What is happening This is mix shift — the performance version of Simpson's paradox. An aggregate percentile is computed over one pooled set of samples drawn from several populations with different distributions. Its value depends on two independent things: how fast each population is, and what share of the samples each contributes. Change the shares and the aggregate moves, even with every population's own distribution held fixed. Make it concrete. Suppose mobile visits are slow and desktop visits are fast, and last quarter the split was 70/30 mobile-heavy. This quarter a desktop-oriented campaign, a seasonal swing, or a native app absorbing your most engaged mobile users moves the split to 50/50. Mobile got slower. Desktop got slower. But a much larger share of the pooled samples now comes from the fast population, and the pooled p75 — which just reads off a rank in the merged sorted list — lands at a better value. The arithmetic is correct and the conclusion "the site got faster" is false. Other mix shifts that produce the same effect: a slow country's traffic drops after you stop advertising there; a heavy page template loses traffic to a light one after a navigation change; bot or synthetic traffic is filtered out and it happened to sit in one tail; you launch in a market with excellent connectivity. The common structure is always a change in *who is measured*, not in *what they experienced*. ## Why you cannot fix it with arithmetic The first thing people reach for is recombining the segment numbers — average the mobile and desktop p75s, or take a traffic-weighted average. Both are wrong, because **percentiles are not additive**. There is no operation on two percentile values that yields the percentile of the combined population; the merged rank depends on the full shape of both distributions, not on one point from each. A traffic-weighted average of percentiles is especially seductive because it looks statistically responsible, and it can be off by a lot when the two distributions differ in shape. The only correct way to get an aggregate percentile is to compute it from the underlying observations — the raw samples, or a distribution representation such as a histogram that can be merged before the percentile is taken. That constraint has a real architectural consequence: if your pipeline stores per-segment summary percentiles and throws away the distribution, you have permanently lost the ability to roll data up, slice it a new way, or answer a question you did not anticipate. Store distributions, derive percentiles at read time. ## Setting the reporting up so this cannot mislead **Segment by default, with shares visible.** The primary view should be per-device-class p75 with each segment's share of traffic next to it, not a single blended number. When a headline moves, the first question — "did the mix change?" — is then answerable in the same glance instead of requiring an investigation. **Choose few, stable, owned segments.** Device class is nearly always the right first cut, because it is where the largest and most persistent performance gap sits. Add region and page template if they carry real ownership. Resist proliferation: fifty segments means every dashboard has something red, nobody is accountable for anything, and the small ones are pure noise anyway. The test for a segment is whether a named team would act on its regression. **Report a mix-normalised comparison for period-over-period claims.** When the question is "did our work make things faster?", hold the mix fixed — compare each segment against itself, or recompute the current period's blend using last period's traffic weights, and label it clearly as what it is. Keep the true blended number too, because that is the one the outside world assesses you on; just never present the two as if they answer the same question. **Alert on segments, not on the headline.** A blended metric can mask a serious regression in a minority population indefinitely. Alerting per segment catches it, and it also stops teams from being paged for a marketing campaign that changed the traffic mix. **Say what the number means every time you publish it.** "p75 LCP across all page views in the window, mobile 62% of traffic (was 71%)" costs one line and pre-empts the wrong inference. Most misreadings of aggregate performance numbers come from a missing denominator rather than a wrong figure. ## The organisational point The deeper issue is that any single headline number becomes a target, and targets get met by whatever means are available — including means that have nothing to do with engineering. A team optimising a blended metric can improve it by shipping a faster site, or by the traffic mix drifting favourably, and the dashboard cannot tell the two apart. Making the mix visible is what keeps the metric honest; making segments the unit of accountability is what makes it actionable. The blended figure is worth keeping for exactly one purpose — it is the bar you are assessed against externally — and it should never be the number a team is asked to move.
- If the mix changed, is the improved site-wide number wrong?No — it is correct, and it is the number the outside world assesses you on, since a page-experience signal really is computed over whoever actually visited. What is wrong is reading it as evidence that engineering work made the site faster. Both statements can hold at once: the pooled experience improved because the audience changed, and the site regressed for every group of users in it.
- What is wrong with a traffic-weighted average of per-segment p75 values?Percentiles are not additive, so no combination of two percentile values reproduces the percentile of the pooled population — the merged rank depends on the whole shape of both distributions. A weighted average looks rigorous and can be substantially off when the segments differ in shape. The only correct route is to merge the underlying samples, or mergeable distributions, and take the percentile afterwards.
- How many segments should a performance dashboard expose by default?Few enough that each has a named owner who would act on its regression — in practice device class first, then region or page template where those map to real teams. Proliferation is self-defeating: every extra dimension multiplies the panels, guarantees something is always red, and pushes segment sizes down to where the percentiles are noise. Extra dimensions belong in ad-hoc analysis, not the default view.
- What has to be true of the data pipeline for this reporting to be possible at all?It has to retain distributions rather than pre-computed percentiles. If you store only per-segment p75 values you cannot roll them up, re-slice along a dimension you did not anticipate, or recompute a mix-normalised comparison. Storing raw samples or mergeable histograms and deriving percentiles at read time is what keeps every one of those questions answerable later.
saying these in an interview costs you the question
- Calls the segmented numbers a bug because the total improved
- Averages per-segment percentiles to rebuild an overall figure
- Reports a blended metric with no traffic shares attached
- Assumes an aggregate improvement always means faster code
- Adds dozens of segments until no one owns any of them