skip to content

A dashboard reports your site's average (mean) Largest Contentful Paint as 2.1 seconds, yet users keep saying pages feel slow. Why can a healthy-looking average LCP hide a real problem, and what should you report instead?

level: juniorimportance: should knowfreq 50%

answer

  1. load times are not bell-shaped
  2. one long tail, no fast tail
  3. which user is at 2.1 seconds?
  4. the mean sits above p75 here
  5. report p50, p75, p95 and a histogram

basics

~20 s

Load times are right-skewed — a fast majority plus a long slow tail — so the mean sits near the fast cluster and hides the users having the worst time. Report percentiles such as p75 and p95 instead, plus the shape of the distribution.

solid answer

~50 s

Page-load timings are not normally distributed. Most visits cluster in a narrow fast band, and a minority — low-end phones, cold caches, weak networks, contended CPUs — stretch out into a long right-hand tail. A single mean collapses that whole shape into one number that describes nobody in particular: it is dragged upward by the tail but still nowhere near it, so it can look acceptable while a meaningful slice of visits is several seconds slower. Percentiles describe actual people. `p75 = 2.4s` means three out of four measured loads finished at or under 2.4 seconds; `p95` tells you what the unlucky twentieth visit experienced. In practice I'd report p50, p75 and p95 side by side, plus a histogram, so I can see whether a change moved the whole distribution or only part of it.

code

javascript · 10 lines
javascript
const lcp = [1.1, 1.2, 1.2, 1.3, 1.4, 1.5, 1.6, 1.8, 2.0, 9.5];

const mean = lcp.reduce((a, b) => a + b, 0) / lcp.length;
const sorted = [...lcp].sort((a, b) => a - b);
const pct = (p) => sorted[Math.ceil((p / 100) * sorted.length) - 1];

console.log('mean', mean.toFixed(2)); // 2.26
console.log('p50', pct(50));          // 1.4
console.log('p75', pct(75));          // 1.8
console.log('p95', pct(95));          // 9.5

go deeper

for a junior

Be ready to say plainly that a mean is pulled around by a few very slow loads and describes no actual user, and that percentiles like p75 do. Knowing what "p75 = 2.4s" means in words is the whole bar here.

for a middle

Explain why load-time data is right-skewed, why that makes the mean sit between the typical case and the bad case, and what a histogram shows that any single statistic cannot — including the bimodal case where the median lands in an empty valley.

for a senior

Show that you drive decisions from the distribution: quantify how many real visits a regression touched, resist trimming the tail that generates support tickets, and pick the percentile that matches the question you are actually answering.

for a principal

Own how performance is reported across an organisation: which statistics appear on the shared dashboard, whether teams are held to a central number or a tail number, and the fact that any single headline figure will be gamed unless the distribution is visible next to it.

## The shape of real load-time data If you plot every measured LCP for a real site, you almost never get a symmetric bell curve. You get a **right-skewed** distribution: a tall cluster of fast loads, then a long tail trailing off to the right. Physically this makes sense. There is a floor on how fast a page can possibly be — you cannot beat the speed of light to the origin, the parse cost of the HTML, the decode of the hero image — but there is no ceiling on how slow it can get. A five-year-old phone on a congested cell network, with a cold cache, behind a captive portal, running three other tabs, can take twenty seconds. Nothing balances that out on the fast side. ## Why the mean fails on skewed data The arithmetic mean adds every value and divides by the count. That works well when values are symmetric around a centre, because then the mean *is* the typical case. On skewed data it is neither the typical case nor the bad case: - It is **pulled upward** by the tail, so it overstates what most visitors see. - It is still **far below** the tail, so it understates what the slowest visitors see. - It is **not robust**: one 30-second outlier in a thousand samples shifts it measurably, so it wobbles for reasons that have nothing to do with your code. Concretely, take ten loads at 1.1, 1.2, 1.2, 1.3, 1.4, 1.5, 1.6, 1.8, 2.0 and 9.5 seconds. The mean is 2.26 s. But nine of the ten loads were under 2.1 s, and the tenth was 9.5 s. There is no user whose experience the number 2.26 describes. Notice too that the mean here is *higher* than the 75th percentile (1.8 s) — on skewed data that is normal, and it is exactly why a mean can simultaneously look mediocre and hide something much worse. ## What a percentile is The *p*-th percentile is the value below which *p* percent of the observations fall. Sort the samples and read off the value at the right rank: ```js const sorted = [...samples].sort((a, b) => a - b); const pct = (p) => sorted[Math.ceil((p / 100) * sorted.length) - 1]; ``` So `p50` (the median) is the middle load, `p75` is the load that three quarters of visits beat, `p95` is the load that only one visit in twenty is worse than. Each of these statements is about *people*, not about arithmetic. That is the property the mean lacks and the reason performance is reported in percentiles across the industry. Two corollaries worth knowing. First, the median is robust — it does not care how extreme the extremes are, only how many of them there are — which is why p50 is a better "typical user" number than the mean. Second, high percentiles cost data: only 5% of your samples sit above p95, so with a small sample that number bounces around. ## What to report instead A good performance report is not one number, it is a small set that describes the shape: - **p50** — the typical visit. Answers "is the site normally fast?" - **p75** — the number Core Web Vitals is assessed on, and a reasonable "most users are at least this fast" bar. - **p95 or p99** — the tail. Answers "how bad is it for the people having the worst time?" - **A histogram or a good/needs-improvement/poor bucket split** — this is what actually shows you whether the distribution is one hump or two. The bucket view matters more than people expect. A bimodal distribution — say, a fast cached path and a slow uncached path — has a p50 sitting in the empty valley between the two humps, describing a visit that essentially never happens. Only the histogram reveals that, and it usually points straight at the cause: a cache-miss path, a device-class split, a slow region, a heavy page template. ## The mistakes this prevents Reporting only a mean leads to three predictable errors. You celebrate a release that lowered the mean while making the tail worse. You ignore a real regression because it only hit 10% of visits and got averaged away. And you argue with support, who hear from exactly the users the mean is hiding — the tail is the part of the distribution that files tickets. Percentiles keep the conversation anchored on how many real people are affected and by how much, which is the only framing that supports a decision about what to fix.

  • If the mean is misleading, why not just report the median and be done with it?
    The median is robust but blind to the tail: you can make a quarter of your visits dramatically worse without moving p50 at all. It answers "is the typical visit fast?" and nothing else. Pairing it with a higher percentile — p75 for the bar you are held to, p95 for the pain — plus a bucket split gives you both the centre and the shape.
  • What does a distribution with two distinct humps usually tell you?
    That two different populations are mixed into one chart, and the summary statistics sit in the valley between them describing nobody. The usual causes are a cached versus uncached path, a device-class split between low-end phones and desktops, or two page templates reported together. The fix is to split the data along the suspected dimension and confirm each hump becomes a single population.
  • Should you strip outliers out of the data before computing these numbers?
    Almost never for user-timing data — the outliers are the users you most need to fix. Trimming is only defensible for values that are not real experiences at all, such as timings from a background tab or an obviously corrupt clock reading, and even then you should log how many you dropped. "Remove the outliers" is usually a way of deleting the evidence.

Put ninety-nine ordinary earners and one billionaire in a room and the average income is enormous, but nobody in the room is rich. The mean describes the room; the percentiles describe the people in it.

saying these in an interview costs you the question

  • Says the average is fine because most users are fast
  • Treats the mean and the median as interchangeable
  • Assumes load times are normally distributed
  • Discards slow samples as outliers or bots
  • Reports a single number with no distribution or counts

context