skip to content

Your real-user monitoring breaks Interaction to Next Paint down by page template, and one template reports a p75 INP of 900ms from only 40 recorded interactions that week. Why should you be careful acting on that number, and what would you do before treating it as a regression?

level: middleimportance: should knowfreq 36%

answer

  1. how many samples above the cut?
  2. p75 of 40 is a handful of events
  3. noise scatters, regressions step
  4. put the count next to the number
  5. forty events is few enough to read

basics

~20 s

A percentile from 40 samples rests on a handful of observations, so it carries a wide margin of error and can swing week to week without anything changing. Widen the window, check the trend and the raw slow samples, and confirm the regression before acting on it.

solid answer

~50 s

With 40 samples, p75 is essentially the tenth-slowest interaction — move two unlucky observations and the number jumps hundreds of milliseconds. That is sampling noise, not evidence. Before I call it a regression I'd do three things. Widen the aggregation window until the sample count is in the hundreds and see whether 900ms persists. Plot the trend across several windows rather than comparing two points, since a real regression appears as a step or a slope and noise appears as scatter. And look at the individual slow interactions — which element, which target, which segment — because a handful of genuinely broken interactions is a real bug even when the percentile around them is unreliable. If it survives all three, it is real. What I would not do is either act on it as fact or dismiss small segments outright, because that is how you never see the templates with few but important users.

go deeper

for a junior

Know that a percentile computed from very few samples is unreliable and can swing on its own, and that you should check how many measurements the number came from before reacting to it.

for a middle

Explain why stability depends on how many observations sit above the cut, why that makes p95 and p99 far hungrier for data than p75, and what widening the window buys and costs.

for a senior

Show the practical workflow: trend over several windows rather than a two-point comparison, sample counts and confidence shown on every breakdown, and going to the individual slow events, which are often decisive even when the aggregate is not.

for a principal

Own the policy split between alerting and investigation — noisy segments must not page anyone, but they must not become invisible either — and decide where synthetic coverage substitutes for field data that will never reach volume.

## Why small samples make percentiles unstable A percentile is a ranked observation, so its precision depends on how many observations sit around that rank. With 40 samples, p75 is roughly the 30th value in sorted order — meaning only about ten samples lie above it. Swap two of those for slightly different values and the reported percentile can move dramatically. The estimate is not wrong exactly; it just has an error bar wide enough to swallow the effect you think you are seeing. Two distinct effects are at work. **Sampling variability**: your 40 interactions are a random draw from the population of interactions that template could receive, and different draws give different percentiles. **Granularity**: with few samples the percentile can only take one of a few discrete values, so it jumps between them rather than moving smoothly. A chart of a small-sample percentile over time looks like a seismograph even when the underlying experience is unchanged. The severity scales with the percentile. Roughly, the number of samples *above* the cut is what stabilises the estimate: p75 leaves 25% above, p95 leaves 5%, p99 leaves 1%. To get a p95 as stable as a given p75 you need something like five times the data, and a p99 twenty times. This is the practical reason the industry reports vitals at p75 rather than deeper in the tail, and it is why per-template or per-country breakdowns of high percentiles are frequently meaningless on all but the largest sites. ## What to do instead **Widen the window.** If a week gives 40 samples, try 28 days. You trade responsiveness for stability: the number stops jumping, but a real change takes longer to show up. That is usually the right trade for a low-traffic segment, and it is exactly why public field datasets use a rolling multi-week window in the first place. **Look at a trend, not a comparison.** Two data points cannot distinguish a step change from scatter. Six or eight consecutive windows can: a regression shows up as a level shift or a slope that persists, noise shows up as values bouncing around a stable centre. Plotting the sample count on the same chart is worth doing, because a percentile that moved at the same moment its sample count collapsed is telling you about traffic, not latency. **Show the uncertainty.** If your tooling can attach a confidence interval to the percentile, do it — "900ms, ±350ms" starts a very different conversation from "900ms". Where that is not available, displaying the sample count next to every percentile is the cheap substitute, and it should be non-negotiable on any breakdown dashboard. A number with no denominator invites people to act on noise. **Go to the raw events.** This is the step people skip, and it is often the most informative. Forty interactions is few enough to inspect individually. If eight of them are the same element on the same template taking over a second, you have found a real bug regardless of what the percentile's error bar says — the percentile was never the evidence, it was the pointer. If the slow ones are scattered across unrelated elements, devices and sessions, the aggregate is probably just noise. **Aggregate upward.** Group the template with similar templates, or report it as part of a section rather than alone, until the bucket has enough traffic to sustain a percentile. You lose specificity; you gain a number you can act on. **Reproduce deliberately.** A slow interaction that is real should be reproducible on a throttled device with the same input. If a targeted attempt to reproduce it succeeds, the sample size question is moot — you have a defect. ## The mistake in the other direction The correct lesson is "weight this number by its uncertainty", not "ignore small segments". Low-traffic segments are frequently the ones that matter most: a new market, a checkout step, an admin surface used by a small number of high-value customers, or a template that is low-traffic *because* it is slow. If your dashboard hides everything below a sample threshold, those areas become permanently invisible and will never be prioritised. The workable policy is to treat small segments as leads rather than verdicts. Filtering them out of *alerting* is reasonable, since you cannot page someone on noise. Filtering them out of *investigation* is not. And for anything you genuinely cannot measure well in the field, use targeted synthetic checks to give the segment a stable signal that field sampling cannot provide at that volume.

  • Roughly how much more data does a stable p95 need compared with a stable p75?
    Around five times as much, because stability comes from the number of observations above the cut: p75 leaves a quarter of samples above it, p95 only a twentieth. A p99 needs roughly twenty times. That ratio is why deep-tail percentiles are only meaningful on high-traffic aggregates, and why per-segment breakdowns should stay at p75 unless the segment is very large.
  • What is the cost of just widening the aggregation window until every segment looks stable?
    Detection latency and dilution. A four-week window mixes pre- and post-release traffic, so a regression is muted and takes weeks to fully appear, and a fix looks like it did nothing. It also smooths over genuinely short-lived incidents. The usual compromise is a short window for alerting on high-traffic aggregates and a long window for reporting on small segments.
  • When would you deliberately keep watching a segment with too little data for a reliable percentile?
    When the segment matters more than its volume — a checkout step, a newly launched market, a template used by a small number of high-value accounts, or one that is low-traffic precisely because it performs badly. For those, drop the percentile as the primary signal and use raw event inspection plus targeted synthetic checks, which give a stable reading that field sampling cannot at that volume.

saying these in an interview costs you the question

  • Treats any reported percentile as fact regardless of sample count
  • Believes percentiles are invalid below some fixed sample count
  • Compares two windows and calls the difference a regression
  • Hides all low-traffic segments from every view, not just alerts
  • Assumes small samples bias the percentile in a predictable direction

context