skip to content

A page visit contains many clicks and key presses, but INP reports one number. How is that value chosen from all the interactions, and why can a single very slow interaction in a long session fail to appear in it?

level: seniorimportance: should knowfreq 38%

answer

  1. not an average, near the maximum
  2. users remember the worst click
  3. one discard per fifty interactions
  4. long sessions collide with noise
  5. rare-and-slow can hide

basics

~20 s

INP reports the slowest interaction of the visit, not an average — but on interaction-heavy pages it discards roughly one outlier per fifty interactions, so on a long session a single freak stall can be the discarded one and never surface.

solid answer

~50 s

INP is a worst-case metric by design: for most visits it is simply the longest interaction latency, because users judge an app by the click that hung, not by the mean. The one refinement is outlier tolerance. On pages with a lot of interaction, the value discards approximately one outlier for every fifty interactions, so a session with a hundred and twenty interactions can drop its two slowest and report the third. That protects the metric from a single unlucky stall — a background process on the user's machine, a garbage-collection pause — turning an otherwise responsive session red. The cost is that a genuinely broken control used once in a long, busy session can be masked. That is a reason to keep per-interaction telemetry alongside the headline metric, not a reason to distrust it: the discarding only applies to heavy sessions, and a control that is reliably slow shows up as soon as more than one user hits it.

go deeper

for a junior

Know that INP reports one value per visit and that it comes from the slowest interactions rather than an average of them.

for a middle

State the rule precisely — the worst interaction, with roughly one discard per fifty interactions — and explain that the allowance protects long sessions from environmental noise.

for a senior

Show you understand the tradeoff in operation: a rare, genuinely broken interaction in a busy session can be masked, so keep per-interaction telemetry with target attribution beside the headline metric.

for a principal

Own the two-level story when reporting upward — a near-worst value per visit, then a high percentile across visits — so the dashboard and individual user complaints can be reconciled instead of used to discredit each other.

## Why worst-case rather than average Averaging interaction latency would be statistically tidy and practically useless. A user who clicks a hundred times and waits 900 ms once does not remember the ninety-nine fast clicks; they remember the moment the app appeared to freeze. INP is built on that observation, so its selection rule reaches for the bad end of the distribution rather than the middle. ## The selection rule For a visit with a modest number of interactions, INP is the **longest** interaction latency in that visit. Simple. For interaction-heavy visits, a small amount of outlier tolerance kicks in: roughly **one interaction is disregarded for every fifty** in the visit, and the value reported is the worst of what remains. Practically: - Up to about 50 interactions → the single worst one is reported. - Around 100 interactions → the worst is dropped, the second-worst reported. - Around 150 interactions → the two worst are dropped, the third-worst reported. The effect is a high percentile of the visit's interaction latencies rather than the strict maximum. ## Why the tolerance exists A browser tab does not have exclusive use of the machine. Over a long session, some interaction will eventually collide with something outside the page's control: a garbage-collection pause, the operating system swapping, another tab waking up, an antivirus scan, a device thermally throttling. Without tolerance, the metric would effectively measure "did anything unlucky happen in the last twenty minutes", and the longer a user stayed, the worse the page would score — punishing exactly the engagement most products want. The tolerance scales with interaction count for the same reason: more interactions mean more chances to be unlucky, so the allowance grows with the exposure. ## The blind spot this creates The cost is real. Consider an admin console where the user performs two hundred routine interactions and once opens a report dialog that takes 1.2 seconds. Two hundred interactions buys roughly four discards, so if the dialog is the only badly slow interaction and the others are comfortable, that 1.2 s can be dropped and the reported INP looks healthy. The user's genuine complaint is invisible in the headline number. Three things keep this from being a serious problem in practice: 1. **It only applies to heavy sessions.** Most visits never reach fifty interactions, so their worst interaction is reported unchanged. 2. **Aggregation rescues it.** The page-level score is a high percentile across many visits. A control that is reliably slow will be the *worst surviving* interaction in plenty of sessions, so it surfaces once more than a handful of users touch it. Only something both slow and extremely rare hides. 3. **You can measure interactions directly.** The headline metric is an aggregate; the Event Timing API still reports every slow interaction individually, so a field-monitoring setup that records per-interaction latency and its target element sees the dialog regardless of whether it won the visit. ## How this stacks with field aggregation It helps to keep two levels of selection straight. **Within a visit**, INP picks a near-worst interaction as described. **Across visits**, field reporting takes a high percentile of those per-visit values over a rolling window. So the published number is a percentile of near-maxima — deliberately pessimistic at the visit level and deliberately robust at the population level. Confusing the two produces claims like "INP is an average", which it is not at either level. ## What to say in an interview Name the rule, name the reason, and name the cost. The strong answer is not just "it takes the worst" — it is "it takes the worst, with an outlier allowance that scales with interaction count, because long sessions inevitably collide with environmental noise, and the tradeoff is that a rare-but-terrible interaction in a busy session can be masked, which is why I keep per-interaction telemetry next to the metric."

  • Why not simply report the strict maximum interaction latency for every visit?
    Because a long session will eventually collide with something outside the page's control — a GC pause, OS swapping, another tab waking up. A strict maximum would make the score deteriorate purely as a function of session length, punishing engaged users and telling you nothing about the page. The outlier allowance keeps the metric about the page.
  • A support ticket describes a 1.5-second dialog, but your INP dashboard is green. How do you reconcile that?
    The two are measuring different things. INP is a near-worst value per visit, then a high percentile across visits — a rare interaction in busy sessions can be discarded or drowned out. Query per-interaction telemetry filtered to that control instead of the headline metric; if it is reliably slow it will be plainly visible there.
  • Does the outlier allowance apply to a visit with only a handful of interactions?
    No. Below roughly fifty interactions there is no discard and the single worst interaction is reported. That is the common case for most page visits, which is why the allowance is easy to overlook — it only shapes the numbers for genuinely interaction-heavy sessions like editors and admin consoles.

saying these in an interview costs you the question

  • Says INP averages the interactions in a visit
  • Thinks every visit reports its strict maximum with no tolerance
  • Claims a slow control can hide from field data entirely
  • Confuses per-visit selection with the cross-visit percentile
  • Believes longer sessions should naturally score worse

context