skip to content

Thresholds and Scoring

A metric number means nothing without the scoring rules behind it. Know why vitals are reported at the 75th percentile, how field and lab scores diverge, and what "passing" actually requires.

on this pageshow

questions

5

Core Web Vitals are reported from real-user field data. What does it actually take for a URL to "pass" the Core Web Vitals assessment?

level: middleimportance: must knowfreq 72%

answer

  1. real visits, not a test you ran
  2. every metric must qualify
  3. the 75th percentile decides
  4. mobile and desktop scored apart
  5. trailing 28-day rolling window

basics

~20 s

A URL passes only when every Core Web Vital — LCP, INP and CLS — sits in the good band at the 75th percentile of real visits, evaluated separately for mobile and desktop over a trailing 28-day window of field data.

solid answer

~50 s

Passing is a conjunction, not a score. For each of the three Core Web Vitals — LCP, INP and CLS — you take the distribution of real visits to that URL and read the 75th percentile; that value must land at or below the metric's `good` boundary. All three must qualify: two good metrics and one in `needs improvement` is a fail, and there is no averaging or partial credit. The evaluation is done per form factor, so a URL can pass on desktop and fail on mobile, and the mobile result is usually the one that matters. The input is field data over a rolling 28-day window from opted-in Chrome users, not a test you ran yourself — which is why a fast local run proves nothing about the assessment. This assumes the post-2024 metric set, where INP replaced First Input Delay.

go deeper

for a junior

Be able to name the three Core Web Vitals and say that the numbers come from real users' visits rather than a test you run, and that each metric is rated good, needs improvement, or poor.

for a middle

Explain the mechanics: fixed boundaries per metric, the value read at the 75th percentile of the visit distribution, all three metrics required, computed separately for mobile and desktop over a rolling 28-day window.

for a senior

Show you can act on it — identify which metric and which form factor is actually failing, resist spending effort on metrics already in the good band, and explain why the slow tail of real devices is what you must move.

for a principal

Own the framing: decide whether the assessment is the right organizational target at all, what internal targets sit alongside it for pages or browsers the dataset never sees, and how to stop teams optimizing a public verdict at the expense of experiences it does not measure.

## What is being assessed Core Web Vitals is a small, fixed set of user-experience metrics that Chrome collects from real page loads and publishes as aggregated field data. As of the 2024 metric set there are three: **LCP** (Largest Contentful Paint, a loading metric), **INP** (Interaction to Next Paint, a responsiveness metric that replaced First Input Delay), and **CLS** (Cumulative Layout Shift, a visual-stability metric). The "assessment" is a single pass/fail verdict derived from those three, and the rules that turn raw numbers into that verdict are what interviewers are probing here. ## Bands, not a score Each metric has two fixed boundaries, which split its range into three bands: **good**, **needs improvement**, and **poor**. The boundaries are the same for every site on the web — they are absolute targets, not a curve relative to your competitors or your own history. (For example, LCP's good boundary sits at 2.5 seconds.) There is no 0–100 composite here and no weighting between the metrics: the output is which band each metric landed in. ## The 75th percentile rule One visit produces one value per metric. A real URL produces thousands of visits with a wide spread — a warm-cache visit on a desktop over fibre and a cold visit on a mid-range Android over a congested mobile network are both in there. Core Web Vitals reports the **75th percentile** of that distribution: the value at or below which three quarters of visits fell. That choice is deliberate. A median would let a large unhappy minority hide behind a fast majority; a very high percentile would be dominated by outliers and network freak events you cannot fix. p75 says the typical experience *and* most of the slow tail must be acceptable, and the slowest quarter is the only part allowed to be bad. One practical consequence trips people up constantly: if 74% of your visits are in the good band for a metric, your p75 is not in the good band, and you fail on that metric. "Mostly green" is not passing. ## All three, or none The verdict is a logical AND. LCP good, CLS good, INP needs-improvement means the URL does not pass. Nothing is averaged, nothing is traded off, and a spectacular LCP does not buy you slack on CLS. In practice this means your effort belongs on the *worst* metric, not the one you find most interesting — improving a metric that is already comfortably good moves the verdict by exactly nothing. ## Segments are assessed separately Field data is bucketed by form factor, and the assessment is computed per bucket. It is entirely normal for a URL to pass on desktop and fail on mobile, because mobile visits carry slower CPUs, slower networks and smaller viewports. Reporting that averages the two together is hiding your actual problem. Treat mobile as the default target and desktop as the easier case, and when someone tells you "we pass Core Web Vitals", ask which form factor they are looking at. ## The window Field reports aggregate a **rolling 28-day window**. Every day the window slides: one day of old visits leaves and one day of new visits enters. That is why a fix does not show up immediately — the window still contains mostly pre-fix visits for weeks — and why the reported value drifts rather than steps. ## Where the data comes from, and who is missing The field dataset is collected from Chrome users who opted into usage reporting, so Safari and Firefox traffic is simply absent, and a URL with too little qualifying traffic gets no record of its own at all. That is not a bug in your instrumentation; it is a property of the dataset. If a URL you care about has no published record, you either read the origin-level aggregate or you collect your own real-user data. ## What to say in an interview The crisp version is: three metrics, each with a fixed good boundary, each read at the 75th percentile of real Chrome visits, all three must be good, computed separately per form factor, over a trailing 28-day window. Then add the judgment: because the rule is a conjunction on a percentile, the only work that changes the verdict is work that pulls the slow tail of your worst metric across a fixed line — which usually means fixing what low-end devices and slow networks experience, not what your laptop experiences.

  • If a URL's metric is comfortably good at p75 already, is there any assessment value in improving it further?
    None for the verdict — the assessment is a threshold check, so extra headroom on an already-good metric changes nothing. The value is defensive: headroom protects you against regressions, seasonal traffic shifts toward slower devices, or a new third-party script eating the margin. But when you are choosing work to move a failing assessment, the only leverage is on the metric that is failing.
  • What happens to the assessment for a page that nobody interacts with, so it never produces an INP sample?
    That page simply has no INP value to evaluate, and the verdict is formed from the metrics that do have data. It is a real situation for content pages a user reads and leaves. Do not read a missing INP as a passing INP — it means the metric was never exercised, and the moment you add interactive elements you are measuring something you have never seen.
  • Why are the band boundaries the same for every site rather than relative to a site's peers or its own history?
    Because the metrics describe the user's experience, and a user's tolerance for a slow load does not scale with how slow your competitors are. Absolute boundaries make the target stable, comparable across sites, and impossible to game by picking a weak peer group. The cost is that some genuinely hard categories of page find the bar harsh.

saying these in an interview costs you the question

  • Says the three metric scores are averaged into one rating
  • Thinks a fast local test run means the URL passes
  • Assumes the median visit determines the rating
  • Believes passing on desktop covers mobile too
  • Expects the field assessment to move the day after a deploy

context

open as a page

A Core Web Vitals field report shows each metric as three bars labelled good, needs improvement and poor. What is being counted in those bars, and where do the band boundaries come from?

level: juniorimportance: should knowfreq 46%

basics

~20 s

Each Core Web Vital has two fixed boundaries that split its values into good, needs improvement and poor. Every real visit produces one value that lands in one band, and the bars show the share of visits in each band.

open as a page

You shipped a change that clearly cut LCP, but a week later the Core Web Vitals field data for that page has barely moved. Why, and what would you look at in the meantime?

level: middleimportance: should knowfreq 48%

basics

~20 s

Field Core Web Vitals are aggregated over a rolling 28-day window, so a week after a deploy roughly three quarters of the samples still come from the old code. The reported value drifts as the window rolls; your own real-user data shows the change the same day.

open as a page

A page scores well when you run a one-off performance test on your laptop, yet its Core Web Vitals field assessment fails. Give the reasons both results can be honest.

level: seniorimportance: should knowfreq 58%

basics

~20 s

A single test measures one device, one network, one location and one page state; the field assessment reads the 75th percentile across every real visit, including slow phones and bad networks. Some vitals also cannot be produced by a load-only test at all.

open as a page

A performance report shows Core Web Vitals field data for one of your pages but labels it origin-level, with no URL-level record available. Why does that happen, and how should you read that data when deciding what to fix?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Field datasets publish a URL-level record only when that URL has enough qualifying visits in the window; low-traffic pages get none and fall back to the origin aggregate, which is traffic-weighted across every page on the site and therefore describes your busiest templates, not this page.

open as a page