A page scores well when you run a one-off performance test on your laptop, yet its Core Web Vitals field assessment fails. Give the reasons both results can be honest.
answer
- one sample versus a whole population
- your laptop sits in the fast quarter
- the percentile ignores your best visits
- no interactions means no INP
- shifts keep accruing after load
basics
~20 sA single test measures one device, one network, one location and one page state; the field assessment reads the 75th percentile across every real visit, including slow phones and bad networks. Some vitals also cannot be produced by a load-only test at all.
solid answer
~40 sThey measure different populations and, for some metrics, different things entirely. Your test is one sample: one CPU, one connection, one geography, one cache state, no consent banner, no logged-in personalization, no user interaction. The field assessment is the 75th percentile over thousands of real visits that include mid-range phones on congested networks far from your servers — and by construction it ignores your fastest three quarters, which is where your laptop sits. On top of that, `INP` requires real interactions, so a load-only run produces no INP value at all, and `CLS` accumulates across the whole visit including shifts caused by scrolling and late-arriving content, which a test that stops at load never sees. The correct reading is that the lab run is a controlled repeatable diagnostic, not a prediction of the assessment.
go deeper
Understand that a test on your own machine is one visit under favourable conditions, while the published numbers come from many real visitors on slower devices and networks.
Explain the mechanics of the gap: a single sample versus a percentile over a population, plus the fact that a load-only run generates no interactions and stops before later layout shifts occur.
Demonstrate the diagnosis: identify the failing metric and form factor, reproduce under representative CPU and network constraints with a real user journey, segment your own field data to find the failing population, and verify at the assessment's percentile.
Decide what role synthetic testing plays organization-wide — regression gate versus source of truth — and make sure teams are held to real-user outcomes rather than to a number produced on developer hardware.
## Two numbers, two populations The most important thing to say out loud is that these are not two measurements of the same quantity that disagree. A one-off test produces **one sample** under conditions you chose. The field assessment produces a **percentile over a population** you did not choose. Even if both are perfectly accurate, there is no reason they should agree, and a senior answer starts there rather than hunting for a bug. Your laptop, on office wi-fi, near your CDN's edge, with a warm DNS and TLS cache, is not a typical visit. In the field distribution it is somewhere in the fast end — and the assessment reads the 75th percentile, which is deliberately positioned to represent the slower part of your audience. You are literally looking at the part of the distribution the rule was designed to ignore. ## What a single run cannot contain - **Device spread.** Real traffic includes mid-range and older phones whose CPUs take several times longer to parse, compile and execute the same JavaScript. Script-heavy pages diverge most here. - **Network spread.** Congested mobile networks, high latency, packet loss, captive portals. Latency-sensitive work like discovering a resource late is punished far harder in the field. - **Geography.** A visitor several thousand kilometres from the nearest edge pays round-trip costs you never see locally. - **Cache state.** Your repeated runs may be warm; a large share of real visits are first-time and cold. - **Page state.** Real users see the consent dialog, the logged-in header, the personalized module, the A/B variant, the ad slot that filled late. Many of these arrive on the page after load and are exactly the things that damage vitals. - **Extensions and third parties in the wild.** Real browsers carry extensions and real sessions hit third-party endpoints that are sometimes slow. ## Two of the vitals are not load-time properties at all This is the part that separates a good answer from a great one. **INP measures responsiveness to real interactions.** A test that loads a page and stops produces no interactions, so it produces no INP value — not a good one, not a bad one, none. Any local confidence you have about interaction latency comes from a proxy or from you clicking around by hand on fast hardware. A page can load beautifully and still respond terribly once a heavy handler runs on a slow phone. **CLS accumulates over the whole visit.** Layout shifts caused by content that arrives as the user scrolls, by an image slot that fills late, by a sticky element that installs after hydration — all of that is in the field number and none of it is in a run that finishes at load. Field CLS is routinely worse than a lab figure for exactly this reason. ## Diagnosis, not despair The practical response is to make your local conditions resemble the failing part of the field distribution rather than your own machine: 1. Look at *which* metric and *which* form factor fails. Almost always it is mobile, and almost always one metric carries the failure. 2. Reproduce under representative constraints: throttle CPU and network to something like a mid-range phone on a mediocre connection, run cold, and exercise the page the way a user does — scroll it, interact with it, accept the consent dialog. 3. Segment your own real-user data to find the population that is failing: device class, connection, country, entry point, logged-in state. The failure is rarely uniform; it is usually concentrated somewhere identifiable. 4. Verify the fix against that population, at the same percentile the assessment uses. ## The reverse case Be ready for the mirror image too: a mediocre local score alongside a passing field assessment. That happens when your audience is concentrated on fast devices and good networks, when most visits are warm-cache returning users, or when the local run happened to hit a cold cache and an unlucky third party. It is a reminder that a single synthetic run is a **diagnostic instrument** — repeatable, controllable, great for comparing two builds under identical conditions — and never an oracle for what real users experience. ## The sentence to land on "My test is one fast sample; the assessment is a percentile over everyone, and two of the three metrics need a real session to exist at all." Everything else in the answer is elaboration of that.
- How would you make a local run better resemble the visits that are failing?Constrain it to look like the failing segment: throttle CPU and network to a mid-range phone on a mediocre connection, start from a cold cache, use a mobile viewport, and drive the page as a user does — accept the consent dialog, scroll to trigger lazy content, interact with the controls. Then compare builds under those fixed conditions rather than against an absolute target.
- Your field CLS is far worse than anything you can reproduce locally. Where do you look first?At everything that happens after load: content that fills in as the user scrolls, late third-party or ad slots, elements that install or resize after hydration, and fonts swapping in on slow connections. Field CLS accrues across the whole visit, so reproduce by scrolling the page slowly on a throttled connection instead of measuring only the initial load.
- Is a lab run useless, then?No — it is the only measurement you can hold constant. Because it fixes device, network and page state, it is ideal for comparing two builds, catching regressions in CI, and attributing a change to a specific commit. Its failure mode is being read as a prediction of field results. Use it for relative comparison, and real-user data for absolute judgment.
saying these in an interview costs you the question
- Assumes a green local run means the field assessment will pass
- Thinks a load-only test can produce an INP value
- Believes field CLS should match a lab measurement
- Ignores form factor when comparing lab and field results
- Blames the field data as inaccurate rather than different