skip to content

A page scores in the high 90s in a local Lighthouse run, but the Core Web Vitals field data for the same page is failing. Explain how both can be true, and how you would reconcile them.

level: middleimportance: must knowfreq 72%

answer

  1. one invented load versus every real one
  2. different populations, sometimes different metrics
  3. the audit never taps anything
  4. field CLS keeps accumulating after load
  5. match the cohort, then reproduce

basics

~20 s

A lab run measures one scripted load on a device, network and cache profile you chose; field data aggregates every real visit on real hardware. Both can be accurate because they describe different populations and, for some metrics, different definitions.

solid answer

~50 s

They are not measuring the same thing. The lab run is one load, on one emulated device, from one location, with a cold cache, no extensions, no logged-in state, no consent banner, no ad blocker and — crucially — no user interaction. Field data is the aggregate of every real session, and the Core Web Vitals verdict is taken from the slow end of that distribution rather than the typical case. There are definitional gaps too: a load audit reports Total Blocking Time, not INP, because it never interacts with the page; and field CLS accumulates over the whole page lifetime, including shifts from scrolling and lazily inserted content that a run ending at load never sees. To reconcile, I first check I am comparing like with like — same URL rather than the whole origin, same form factor — then segment the field data to find who is slow, rebuild those conditions in a lab run, and only then look for a cause.

go deeper

for a junior

Know that a test you run yourself uses conditions you chose, while field data comes from real devices, and be able to name two conditions real users have that a test does not.

for a middle

Explain the concrete mismatches: throttling profile, cold versus warm cache, missing third-party and consent scripts, no interactions during a load audit, and layout shift that keeps accumulating in the field.

for a senior

Demonstrate the reconciliation path end to end — confirm you are comparing the same scope, find the failing cohort, rebuild those conditions, drive real interactions, and verify with field data after the deploy window has rolled.

for a principal

Be ready to argue which number the organisation is accountable for. A lab score is a debugging instrument, not a goal; making it a target invites work that moves the score without moving any user.

## Why a green lab score and red field data are compatible The instinct is that one of the two numbers must be wrong. Usually neither is. They answer different questions about different populations, and for some metrics they are not even measuring the same quantity. ## The environment gap A lab run pins the environment. A mobile audit uses an emulated device profile with CPU and network throttling; a run on a laptop with no throttling is faster still. Real visitors bring conditions the run never had: - **Devices.** Mid and low-tier phones parse and execute JavaScript several times slower than a developer machine, and a thermally throttled or multi-tasking phone is slower again. - **Networks.** Real latency, packet loss and captive portals are worse than a clean simulated profile. - **Distance.** A run from your own region hides the time to first byte that a user on another continent pays. - **Page state.** Real sessions may be logged in, personalised, bucketed into an A/B variant, or carrying a consent banner, a chat widget and tag-manager scripts that never fire in your test. - **Client software.** Extensions, ad blockers and injected content change both what loads and how much main-thread work there is. ## The definitional gap Even with an identical environment, some metrics cannot match. - **INP.** A load audit performs no genuine interactions, so it cannot produce an Interaction to Next Paint value. Lighthouse reports **Total Blocking Time** instead, a lab proxy for main-thread congestion during load. TBT can be modest while real INP is dreadful, because the expensive handlers run when a user taps a filter, opens a menu, or types — long after the audit stopped. - **CLS.** In the field, layout shift accumulates across the entire page lifetime, including shifts caused by scrolling into lazily loaded images, late ads, and content inserted after the load event. A run that stops shortly after load records only the early shifts. - **LCP element.** The element that paints last can differ between the two. A returning user with a warm cache, a personalised hero, a cookie banner covering the viewport, or a different viewport size can all move the LCP candidate to something else entirely. - **Navigation type.** Field data contains restores from the back/forward cache, prerendered navigations and reloads. Synthetic runs are cold navigations only. ## The population gap A lab run yields a single value. Field data yields a distribution across every session, and the Core Web Vitals verdict is taken from the slow end of it, not the middle. A site whose typical visit is fast can still fail if a meaningful slice of sessions — old Android devices, a distant market, one heavy template — is far slower. One green run tells you nothing about the shape of that tail. ## Reconciling them, in order 1. **Compare like with like.** Is the field number for this URL or for the whole origin? Origin-level data pools every page, so a fast landing page can be dragged down by an unrelated section. Also match form factor: phone data against a phone profile, not a desktop run. 2. **Check freshness.** Public field datasets aggregate a rolling multi-week window. A fix shipped a week ago is only partially represented, and a regression that started yesterday is barely visible yet. 3. **Find who is slow.** Split the field data by device class, connection and geography until the failing slice appears. "Everyone is slow" and "the oldest Android devices in one market are slow" call for completely different fixes. 4. **Rebuild those conditions.** Throttle to a comparable device and network, test from a comparable region, and include the things real users get: consent banner accepted or not, tag manager live, logged-in state if applicable, third-party scripts unblocked. 5. **Test the state, not just the cold load.** If the failing metric is INP, drive the interactions users actually perform. If it is CLS, scroll the page and let the lazily loaded content arrive. 6. **Only now look for a cause.** With a reproduction in hand, the lab trace tells you which request was discovered late or which task blocked the thread. ## What not to conclude Do not "fix" the disagreement by tuning until the lab score is 100. The score is a weighted blend of lab metrics under a fixed profile; it can be improved in ways that do not touch what any user experiences. And do not dismiss the field data as noise — it is the population you actually serve. Where the two disagree, the field number defines the problem and the lab run is the instrument you use to explain it.

  • Which lab metric is the closest proxy for field INP, and where does the proxy break down?
    Total Blocking Time is the usual stand-in: it sums main-thread blocking during load, so a page with heavy hydration or a huge script tends to show both a high TBT and a poor INP. The proxy breaks down in both directions — an app can load lightly and then run an expensive handler on every keystroke, or load heavily and be perfectly responsive once idle, because INP is dominated by post-load interaction work.
  • You ship an LCP fix and the field report is unchanged two days later. Is the fix wrong?
    Not necessarily. Public field datasets aggregate a rolling window of several weeks, so a change is diluted by older sessions and only fully visible once the window has rolled over. Your own RUM, which you can slice by deploy time, moves far sooner. Before assuming failure, confirm the fix reached production for the affected cohort and check the newest days in your own data.
  • Field CLS is much worse than lab CLS on the same page. What is the usual cause?
    Shifts that happen after the point where the run stopped. Real users scroll, and lazily loaded images without reserved dimensions, late-arriving ads, injected banners and font swaps all move content further down the page. Field CLS covers the whole page lifetime, so it captures those; a load audit that ends shortly after load does not.
  • Would you ever ignore a bad field number for a page?
    Only after checking the sample. A page with very few eligible sessions can produce an unstable number that swings on a handful of unlucky visits, and origin-level data attributed to a page you did not change is a reporting artefact rather than a regression. That is a reason to look at the right slice, not a reason to dismiss the dataset.

saying these in an interview costs you the question

  • Claims the field data must be stale or broken
  • Tries to fix the disagreement by chasing a score of 100
  • Assumes a load audit measures interaction responsiveness
  • Compares origin-level field data against one URL's lab run
  • Ignores device and network differences between test and users

context