skip to content

Lab vs Real-User Monitoring

Lab tests are repeatable but synthetic; field data is messy but real. Interviewers want to hear when you trust each one and what you do when the two disagree.

on this pageshow

questions

5

In web performance work, what is the difference between lab (synthetic) data and field data (real-user monitoring), and what question is each one good at answering?

level: juniorimportance: must knowfreq 65%

answer

  1. one load you made vs many they made
  2. repeatable versus representative
  3. cause versus effect
  4. cold cache, pinned device, no interaction
  5. field covers users you never test

basics

~20 s

Lab data comes from a page load you generate yourself on a chosen device, network and cache profile, so it is repeatable and debuggable. Field data comes from real visitors' browsers, so it is messy but describes what users actually experienced.

solid answer

~50 s

Lab data is synthetic: a tool loads the URL on a machine and profile you control — fixed CPU and network throttling, usually a cold cache, no extensions, no real interaction — and hands you a full trace. Because you control every variable, it is the right tool for *why* a page is slow and for comparing a before and an after. Field data is collected inside real visitors' browsers as they use the site and aggregated across all of them, so it covers devices, networks, geographies, cache states and page states you would never think to script. It is the right tool for *whether* users are actually slow and *who* is affected. The rule of thumb: field data tells you there is a problem and how big it is, lab data tells you what to fix. Neither replaces the other, and a green lab score is not evidence that users are having a good time.

go deeper

for a junior

Be able to state plainly that lab data is a page load you generated under conditions you chose, and field data comes from real visitors. Give one example of a question each answers.

for a middle

Explain the mechanics behind the difference: fixed throttling and cold cache versus real devices, caches and interactions, and why one is a single sample while the other is a distribution over a population.

for a senior

Show the working loop — field data identifies the affected cohort, a lab reproduction matched to that cohort finds the cause, and field data confirms the fix. Say what you do when the reproduction does not reproduce.

for a principal

Own which signal is authoritative for what. Argue why a performance goal should be stated in field terms while pre-release gating is lab-based, and where reporting the wrong one distorts team behaviour.

## Two ways to get a performance number Every performance number a team looks at comes from one of two places. **Lab data**, also called synthetic data, is produced by deliberately loading the page yourself. A tool opens the URL on a machine you control, with a device profile, a network profile and a cache state that you chose, records everything that happened, and prints metrics. Scheduled synthetic monitoring, a local audit run, and a page load you record by hand are all lab data. **Field data**, also called RUM (real-user monitoring), is produced by real visitors. Measurement code runs inside their browsers while they use the site normally, and each visit reports what it observed to a collector, where the values are aggregated into a distribution across many sessions. The distinction is not about which tool you use — it is about who generated the page load. If you generated it, it is lab. If a real person on their own device generated it, it is field. ## What a lab run buys you - **Repeatability.** The device, the network and the cache state are pinned, so two runs are comparable and a change in the number is attributable to a change in the page. - **Attribution.** A synthetic run can keep a full trace: the request waterfall, which resource was render-blocking, which element painted last, where the main thread was busy. Field data usually carries only the final numbers plus a little context. - **It works before release.** You can measure a branch, a preview deployment or a staging environment — anywhere with no real traffic at all. - **It works on pages nobody visits.** A rarely used checkout step still gets measured. ## What a lab run cannot see A scripted load is one page view under conditions you invented. It does not include the visitor's five-year-old phone, their congested mobile network, their ad blocker or their extensions, their warm cache from yesterday, their logged-in personalised page, the consent banner they see and you do not, the A/B variant they were bucketed into, their distance from your servers, or the fifteen other tabs competing for their CPU. Above all, a load test contains no real interaction: nobody taps a menu or types in a field, so responsiveness under real use is simply absent from the recording. ## What field data buys you Field data is the only evidence of what users actually got. It covers the whole population — every device class, network, country, entry page and cache state — and it accumulates over the real lifetime of a page rather than stopping when a script decides the load is over. Because it is a distribution rather than a single value, it also shows how bad the unlucky visits are, not just the typical one. Business decisions and performance goals should be anchored to it. ## What field data cannot do Field data is an effect, not a cause. It tells you that the metric is bad for phone users in a particular market; it rarely tells you which request arrived late or which script blocked the thread. It needs real traffic, so it says nothing about a page before launch or about a page with a handful of visits a week. It arrives after the fact, often aggregated over days or weeks, so it is a slow feedback loop. And it is only as complete as your instrumentation: sessions where the measurement code never ran contribute nothing. ## Which question goes where | Question | Source | | --- | --- | | Are our users having a slow experience? | Field | | Which cohort is worst off? | Field | | Did this pull request make the page heavier or slower? | Lab | | Why is this page slow — what is on the critical path? | Lab | | Did the fix we shipped last month help real users? | Field | | Is the unreleased redesign faster than what we have? | Lab | ## How they fit together in practice The healthy loop runs in one direction and then the other. Field data raises the alarm and names the cohort. You then reproduce that cohort in the lab — throttle to a comparable device and network, test from a comparable location, start from a comparable cache state — and use the trace to find the cause. You fix it, verify the fix in the lab because that feedback is immediate, ship it, and then wait for field data to confirm that real users moved. If the field number does not move, the lab reproduction was not the real cohort's problem. ## Where people go wrong The two classic mistakes are treating a lab score as the user experience — it is one synthetic sample of a population you invented — and treating field data as self-explanatory, then guessing at causes because the dashboard has no trace attached. Both signals are cheap; the skill is knowing which one you are allowed to draw a conclusion from.

  • Your synthetic tests run from a machine in the same region as your servers. What does that setup hide?
    Everything distance-related. Time to first byte, connection setup and any uncached origin fetch will look far better than they do for a user on another continent or a congested mobile network. A test run next to the origin measures the application, not the delivery path, so it can be consistently green while overseas users wait seconds for the first byte.
  • Is there a case where you would trust lab data over field data for a decision?
    Yes — whenever the decision is about a change rather than a state. Before a release there is no field data at all, so a controlled before-and-after run is the only evidence. The same goes for low-traffic pages, where the field sample is too small to say anything, and for isolating one variable, which real traffic never lets you do.
  • Name a condition real users hit routinely that a standard synthetic load test never reproduces.
    Several: a warm HTTP cache from an earlier visit, a logged-in and personalised page, a consent or cookie banner shown only in some regions, an ad blocker removing third-party scripts, back/forward navigations, and a device that is thermally throttled or sharing CPU with other tabs. Each can move a metric in either direction.

saying these in an interview costs you the question

  • Says a high lab score proves users have a fast experience
  • Thinks field data is just an average of synthetic runs
  • Believes lab tools are unnecessary once RUM exists
  • Assumes a local test on a laptop represents a phone user
  • Cannot say what a synthetic run is unable to observe

context

open as a page

A page scores in the high 90s in a local Lighthouse run, but the Core Web Vitals field data for the same page is failing. Explain how both can be true, and how you would reconcile them.

level: middleimportance: must knowfreq 72%

basics

~20 s

A lab run measures one scripted load on a device, network and cache profile you chose; field data aggregates every real visit on real hardware. Both can be accurate because they describe different populations and, for some metrics, different definitions.

open as a page

What is the Chrome UX Report (CrUX), where does its data come from, and what are its limits as a source of field performance data for a site you own?

level: middleimportance: should knowfreq 50%

basics

~20 s

CrUX is Google's public dataset of real-user Core Web Vitals, gathered from opted-in Chrome users on desktop and Android and aggregated over a rolling 28-day window per origin, and per URL where a page has enough traffic. It reports what happened, never why.

open as a page

Your real-user monitoring dashboard shows healthy loading metrics, yet the same pages have a high abandonment rate and users report the site as slow. How can RUM systematically under-report the worst experiences, and what would you change?

level: seniorimportance: should knowfreq 40%

basics

~20 s

RUM only records sessions where the measurement code loaded and the page survived long enough to send a report, so abandoned loads, blocked scripts and killed tabs contribute nothing. That biases the dataset toward users who already had a decent experience.

open as a page

You own web performance for a large site. How would you divide responsibility between scheduled synthetic monitoring, your own real-user monitoring, and public field data, and what does each one fail at?

level: principalimportance: should knowfreq 36%

basics

~20 s

Use scheduled synthetic runs to catch regressions on known journeys under controlled conditions, first-party RUM as the authoritative measure of what users actually get, and public field data as the external scorecard. Each covers a blind spot of the other two.

open as a page