You own web performance for a large site. How would you divide responsibility between scheduled synthetic monitoring, your own real-user monitoring, and public field data, and what does each one fail at?
answer
- assign decisions to signals, not tools
- field defines the goal
- synthetic guards the deploy loop
- public data is the scorecard, not the console
- a score as a KPI invites gaming
basics
~20 sUse scheduled synthetic runs to catch regressions on known journeys under controlled conditions, first-party RUM as the authoritative measure of what users actually get, and public field data as the external scorecard. Each covers a blind spot of the other two.
solid answer
~50 sI would make first-party RUM the source of truth for goals: it is the only signal that measures every real session and can be sliced by template, release, device and market, so performance targets and any alerting hang off it. Scheduled synthetic monitoring is the fast, reproducible early-warning system — same journeys, same profile, every deploy — and it is the only thing that works pre-release, on low-traffic pages, and on flows behind a login. Public field data is the external scorecard: comparable across companies, feeding the assessments other people judge us by, worth watching but never a debugging tool. The failure modes are as important as the coverage. Synthetic drifts from reality and tempts teams to optimise a score; RUM tells you a number moved but rarely why, and it lags the deploy; public data lags by weeks and has no attribution at all. The strategy is to state which signal is authoritative for which decision, and never let a lab score become the goal.
go deeper
Know that the three sources exist and answer different questions: scheduled tests you run, measurements from your own users, and public data about real users that anyone can read.
Explain what each signal can and cannot do — reproducibility and pre-release coverage from synthetic, population coverage and slicing from RUM, comparability and lag from public data.
Show how you would wire them together operationally: synthetic as the fast post-deploy check, RUM as the goal and the segment finder, lab traces for diagnosis, and coverage of the measurement itself as a tracked property.
Own the incentive design as much as the tooling. Decide which signal each decision hangs off, defend why the objective is stated in field terms, and be explicit about where you are choosing not to measure and what that risks.
## Start from the decisions, not the tools A measurement strategy is not a list of tools; it is an assignment of decisions to signals. There are roughly four decisions a large site has to make repeatedly: 1. **Are we meeting our performance goal for real users?** 2. **Did this change make it worse?** 3. **Why is this slow?** 4. **How do we compare with the rest of the market, and with the yardstick our distribution channels use?** Each signal answers one or two of those well and the rest badly. ## First-party RUM: the source of truth RUM you collect yourself is the only signal that covers every real session and can be cut the way your organisation is actually shaped: by page template, by release, by device class, by market, by logged-in state, by experiment arm. That makes it the natural home for the goal itself. If a performance target is written down anywhere, it should be a field statement — "the checkout template meets its loading target for the slow end of phone sessions" — because that is the only phrasing that means something to a user. What it fails at: attribution and latency. A dashboard says the number moved; it does not usually say which resource arrived late. Good RUM narrows that by carrying context with each sample — the identity of the element that produced the metric, navigation type, whether the session was truncated — but it will not replace a trace. It also needs traffic, so new and low-volume pages are invisible, and it only ever tells you about code that has already shipped to real users. ## Scheduled synthetic monitoring: the early warning Synthetic runs are the fast loop. Because the device, network and page state are pinned, a change in the number is attributable to a change in the site, and the answer arrives within minutes of a deploy rather than days later. Synthetic is also the *only* option in three situations that matter at scale: before release, on pages with too little traffic for field statistics, and on flows behind authentication that public datasets never see. Running the same handful of critical journeys on a schedule, from a few locations that match where the audience actually is, catches a large share of regressions before users meet them. What it fails at: representativeness and incentives. The profile you pinned is an invention, and it drifts from your real audience as devices, third parties and markets change. Third parties that are stubbed or blocked in the test environment do not exist in the run but do exist for users. And a synthetic score is dangerously easy to turn into a target — the moment a composite score becomes a team KPI, work flows toward moving the score rather than moving the user experience. ## Public field data: the external scorecard Public field data is comparable across organisations, costs nothing to obtain, and is what external assessments are computed from — which makes it a business input rather than an engineering one. It is also the only way to see a competitor's real numbers, which is occasionally decisive in an argument about investment. What it fails at: everything operational. It lags by weeks because of its aggregation window, it covers a browser-restricted and opt-in population, it excludes anything behind a login, it drops long-tail URLs to origin-level aggregates, and it carries no attribution whatsoever. Watching it as a control panel produces teams that discover problems a month late and cannot explain them. ## The division of labour | Decision | Authoritative signal | | --- | --- | | Are we meeting the goal? | First-party RUM | | Did this deploy regress? | Synthetic, confirmed later by RUM | | Why is it slow? | Lab trace, cohort-matched to the RUM segment | | How do we look externally? | Public field data | ## Where strategies fail - **Goodharting the lab score.** A composite lab score is a diagnostic aid. Made a KPI, it rewards changes that suit the test profile — deferring work the test does not exercise, stripping content the test happens to weight — and none of it reaches users. - **Instrumenting everything and acting on nothing.** Coverage without ownership produces dashboards nobody reads. Each metric should have a team that owns it and a decision it feeds. - **Alerting on thin slices.** Field alerts on low-traffic pages fire on noise; teams then learn to ignore alerts. Alert on aggregates with enough volume to be stable, and let synthetic carry the per-page early warning. - **No attribution in RUM.** Collecting only the headline number guarantees that every investigation restarts from zero. The context you attach at collection time is what makes the field signal actionable. - **Ignoring measurement coverage.** If a meaningful share of navigations never reports, the field numbers describe survivors. Coverage deserves its own tracked metric. ## The one-sentence version Field data defines the goal and the problem, synthetic guards the pipeline and gives fast feedback, public data keeps you honest against the outside world — and no lab number is ever allowed to be the objective.
- Why is a composite lab performance score a poor team objective?Because it is computed from a fixed test profile using a fixed weighting, so it can be improved by changes that suit the test rather than the user — deferring work the run never exercises, or trimming what the weighting happens to punish. It is also a single sample, blind to the distribution across real devices. Keep it as a diagnostic, and set objectives in field terms.
- Where would you deliberately choose not to invest in measurement?On pages with too little traffic for field statistics and no business weight — internal utilities, rarely used legal pages, low-value long tail. Cover them with one scheduled synthetic run for gross regressions and spend the instrumentation and alerting budget on the few templates carrying most sessions and revenue. Measurement has an ongoing cost in noise and attention, not just in storage.
- How do you keep synthetic monitoring from drifting away from your real audience?Re-derive the test profile from field data on a schedule. Take the device classes, connection characteristics and locations that dominate your real traffic and set the run's throttling and test locations to match the slow end of that population, not the median. Also keep third parties and consent flows enabled in the run, since disabling them quietly makes the test a different page from the one users get.
- Two teams disagree about whether a release regressed performance. Which signal settles it?Field data sliced by release, once enough sessions have accumulated — that is the only evidence about real users. Synthetic settles it faster and is usually right about direction, so it is the practical tiebreaker in the hours after a deploy, but if the two disagree the field number wins and the synthetic profile is what needs re-examining.
saying these in an interview costs you the question
- Makes a synthetic score the team's performance KPI
- Relies on public field data as the day-to-day console
- Alerts on field metrics for low-traffic pages
- Collects headline metrics with no attribution context
- Treats more dashboards as equivalent to better measurement