skip to content

Measurement and Monitoring

Optimizing without measuring is guessing. This group covers gathering performance data in the lab and from real users, reading it honestly, and stopping regressions before they ship.

on this pageshow

explore

questions

26

What is a performance budget in a web project's CI pipeline, and what has to be true of it before it actually prevents regressions?

level: juniorimportance: must knowfreq 50%

answer

  1. a number agreed before the work
  2. checked automatically on every change
  3. report versus blocked merge
  4. bytes of output, or a measured metric
  5. measure today, add headroom, ratchet down

basics

~20 s

A performance budget is a limit agreed in advance on a measurable property of the built site — bytes shipped, request count, or a metric such as LCP — that a CI job checks on every pull request and fails the build when exceeded.

solid answer

~50 s

A performance budget turns a vague goal like "the site should feel fast" into a specific number that a machine can check. Budgets come in two families: **quantity budgets** over the build output (compressed bytes of initial JavaScript, number of blocking requests, image weight per route) and **metric budgets** over a measured page load (LCP, total blocking time, layout shift). The budget only works if three things are true: it runs automatically on every pull request, it is measured against a production build of a fixed set of routes, and the job exits non-zero so the merge is blocked. A budget that only posts a comment or updates a dashboard is a report, not a budget — it gets skimmed and the regression lands anyway. Starting numbers usually come from measuring today's site and adding a little headroom, then tightening over time.

code

json · 6 lines
json
{
  "size-limit": [
    { "name": "main bundle", "path": "dist/main.js", "limit": "120 kB" },
    { "name": "initial CSS", "path": "dist/*.css", "limit": "20 kB" }
  ]
}

go deeper

for a junior

Be ready to define a budget as a pre-agreed limit checked automatically on every change, and to name two things you can put a number on: compressed bytes of JavaScript shipped, and a measured metric such as LCP.

for a middle

Expect to explain where the check sits in the pipeline, that it must run over a production build of named routes, and why the job has to exit non-zero rather than post a comment that people skim.

for a senior

Show that you can pick starting numbers from current measurements plus headroom, separate deterministic byte assertions from noisy measured ones, and make the failure message name the file that grew so the author can act without asking you.

for a principal

Own the policy side: who sets a budget, who may raise one and what they owe in exchange, how budgets ratchet downward as the site improves, and what you do when a valuable feature genuinely does not fit.

## What a budget actually is A performance budget is a number, agreed before the work starts, that some measurable property of the shipped site is not allowed to exceed. "Initial JavaScript on the product page stays under 170 kB compressed" is a budget. "We care about performance" is not, because nothing can check it. The point of putting it in CI is that performance decays by accident. Nobody ships a pull request called "make the site slower"; they ship a date picker, an analytics tag, a polyfill, an icon set. Each one is defensible on its own and each one costs a few kilobytes or a few milliseconds. Six months of individually-reasonable pull requests is how a fast site becomes a slow one. A budget is the mechanism that makes the cost visible at the moment someone can still choose differently — in review, not six months later in a performance sprint. ## Two families of budget **Quantity budgets** are asserted over the build artifacts. Compressed size of the entry chunk, total JavaScript on a route, number of render-blocking requests, total image bytes, size of the CSS in the head. These are deterministic: the same commit produces the same number on any machine, so the check never flakes and the failure is easy to attribute — a named file grew by a named amount. **Metric budgets** are asserted over a measured page load: largest contentful paint, total blocking time, cumulative layout shift, time to first byte. These are closer to what a user experiences, but they are measurements, so they carry run-to-run noise and need care (repeat runs, fixed throttling) before they can block a merge. Most teams run both, because they answer different questions. A byte budget catches the cause; a metric budget catches the effect, including effects that ship no new bytes at all — a hero image that stopped being preloaded, a font that started blocking, an animation that began thrashing layout. ## Where the check runs The check belongs in the pull-request pipeline, after a **production build**. Running it against a development build measures the wrong artifact entirely: unminified code, no tree shaking, dev-only warnings. It should measure a **fixed list of representative routes** — typically the entry page plus the one or two routes that carry the most traffic or the most code — because a budget on "the site" cannot be attributed to anything, and per-route budgets can be owned by the team that owns the route. A minimal quantity check needs very little machinery. A tool such as `size-limit` reads a config listing paths and limits, measures the compressed output, and exits non-zero when a path is over: ```json { "size-limit": [ { "name": "main bundle", "path": "dist/main.js", "limit": "120 kB" } ] } ``` The exit code is the whole point. In a hosted CI system, a job that exits non-zero is a failed check, and a failed required check blocks the merge. ## Choosing the first numbers The common mistake is picking an aspirational number, watching every pull request fail, and switching the check off within a fortnight. The workable sequence is: measure what the site does today, set the budget slightly above it so the current state passes, and then **ratchet** — whenever a change lowers the real number, lower the budget to match, so the improvement cannot silently be spent again. The opposite mistake is setting the budget at exactly today's value. Then ordinary noise, or a legitimate three-kilobyte change, fails the build, and the team learns that failures are meaningless. ## Warn versus fail Not every assertion deserves to block a merge. A useful split is to **fail** on assertions that are deterministic and clearly owned — byte counts, request counts — and to **warn** on assertions that are measured and noisy while you are still learning their spread. A warning is visible in review and can be argued about; a failure stops the line. Reserving failure for checks you trust completely is what keeps people from reflexively re-running or bypassing the job. ## What a budget cannot do A budget is a ceiling, not a target. Passing it means the change did not make things worse than the agreed limit; it does not mean the page is fast, and it does not tell you what real users on real devices are experiencing. It also cannot decide policy for you: someone still has to answer what happens when a genuinely valuable feature does not fit. Answering that in advance — who may raise a budget, and what they must do in exchange — is what separates a budget that survives a year from a check that quietly gets deleted.

  • Why measure the budget against a production build rather than the development server?
    A development build is a different artifact: unminified, unsplit, carrying dev-only warnings and no tree shaking. Its numbers are both much larger and unrelated to what users download, so a budget over it either passes trivially or fails permanently. The check has to run over exactly the artifact that would be deployed.
  • Where would you set the very first budget numbers for a site that has never had one?
    Measure the current production build, then set each budget a little above what it does today so the existing state passes. That makes the first green build honest. From there, ratchet: whenever a change brings the real number down, lower the budget to match, so the gain cannot be spent again by accident.
  • If you could only afford one budget assertion on a marketing site, which would you pick?
    Compressed bytes of JavaScript on the entry route. It is deterministic, so it never flakes; it is the thing that grows by accident most often; and it drives both download time and main-thread work, so it correlates with the metrics you actually care about while being far cheaper to check than measuring them.

saying these in an interview costs you the question

  • A budget is a dashboard target, not something enforced
  • One global budget covers the whole site
  • Setting the budget to today's exact number
  • Measuring the development build instead of the production output
  • Passing the budget means the page is fast

context

open as a page

In web performance work, what is the difference between lab (synthetic) data and field data (real-user monitoring), and what question is each one good at answering?

level: juniorimportance: must knowfreq 65%

basics

~20 s

Lab data comes from a page load you generate yourself on a chosen device, network and cache profile, so it is repeatable and debuggable. Field data comes from real visitors' browsers, so it is messy but describes what users actually experienced.

open as a page

A page scores in the high 90s in a local Lighthouse run, but the Core Web Vitals field data for the same page is failing. Explain how both can be true, and how you would reconcile them.

level: middleimportance: must knowfreq 72%

basics

~20 s

A lab run measures one scripted load on a device, network and cache profile you chose; field data aggregates every real visit on real hardware. Both can be accurate because they describe different populations and, for some metrics, different definitions.

open as a page

Core Web Vitals are assessed at the 75th percentile of visits rather than at the median or the 95th. What does "p75 LCP is 2.4 seconds" actually tell you about your users, and why is p75 the chosen cut point?

level: middleimportance: must knowfreq 68%

basics

~20 s

A p75 LCP of 2.4 seconds means three quarters of measured page loads reached their largest contentful paint at or before 2.4 seconds, and one quarter were slower. p75 is chosen because it represents most visits while staying far enough from the extreme tail to be stable and actionable.

open as a page

An analytics module creates a PerformanceObserver for the 'largest-contentful-paint' entry type, but on fast page loads its callback never fires. Which observe() option was most likely omitted, and what does that option do?

level: middleimportance: must knowfreq 58%

basics

~20 s

The observe() call omitted buffered: true. A PerformanceObserver normally only sees entries created after it registers, so an observer starting late misses events the browser already recorded. buffered: true replays matching entries already sitting in the performance timeline, and it only works with the single-type form of observe().

open as a page

A page takes about four seconds before it is usable. Looking at a browser performance trace of that load, how do you tell whether the main thread was busy or sitting idle waiting on the network, and how does the answer change which fix you reach for?

level: middleimportance: must knowfreq 56%

basics

~20 s

Compare the main-thread track with the network track: solid scripting blocks mean the CPU is the bottleneck, while an idle main thread under a long request bar means you are waiting on bytes. Each points at a different family of fixes.

open as a page

A stakeholder says a page in your web app "feels slow". Walk through the profiling workflow you would follow in the browser, from that complaint to a verified fix.

level: middleimportance: must knowfreq 68%

basics

~20 s

Turn the complaint into one reproducible scenario and one number, record a browser performance trace under representative conditions, attack the single dominant cost the trace shows, then re-record the same scenario and compare. One change per measurement.

open as a page

A CI job that measures page-load metrics on every pull request fails about one run in five with no real regression, and developers have started just re-running it. What causes that variance, and how do you make the check trustworthy again?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Shared CI runners give inconsistent CPU, and live networks and third-party scripts add more noise, so single-run measurements scatter. Fix it by running N times and asserting on the median, pinning throttling, serving locally with third parties stubbed, and failing only on deterministic assertions.

open as a page

A dashboard reports your site's average (mean) Largest Contentful Paint as 2.1 seconds, yet users keep saying pages feel slow. Why can a healthy-looking average LCP hide a real problem, and what should you report instead?

level: juniorimportance: should knowfreq 50%

basics

~20 s

Load times are right-skewed — a fast majority plus a long slow tail — so the mean sits near the fast cluster and hides the users having the worst time. Report percentiles such as p75 and p95 instead, plus the shape of the distribution.

open as a page

In a browser, how do you use performance.mark() and performance.measure() to time a block of your own application code, and where do the results end up?

level: juniorimportance: should knowfreq 52%

basics

~10 s

performance.mark() records a named timestamp; performance.measure() creates a named duration between two marks. Both become entries in the browser's performance timeline, readable with performance.getEntriesByType('mark'/'measure') or a PerformanceObserver, and visible in a DevTools performance recording.

open as a page

Before recording a page-load trace in the Chrome DevTools Performance panel, which recording conditions do you set, and why is an unthrottled run on a developer laptop with a warm cache a misleading baseline?

level: juniorimportance: should knowfreq 48%

basics

~20 s

Record a production build with CPU throttling, network throttling and the cache disabled, in a clean browser profile with extensions off, using the reload-and-record option so the trace covers startup. Otherwise the trace measures your laptop, not your users.

open as a page

Should a CI performance budget be an absolute limit or a comparison against the base branch's measurement? Explain the tradeoff, and how a team keeps the number from drifting upward over time.

level: middleimportance: should knowfreq 38%

basics

~20 s

Use both. An absolute limit expresses the contract with users but blames whoever happens to cross it; a delta against the base branch attributes the cost to the pull request that added it but permits unlimited slow creep. Together with a ratchet, they close each other's gap.

open as a page

A team blocks merges on Lighthouse's overall performance score staying at or above a fixed number. What does that gate miss, and what would you assert on instead?

level: middleimportance: should knowfreq 45%

basics

~20 s

A single composite score blends several lab metrics, so opposing regressions cancel out, a small change can swing the score, and a failure never says what broke. Assert on individual metrics and on byte counts instead, and keep the score as a summary only.

open as a page

What is the Chrome UX Report (CrUX), where does its data come from, and what are its limits as a source of field performance data for a site you own?

level: middleimportance: should knowfreq 50%

basics

~20 s

CrUX is Google's public dataset of real-user Core Web Vitals, gathered from opted-in Chrome users on desktop and Android and aggregated over a rolling 28-day window per origin, and per URL where a page has enough traffic. It reports what happened, never why.

open as a page

Your real-user monitoring breaks Interaction to Next Paint down by page template, and one template reports a p75 INP of 900ms from only 40 recorded interactions that week. Why should you be careful acting on that number, and what would you do before treating it as a regression?

level: middleimportance: should knowfreq 36%

basics

~20 s

A percentile from 40 samples rests on a handful of observations, so it carries a wide margin of error and can swing week to week without anything changing. Widen the window, check the trend and the raw slow samples, and confirm the regression before acting on it.

open as a page

How do you read a document's server response time and DOMContentLoaded timing from the browser's Navigation Timing API, and why is the legacy performance.timing object a poor substitute?

level: middleimportance: should knowfreq 40%

basics

~20 s

Read the single PerformanceNavigationTiming entry from performance.getEntriesByType('navigation'). Its timestamps are already relative to the document's time origin, so responseStart is the byte-arrival time and domContentLoadedEventEnd is that milestone directly. The deprecated performance.timing returned Unix-epoch values requiring manual subtraction and could not be observed.

open as a page

In a browser, PerformanceResourceTiming entries for third-party scripts and images report 0 for connectStart, requestStart, responseStart and transferSize, while same-origin entries are fully populated. Why, and what makes the real numbers visible?

level: middleimportance: should knowfreq 42%

basics

~20 s

Detailed Resource Timing attributes are cross-origin restricted: for a cross-origin resource the browser zeroes the connection, request and size fields unless the response carries a Timing-Allow-Origin header naming your origin (or *). Only startTime, responseEnd, duration and the resource name remain trustworthy without it.

open as a page

In a Chrome DevTools performance trace, the main-thread flame chart shows one very wide bar whose children fill almost its entire width. What does a bar's width mean here, and how do the Self time and Total time figures tell you which function is actually worth optimizing?

level: middleimportance: should knowfreq 54%

basics

~20 s

Width in a flame chart is wall-clock duration and depth is call nesting. Total time counts a function plus everything it called; Self time counts only its own work. Optimize where Self time concentrates, not the wide parent containing it.

open as a page

Your real-user monitoring dashboard shows healthy loading metrics, yet the same pages have a high abandonment rate and users report the site as slow. How can RUM systematically under-report the worst experiences, and what would you change?

level: seniorimportance: should knowfreq 40%

basics

~20 s

RUM only records sessions where the measurement code loaded and the page survived long enough to send a report, so abandoned loads, blocked scripts and killed tabs contribute nothing. That biases the dataset toward users who already had a decent experience.

open as a page

After a release, your real-user monitoring shows p75 LCP improved from 3.1s to 2.6s while p95 LCP got worse, from 6.2s to 8.0s, and support tickets about slowness went up. How do you interpret that, and how would you decide whether the release was a win?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The release helped typical loads and hurt the slowest ones — a classic tradeoff that shifts cost onto weak devices and networks. Decide by segmenting the distribution to find which population regressed and how many visits it covers, not by comparing two headline numbers.

open as a page

A RUM script observes 'largest-contentful-paint' and 'event' entries in the browser and POSTs each entry to an analytics endpoint as it arrives. Why does that produce wrong Core Web Vitals numbers, and what is the correct collection pattern?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Vitals are not single events. LCP emits successive candidates until the first interaction, and INP is derived by grouping many 'event' entries by interactionId, so per-entry reporting sends drafts, not results. Accumulate in memory, finalise once, and send with sendBeacon when the page is hidden.

open as a page

A PerformanceObserver for the 'longtask' entry type reports many 50ms-plus tasks in production, but the entries do not reveal which code is responsible. What does a longtask entry actually give you, and what does the 'long-animation-frame' entry type add?

level: seniorimportance: should knowfreq 30%

basics

~20 s

A longtask entry gives you only a start time, a duration and a coarse attribution to the containing frame or iframe — never a script or function. The newer long-animation-frame entry type reports the whole slow frame, including a scripts array with each script's source URL, invoker and duration.

open as a page

After a code change, your second performance trace of the same page load is 400 ms faster than the first. What could make that comparison wrong, and how would you establish that your change actually caused the improvement?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A single before/after pair proves little: run-to-run variance, cache state, background load and different code paths all move the number. Re-run both builds several times under identical interleaved conditions, compare medians, and check that the specific cost you targeted actually shrank.

open as a page

You are introducing performance budgets across a dozen product teams' repositories. How do you roll them out so the checks are still enforced a year later rather than disabled or routinely bypassed?

level: principalimportance: should knowfreq 28%

basics

~20 s

Treat the gate's credibility as the thing being managed. Start in warn mode with budgets derived from today's measurements, make the check fast and low-noise before it blocks anything, give each budget a named owner and an exemption path with an expiry, then track bypass rate as a health signal.

open as a page

You own web performance for a large site. How would you divide responsibility between scheduled synthetic monitoring, your own real-user monitoring, and public field data, and what does each one fail at?

level: principalimportance: should knowfreq 36%

basics

~20 s

Use scheduled synthetic runs to catch regressions on known journeys under controlled conditions, first-party RUM as the authoritative measure of what users actually get, and public field data as the external scorecard. Each covers a blind spot of the other two.

open as a page

Your site-wide p75 LCP improved over the last quarter, but when you split the same data by device class, p75 LCP got worse for mobile and worse for desktop. As the person who owns performance reporting for the organisation, how do you explain that, and how would you set reporting up so leadership is not misled by it again?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

The traffic mix shifted toward the faster segment, so the blended number improved while both segments regressed — a mix-shift effect. Fix the reporting, not the arithmetic: publish segmented percentiles with their traffic shares beside the headline, and never derive an aggregate by combining per-segment percentiles.

open as a page