A CI job that measures page-load metrics on every pull request fails about one run in five with no real regression, and developers have started just re-running it. What causes that variance, and how do you make the check trustworthy again?
answer
- the re-run habit is the finding
- shared runners, variable CPU
- several runs, take the middle one
- pin the environment, stub third parties
- fail on bytes, warn on timings
basics
~20 sShared CI runners give inconsistent CPU, and live networks and third-party scripts add more noise, so single-run measurements scatter. Fix it by running N times and asserting on the median, pinning throttling, serving locally with third parties stubbed, and failing only on deterministic assertions.
solid answer
~50 sThe re-run habit is the real finding: the check has already lost its authority, and every failure from now on is assumed to be noise. The variance comes from measuring on shared runners whose CPU speed varies with whatever else is on the host, over a live network, against a preview deploy that may be cold, with third-party scripts that respond differently each run. I would attack it in that order: run the page several times and assert on the **median** rather than a single run; apply explicit CPU and network throttling so the numbers do not track the runner's mood; serve a production build from a local static server instead of a remote preview; and block or stub third-party requests. Then split the assertions — **fail** on deterministic byte budgets, **warn** on measured metrics — and set metric thresholds with headroom sized from the observed spread, not from the best run I ever saw.
code
json · 11 lines{
"ci": {
"collect": { "numberOfRuns": 5 },
"assert": {
"assertions": {
"largest-contentful-paint": ["warn", { "maxNumericValue": 2500 }],
"total-blocking-time": ["error", { "maxNumericValue": 300 }]
}
}
}
}go deeper
Know that measured page timings vary between runs on shared CI machines, and that repeating the measurement and using the middle value is the standard defence.
Be ready to list the concrete sources of variance — runner CPU contention, live network and cold starts, third-party scripts, non-deterministic page content — and the fix that removes each one.
Show that you quantify the noise floor by running an unchanged commit repeatedly, size thresholds from that spread, and separate deterministic assertions that may block merges from measured ones that only warn.
Treat gate credibility as the asset being managed: track the check's false-failure rate and bypass rate as first-class signals, and be willing to spend runtime and infrastructure to keep a blocking check believable.
## Start by naming the actual failure The symptom to react to is not the 20% failure rate — it is the re-run. Once a team learns that failures are usually noise, the check stops being a gate and becomes a toll booth. A real regression will now be re-run twice, shrugged at, and merged. Any fix has to restore trust, which means the target is not "fewer failures" but "a failure means something". It helps to measure the check itself: run the same unchanged commit through the job ten times and record the spread. That number — the noise floor — is the input to every threshold decision that follows, and it is the evidence that turns "the perf check is flaky" into a fixable engineering problem. ## Where the noise comes from **The runner.** Hosted CI machines are shared and virtualised. Available CPU depends on co-tenants, so the same page can take 40% longer on one run than another. Main-thread-sensitive metrics — blocking time, interaction latency, anything gated on script execution — inherit that variance directly. **The network.** If the job loads a page over the internet, it inherits DNS, TLS, route weather and origin cold starts. If it points at a per-pull-request preview deployment, the first request may hit a cold instance and pay startup cost that has nothing to do with the change. **Third parties.** Analytics, consent banners, tag managers and fonts from other origins respond at different speeds each run, sometimes fail, and sometimes ship different payloads than they did an hour ago. They inject variance the pull request cannot control. **The page itself.** Non-deterministic content — a carousel of the newest items, an A/B assignment, a personalised block, a randomly ordered ad slot — changes what is even being measured between runs. ## Repeat and take the median The single highest-value change is to stop trusting one sample. Load the page N times (three to five is typical) and assert on the **median** run rather than the mean or the best. The median is the right statistic because load-time distributions are right-skewed: one unlucky run with a garbage-collection pause or a noisy neighbour drags a mean upward, while the median ignores it. Taking the best run instead is tempting and wrong — it hides real regressions behind one lucky sample. Most CI-oriented audit runners support this directly; in a Lighthouse CI config it is the collection run count, with each assertion carrying its own severity: ```json { "ci": { "collect": { "numberOfRuns": 5 }, "assert": { "assertions": { "largest-contentful-paint": ["warn", { "maxNumericValue": 2500 }], "total-blocking-time": ["error", { "maxNumericValue": 300 }] } } } } ``` The cost is linear runtime, which is the trade: five runs of one route usually beats one run of five routes when you are trying to make a gate credible. ## Remove variance rather than averaging it away Repetition narrows the spread; elimination is better where it is available. - **Throttle explicitly.** Apply a fixed CPU slowdown multiplier and a fixed network profile so the measurement describes a simulated device, not the runner's current mood. Without it, your numbers track infrastructure. - **Serve locally.** Build once, serve the static output from the same machine, and measure that. It removes the network path and cold starts, and it makes the job reproducible on a laptop. - **Stub third parties.** Block external origins or serve recorded responses. You lose fidelity to the real page — worth remembering when someone asks why the field data disagrees — but you gain a signal that reflects only the code under review. - **Pin the content.** Point the job at a fixed fixture route or a seeded dataset so the page is the same page every time. ## Split deterministic from measured After all that, some noise remains, so grade the assertions by how much you trust them. Byte counts, request counts and chunk sizes come from the build and are bit-for-bit reproducible: those can **fail** the build with a tight threshold. Measured timings still scatter: run them at **warn** level, or fail them only at a threshold set well outside the observed noise floor — if the unchanged-commit spread is ±150 ms, a threshold 100 ms above today's median guarantees false failures forever. ## Catch the drift you deliberately let through Loosening metric thresholds creates a gap: a series of changes each within the noise band can add up to a real regression that no single pull request ever tripped. Close it by recording every run's numbers against the commit and watching the trend on the main branch over weeks. The per-pull-request gate catches cliffs; the trend line catches creep. Together they cover what neither does alone — and the trend chart is also where you notice that the check's own noise floor has started growing, which usually means the runner fleet changed under you.
- Why assert on the median of N runs rather than the mean or the fastest run?Load-time distributions are right-skewed: an unlucky run with a scheduling stall or garbage collection pulls the mean up, so a mean-based gate fails on noise. The fastest run has the opposite flaw — it reports the luckiest conditions and hides real regressions. The median is stable against single outliers in either direction.
- What do you lose by blocking third-party scripts during the CI measurement?Fidelity. Real users load the consent banner, the tag manager and the analytics beacon, and those often dominate main-thread time, so the CI numbers will be optimistic and will not match field data. That is an acceptable trade for a gate — you are measuring the change under review — as long as third-party cost is tracked somewhere else.
- How would you decide whether a metric assertion is stable enough to block merges?Measure it. Run the same unchanged commit through the job ten or more times and look at the spread. If the range is a small fraction of the budget's headroom, it can fail the build; if a threshold that catches real regressions sits inside the noise band, it stays at warn level until the environment is made quieter.
- The job now takes eight minutes because of repeat runs. How do you keep it affordable?Narrow the scope before cutting the repeats — measure one or two representative routes rather than every page, and run the full sweep on the main branch nightly. Keep the cheap deterministic byte checks on every pull request, since they run in seconds, and reserve the repeated page measurements for the routes where a regression would actually matter.
saying these in an interview costs you the question
- Just re-run the job until it goes green
- One measurement per pull request is enough
- Take the fastest run to avoid false failures
- Raise the threshold until nothing fails
- Measure against a live preview deployment for realism