skip to content

A team blocks merges on Lighthouse's overall performance score staying at or above a fixed number. What does that gate miss, and what would you assert on instead?

level: middleimportance: should knowfreq 45%

answer

  1. one number, several inputs
  2. opposing changes net out
  3. steep curves, uneven sensitivity
  4. the author cannot act on it
  5. gate on specifics, chart the composite

basics

~20 s

A single composite score blends several lab metrics, so opposing regressions cancel out, a small change can swing the score, and a failure never says what broke. Assert on individual metrics and on byte counts instead, and keep the score as a summary only.

solid answer

~50 s

The overall score is a weighted blend of several lab metrics, so it is a bad gate for three reasons. First, it **cancels**: a bundle that grew and an image that shrank can net out to the same score while the page got worse for anyone on a slow CPU. Second, it is **non-linear** — the score curves are steep in the middle, so a change worth 100 ms can move the score several points on one page and none on another. Third, it is **undiagnosable**: "score dropped from 92 to 88" tells the pull-request author nothing they can act on. I would assert on the specific things I care about — largest contentful paint, total blocking time, layout shift, and compressed bytes per route — each with its own threshold, so a failure names the thing that regressed. The composite score is still useful as a trend line on a dashboard; it is just the wrong unit for a merge gate.

code

json · 9 lines
json
{
  "resourceSizes": [
    { "resourceType": "script", "budget": 170 },
    { "resourceType": "total", "budget": 450 }
  ],
  "timings": [
    { "metric": "largest-contentful-paint", "budget": 2500 }
  ]
}

go deeper

for a junior

Know that a composite performance score is several measurements blended into one number, and that a lower score does not tell you which part of the page got worse.

for a middle

Be ready to explain cancellation and the non-linear scoring curve, and to list the specific assertions you would gate on instead — per-route byte counts plus individual metrics.

for a senior

Demonstrate that you choose assertions by whether a failure is attributable and actionable by the pull-request author, and that you know which metrics a single lab load can and cannot produce.

for a principal

Frame it as a measurement-design problem: aggregates hide regressions behind unrelated gains, so define which quantity each team is accountable for and keep composites in the reporting layer rather than the enforcement layer.

## What the composite score is A synthetic audit runs the page once under simulated throttling and reports several lab metrics, then combines them into one 0–100 number using fixed weights and a scoring curve per metric. The weights are not equal — main-thread blocking work is weighted heavily, paint timing moderately, and so on — and each metric's raw value is mapped through a curve rather than a straight line before being weighted. That design is deliberate and reasonable for its purpose: giving a non-specialist one number to look at. It is exactly the wrong shape for a merge gate. ## Why a blended number gates badly **Regressions cancel.** Because the score is a sum of weighted parts, two changes moving in opposite directions can net to zero. A pull request that adds 80 kB of JavaScript and also happens to swap a hero JPEG for a smaller AVIF can leave the score untouched while making the page materially worse on a low-end phone — the added script costs main-thread time that a cheap device pays for, and the image saving does not compensate for it. **The mapping is non-linear.** The curves are steep in their middle range and flat at the ends. A page already scoring in the high nineties can absorb a real regression with barely a point of movement, while a page sitting on a steep part of a curve can lose five points for a change nobody would call significant. The same millisecond cost produces a different score delta depending on where the page already sits, so a fixed score threshold enforces a different amount of strictness on every route. **It cannot be acted on.** The author of the failing pull request sees a number went down. They do not see which metric moved, by how much, or which resource caused it. In practice that means either a long manual investigation or — far more commonly — a re-run and a shrug. **It moves for reasons that are not the change.** Because the score aggregates measured values, it inherits all of their run-to-run variance, and compounds it: noise on any one of the constituent metrics can push the total across the threshold. ## What to assert instead Budget the specific quantities you actually care about, each with its own threshold and its own owner: - **Compressed bytes** per route, split by resource type — script, stylesheet, image, font. Deterministic, so it never flakes, and it names the file that grew. - **Largest contentful paint**, as the load-speed proxy that maps to what the user waits for. - **Total blocking time**, as the lab stand-in for responsiveness. A single lab page load cannot measure INP — the responsiveness Core Web Vital since March 2024 — because INP requires real user interactions, so total blocking time is what a synthetic run can offer in its place. - **Cumulative layout shift**, which is cheap to measure and catches a whole class of visual regressions that no byte budget sees. Each assertion answers a question, so each failure carries a diagnosis. "script bytes on /product went from 168 kB to 214 kB" is a review comment that fixes itself. ```json { "resourceSizes": [ { "resourceType": "script", "budget": 170 }, { "resourceType": "total", "budget": 450 } ], "timings": [ { "metric": "largest-contentful-paint", "budget": 2500 } ] } ``` ## Keeping the score around anyway None of this means the composite is worthless. It is a decent single number for a trend chart, for a management-facing report, and for spotting that something changed on a route nobody is watching closely. The rule is about *role*: composites are for **noticing**, specifics are for **gating**. A gate needs to be attributable to a change and actionable by its author, and a weighted blend is neither. ## The related trap The same argument applies to gating on any single aggregate — an average across routes, a total across all pages, one "performance grade" from a vendor. Aggregation is what lets a regression hide behind an unrelated improvement. Assert on the narrowest quantity that still means something to a user, and let the aggregate be a dashboard.

  • If the composite score is a poor gate, is it worth computing at all in CI?
    Yes, as a recorded trend rather than an assertion. It is a cheap single number for spotting drift on routes nobody watches closely, and for reporting outside the team. Store it per commit, chart it, and let the specific metric and byte assertions be the things that can fail the build.
  • Why does a lab audit report total blocking time rather than INP?
    INP measures the latency of real user interactions across a whole page visit, so it needs someone actually clicking and typing. A synthetic run loads the page once with no interaction, so it reports total blocking time — how long the main thread was blocked by long tasks during load — as the closest available proxy for how responsive the page would feel.
  • A pull request improves one metric and worsens another. How should the gate behave?
    Both assertions are evaluated independently, so the improved metric passes and the regressed one fails — which is the desired behaviour. Trading responsiveness for paint speed may still be the right call, but it should be an explicit decision made in review with the numbers on the table, not something a blended score silently absorbs.

saying these in an interview costs you the question

  • A single score is the simplest thing to gate on
  • Score deltas map linearly to real user impact
  • One high score means every route is fine
  • Aggregate metrics are fine as long as the threshold is strict
  • A synthetic run can measure INP directly

context