skip to content

What is a performance budget in a web project's CI pipeline, and what has to be true of it before it actually prevents regressions?

level: juniorimportance: must knowfreq 50%

answer

  1. a number agreed before the work
  2. checked automatically on every change
  3. report versus blocked merge
  4. bytes of output, or a measured metric
  5. measure today, add headroom, ratchet down

basics

~20 s

A performance budget is a limit agreed in advance on a measurable property of the built site — bytes shipped, request count, or a metric such as LCP — that a CI job checks on every pull request and fails the build when exceeded.

solid answer

~50 s

A performance budget turns a vague goal like "the site should feel fast" into a specific number that a machine can check. Budgets come in two families: **quantity budgets** over the build output (compressed bytes of initial JavaScript, number of blocking requests, image weight per route) and **metric budgets** over a measured page load (LCP, total blocking time, layout shift). The budget only works if three things are true: it runs automatically on every pull request, it is measured against a production build of a fixed set of routes, and the job exits non-zero so the merge is blocked. A budget that only posts a comment or updates a dashboard is a report, not a budget — it gets skimmed and the regression lands anyway. Starting numbers usually come from measuring today's site and adding a little headroom, then tightening over time.

code

json · 6 lines
json
{
  "size-limit": [
    { "name": "main bundle", "path": "dist/main.js", "limit": "120 kB" },
    { "name": "initial CSS", "path": "dist/*.css", "limit": "20 kB" }
  ]
}

go deeper

for a junior

Be ready to define a budget as a pre-agreed limit checked automatically on every change, and to name two things you can put a number on: compressed bytes of JavaScript shipped, and a measured metric such as LCP.

for a middle

Expect to explain where the check sits in the pipeline, that it must run over a production build of named routes, and why the job has to exit non-zero rather than post a comment that people skim.

for a senior

Show that you can pick starting numbers from current measurements plus headroom, separate deterministic byte assertions from noisy measured ones, and make the failure message name the file that grew so the author can act without asking you.

for a principal

Own the policy side: who sets a budget, who may raise one and what they owe in exchange, how budgets ratchet downward as the site improves, and what you do when a valuable feature genuinely does not fit.

## What a budget actually is A performance budget is a number, agreed before the work starts, that some measurable property of the shipped site is not allowed to exceed. "Initial JavaScript on the product page stays under 170 kB compressed" is a budget. "We care about performance" is not, because nothing can check it. The point of putting it in CI is that performance decays by accident. Nobody ships a pull request called "make the site slower"; they ship a date picker, an analytics tag, a polyfill, an icon set. Each one is defensible on its own and each one costs a few kilobytes or a few milliseconds. Six months of individually-reasonable pull requests is how a fast site becomes a slow one. A budget is the mechanism that makes the cost visible at the moment someone can still choose differently — in review, not six months later in a performance sprint. ## Two families of budget **Quantity budgets** are asserted over the build artifacts. Compressed size of the entry chunk, total JavaScript on a route, number of render-blocking requests, total image bytes, size of the CSS in the head. These are deterministic: the same commit produces the same number on any machine, so the check never flakes and the failure is easy to attribute — a named file grew by a named amount. **Metric budgets** are asserted over a measured page load: largest contentful paint, total blocking time, cumulative layout shift, time to first byte. These are closer to what a user experiences, but they are measurements, so they carry run-to-run noise and need care (repeat runs, fixed throttling) before they can block a merge. Most teams run both, because they answer different questions. A byte budget catches the cause; a metric budget catches the effect, including effects that ship no new bytes at all — a hero image that stopped being preloaded, a font that started blocking, an animation that began thrashing layout. ## Where the check runs The check belongs in the pull-request pipeline, after a **production build**. Running it against a development build measures the wrong artifact entirely: unminified code, no tree shaking, dev-only warnings. It should measure a **fixed list of representative routes** — typically the entry page plus the one or two routes that carry the most traffic or the most code — because a budget on "the site" cannot be attributed to anything, and per-route budgets can be owned by the team that owns the route. A minimal quantity check needs very little machinery. A tool such as `size-limit` reads a config listing paths and limits, measures the compressed output, and exits non-zero when a path is over: ```json { "size-limit": [ { "name": "main bundle", "path": "dist/main.js", "limit": "120 kB" } ] } ``` The exit code is the whole point. In a hosted CI system, a job that exits non-zero is a failed check, and a failed required check blocks the merge. ## Choosing the first numbers The common mistake is picking an aspirational number, watching every pull request fail, and switching the check off within a fortnight. The workable sequence is: measure what the site does today, set the budget slightly above it so the current state passes, and then **ratchet** — whenever a change lowers the real number, lower the budget to match, so the improvement cannot silently be spent again. The opposite mistake is setting the budget at exactly today's value. Then ordinary noise, or a legitimate three-kilobyte change, fails the build, and the team learns that failures are meaningless. ## Warn versus fail Not every assertion deserves to block a merge. A useful split is to **fail** on assertions that are deterministic and clearly owned — byte counts, request counts — and to **warn** on assertions that are measured and noisy while you are still learning their spread. A warning is visible in review and can be argued about; a failure stops the line. Reserving failure for checks you trust completely is what keeps people from reflexively re-running or bypassing the job. ## What a budget cannot do A budget is a ceiling, not a target. Passing it means the change did not make things worse than the agreed limit; it does not mean the page is fast, and it does not tell you what real users on real devices are experiencing. It also cannot decide policy for you: someone still has to answer what happens when a genuinely valuable feature does not fit. Answering that in advance — who may raise a budget, and what they must do in exchange — is what separates a budget that survives a year from a check that quietly gets deleted.

  • Why measure the budget against a production build rather than the development server?
    A development build is a different artifact: unminified, unsplit, carrying dev-only warnings and no tree shaking. Its numbers are both much larger and unrelated to what users download, so a budget over it either passes trivially or fails permanently. The check has to run over exactly the artifact that would be deployed.
  • Where would you set the very first budget numbers for a site that has never had one?
    Measure the current production build, then set each budget a little above what it does today so the existing state passes. That makes the first green build honest. From there, ratchet: whenever a change brings the real number down, lower the budget to match, so the gain cannot be spent again by accident.
  • If you could only afford one budget assertion on a marketing site, which would you pick?
    Compressed bytes of JavaScript on the entry route. It is deterministic, so it never flakes; it is the thing that grows by accident most often; and it drives both download time and main-thread work, so it correlates with the metrics you actually care about while being far cheaper to check than measuring them.

saying these in an interview costs you the question

  • A budget is a dashboard target, not something enforced
  • One global budget covers the whole site
  • Setting the budget to today's exact number
  • Measuring the development build instead of the production output
  • Passing the budget means the page is fast

context