skip to content

A screenshot-diffing tool exposes both a per-pixel colour tolerance and a limit on how many pixels may differ (for example Playwright's `threshold` alongside `maxDiffPixels` / `maxDiffPixelRatio`). What does each knob control, and how can a generous pixel budget hide a real regression?

level: seniorimportance: should knowfreq 48%

answer

  1. one knob per pixel, one per image
  2. noise scatters, regressions cluster
  3. do the arithmetic on the ratio
  4. per-screenshot, not global, slack

basics

~20 s

The colour tolerance decides whether one pixel counts as different at all; the pixel count or ratio decides how many such pixels the whole image may contain before the test fails. A budget large enough to absorb rendering noise is also large enough to hide a small, genuinely broken region.

solid answer

~50 s

They act at different stages of the same pipeline. The per-pixel tolerance is applied inside the comparison: for each coordinate, how far apart must the two colours be before this pixel is called different. Playwright names that `threshold`, and in recent 1.4x releases it defaults to 0.2 on a 0–1 scale. The pixel budget is applied after the count: `maxDiffPixels` is an absolute number of differing pixels the image may contain, `maxDiffPixelRatio` the same as a fraction of the frame. The danger is that both are blunt and global. A one percent ratio on a 1280×800 screenshot allows over ten thousand differing pixels — comfortably enough to hide a misaligned icon, a wrong badge colour or a truncated label, because a real regression is usually *localised* while noise is *scattered*. So budgets should be small, set per screenshot rather than globally, and raised only with a named reason.

go deeper

for a junior

Know that a visual comparison has a tolerance, and that it exists because identical-looking renders differ slightly; do not silently raise it to make a failing test pass.

for a middle

Explain the two stages clearly: a colour distance decides whether a pixel counts as different, then a count or ratio decides whether the image as a whole fails, and say what each one lets through.

for a senior

Show the arithmetic and the asymmetry — scattered noise versus a localised defect — and describe a policy: strict defaults, per-screenshot exceptions, region exclusion instead of blanket loosening.

for a principal

Own the long-run risk that thresholds ratchet outward one incident at a time until the suite cannot fail, and put a check in place — periodic deliberate breakage, review of every budget increase — that keeps it honest.

## Two knobs, two stages Every screenshot comparator has the same two-stage shape, whatever it calls the settings. **Stage one — is this pixel different?** For each coordinate the comparator computes a colour distance and compares it to a per-pixel tolerance. Below the tolerance the pixel is identical for the purposes of the test; above it, the pixel is counted and painted into the diff image. Playwright surfaces this as the `threshold` option on its screenshot assertion (0 = strict, 1 = anything passes); BackstopJS-style tools expose an equivalent via their resemble-based comparator. **Stage two — are there too many different pixels?** The count from stage one is compared to a budget. Playwright offers two forms: `maxDiffPixels` (an absolute pixel count) and `maxDiffPixelRatio` (a fraction of total pixels). Exceed it and the assertion fails. The knobs are not interchangeable. Raising the colour tolerance makes every pixel more forgiving everywhere — including in the flat interior of a button whose brand colour is now wrong. Raising the pixel budget keeps per-pixel strictness but allows more of the frame to change — including one small, entirely broken region. ## Why a generous budget is dangerous The key asymmetry: **noise is scattered, regressions are localised.** Anti-aliasing jitter sprays a handful of pixels along every edge in the image. A real defect — a control shifted eight pixels, an icon that failed to load, a badge rendered in the wrong colour, a label truncated with an ellipsis — occupies a compact block. Do the arithmetic before you accept a ratio. A 1280×800 screenshot is 1,024,000 pixels; `maxDiffPixelRatio: 0.01` permits 10,240 differing pixels. A 100×100 icon that disappears entirely is 10,000 pixels — it passes. That is the whole trap: the number looks conservative as a percentage and is enormous as an area. Ratios also scale wrongly with page size. The same ratio on a tall full-page screenshot allows several times more absolute damage than on a component screenshot, so "one percent everywhere" silently means "almost nothing is checked" on your largest pages. ## Setting a policy that actually holds - **Default to strict and specific.** Start with a near-zero budget and tighten the colour tolerance as far as your environment allows. A suite whose baselines are produced in the same environment that runs the tests can often sit at zero or a handful of pixels. - **Set budgets per screenshot, not globally.** A global loosening is applied to your most important assertions along with your noisiest one. If one screenshot genuinely needs slack, give the slack to that screenshot. - **Prefer absolute counts over ratios for large frames**, or split a long page into several smaller screenshots so a budget means the same thing everywhere. - **Prefer excluding a region to loosening the whole frame.** If a specific area is legitimately variable, keep the rest strict and take that area out of the comparison. - **Treat a budget increase as a code change with a reason.** "Bumped to 2% to get the build green" is the commit message that ends a suite's usefulness; require the diff artefact and a sentence about what produced the noise. ## The failure mode to name in an interview A suite that is never red is not necessarily a suite that is passing. The most common end state of a visual regression project is not flakiness — it is a set of thresholds that were widened one emergency at a time until nothing can fail. The diagnosis is easy and worth mentioning: deliberately break a component, run the suite, and see whether it goes red. If a known regression passes, your budget is fiction.

  • How would you verify that your current thresholds still catch anything?
    Mutation-test the suite: introduce a deliberate visual defect — nudge a control by a few pixels, change a token colour, hide an icon — and confirm the relevant screenshot fails. Do it for a small localised change, not a full-page one, since that is exactly the class a loose budget swallows. If the suite stays green, the threshold is decorative and should be tightened or the assertion narrowed.
  • When would you exclude a region from the comparison instead of raising the budget?
    Whenever the variability is confined to a known area — a live counter, an avatar, an embedded map. Excluding that region keeps every other pixel under a strict policy, whereas raising the budget spends the same tolerance across the whole frame and would let an unrelated defect of the same size pass unnoticed.
  • Why can the same ratio be reasonable on a component screenshot and reckless on a full-page one?
    A ratio is a fraction of total pixels, so its absolute allowance grows with the frame. One percent of a small component image is a few hundred pixels; one percent of a tall full-page capture can be tens of thousands — a large visible area. Either use absolute counts, or capture smaller, targeted screenshots so the budget means a consistent amount of damage.

saying these in an interview costs you the question

  • Set one global threshold and forget about it
  • One percent of pixels is obviously a small change
  • Widen the threshold whenever the build is red
  • Colour tolerance and pixel budget do the same thing
  • A green visual suite proves nothing regressed

context