In a visual regression suite, why does requiring two screenshots of the same page to match exactly fail almost immediately, and what does a perceptual per-pixel comparison do instead?
answer
- identical-looking, not identical bytes
- anti-aliased edges vary per machine
- colour distance, not equality
- decode pixels, never hash the file
basics
~20 sByte-equal screenshots almost never repeat: GPU rasterisation, text shaping and anti-aliasing shift pixel values without changing what a human sees. Visual tools instead decode both images, score each pixel pair with a perceptual colour distance, and count only differences above a tolerance.
solid answer
~50 sExact equality treats a screenshot as an opaque blob, and two renders of an unchanged page are rarely identical at that level: sub-pixel anti-aliasing, GPU rasterisation and text shaping move individual channel values by a few units. So a visual tool decodes both images and walks them pixel by pixel, scoring each pair with a colour-distance metric rather than an equality test. `pixelmatch`, the comparator Playwright uses for PNG screenshots, measures that distance in the YIQ colour space, which weights luminance closer to the way human vision does, and it separately detects anti-aliased pixels — a pixel sitting between contrasting neighbours — which it does not count by default. The output is a count of genuinely differing pixels plus a diff image highlighting them. That count is what your threshold policy applies to; exact equality gives you no dial at all, which is why it yields a permanently red build.
code
javascript · 20 linesimport { PNG } from 'pngjs';
import pixelmatch from 'pixelmatch';
import fs from 'node:fs';
const baseline = PNG.sync.read(fs.readFileSync('baseline.png'));
const actual = PNG.sync.read(fs.readFileSync('actual.png'));
const { width, height } = baseline;
const diff = new PNG({ width, height });
const changedPixels = pixelmatch(
baseline.data,
actual.data,
diff.data,
width,
height,
{ threshold: 0.2 }
);
fs.writeFileSync('diff.png', PNG.sync.write(diff));
console.log(`${changedPixels} pixels differ`);go deeper
Know that a visual test compares a fresh screenshot with a saved baseline image, and that the tools compare decoded pixels with a tolerance rather than checking that the two files are identical.
Be ready to explain why identical-looking renders produce different pixels — anti-aliasing, GPU rasterisation, text shaping — and what a colour-distance comparison plus anti-alias detection gives you that an equality check cannot.
Show that you treat the differing-pixel count as the input to a policy you own, and that you can debug a suite failing on noise without simply widening tolerance until it goes quiet.
Frame the tradeoff: a comparator that is too strict costs engineering hours in triage, one that is too loose silently stops catching regressions, and neither failure is visible from the build's colour alone.
## The three steps of a visual test A visual regression test does three things: **capture** an image of a page or component, **compare** it against a stored baseline image, and **decide** whether the difference counts. The comparison step is where the engineering lives, and the naive version — "are these two files identical?" — collapses on the first run on a different machine. ## Why exact equality fails A screenshot is the result of a long rasterisation pipeline. Text is shaped and hinted, edges and curves are anti-aliased, and parts of the page may be composited on the GPU. None of those stages promise bit-identical output between runs, machines or driver versions; they promise something that *looks* the same. A single anti-aliased glyph edge can differ by a few units per channel and still be visually indistinguishable. On top of that, comparing the encoded file rather than the decoded pixels adds noise that has nothing to do with the render: PNG is lossless, but encoder version, filter choice and metadata chunks can all change the bytes of a file whose decoded pixels are identical. So the first rule is: **decode, then compare pixels — never hash the file.** ## What a perceptual comparison does A per-pixel comparator lines the two decoded bitmaps up (they must have the same dimensions) and, for each coordinate, computes how *different* the two colours are rather than whether they are equal. `pixelmatch`, the library behind Playwright's PNG screenshot comparison, converts pixels to YIQ and takes a weighted distance in which the luminance term dominates — a shift the eye barely registers scores low, while a change of hue or brightness that a person would notice scores high. A per-pixel tolerance decides how large that distance must be before the pixel is called different. ```js import pixelmatch from 'pixelmatch'; // returns the number of pixels judged different const changed = pixelmatch( baseline.data, actual.data, diff.data, width, height, { threshold: 0.2 } ); ``` The comparator also does something structural: it classifies **anti-aliased** pixels. A pixel is treated as anti-aliasing when it is a blend between contrasting neighbours, and by default those pixels are excluded from the count. This matters because edge noise is exactly the noise you want to discard, and discarding it with a *global* colour tolerance would mean also tolerating real colour changes in flat areas. The result is two artefacts: a number (how many pixels differ) and a diff image (where they differ, usually painted in a high-contrast colour over a faded copy of the baseline). ## Why this matters for the build Because the comparison produces a *number*, you can hold a policy over it: allow N differing pixels, or a ratio of the frame. Exact equality gives you a boolean with no dial, so any harmless rendering jitter fails the build; teams that start there quickly stop trusting the suite. Note the flip side: the dial is where regressions hide, so a perceptual comparator is not a licence to loosen tolerance until the suite is quiet. ## Two practical constraints - **Dimensions must match.** If the actual image is one pixel taller than the baseline, most comparators refuse to diff at all, or every row below the shift misaligns and the diff reads as near-total. Failing loudly on a size mismatch is better than a meaningless 90% diff. - **Never store baselines in a lossy format.** A JPEG baseline bakes compression artefacts into the reference, and those artefacts differ per encode; PNG (or another lossless format) is the only sane choice. ## What it is not A perceptual comparator has no idea what a button or a heading is. It does not match shapes, detect that an element moved, or understand layout. It compares two grids of colours at the same coordinates — which is why a one-pixel vertical shift of a long page can produce a huge diff even though nothing "changed".
- Why do comparators detect anti-aliased pixels separately instead of just raising the global colour tolerance?Anti-aliasing noise sits on edges and can be a large colour difference while the shape is unchanged. A global tolerance high enough to absorb it would also absorb real colour regressions in flat regions — a wrong brand colour, a wrong disabled state. Anti-alias detection is local and structural, so it discards edge noise while keeping strict comparison everywhere else.
- The actual screenshot comes back one pixel taller than the baseline. What does a per-pixel comparator do?Nothing useful. Most tools fail immediately with a size mismatch, and that is the right behaviour: if you pad or crop instead, every row below the shift misaligns and the diff reports most of the image as changed, which tells the reviewer nothing. Treat a dimension change as its own failure mode and investigate what resized the frame.
saying these in an interview costs you the question
- Two screenshots of the same page are always byte-identical
- Any non-zero pixel difference means a real visual bug
- Store baselines as JPEG to keep the repository small
- Perceptual diffing understands layout or elements moving
- Loosening the colour tolerance costs nothing