skip to content

A visual regression check fails in a CI pipeline with a message saying the screenshot comparison did not match. Which images does the job need to publish for someone to triage that failure, and why is the highlighted diff image alone not enough?

level: juniorimportance: should knowfreq 42%

answer

  1. three images, not one
  2. the diff is a mask, not content
  3. CI produced the failing pixels
  4. upload artefacts even when the step failed

basics

~20 s

Publish three images per failing comparison: the committed baseline, the actual screenshot from this run, and the highlighted diff, plus the differing-pixel count. The diff shows where pixels changed but not what the component looked like before or what it looks like now.

solid answer

~50 s

The diff image is a mask — it tells you which coordinates differ, usually painted in a loud colour, but it does not show either of the two pictures it was computed from. To decide whether a failure is an intended redesign, a regression, or capture noise, a reviewer needs all three artefacts side by side: the baseline (what we said it should look like), the actual image (what this run produced), and the diff (where they disagree). Publish them as build artefacts keyed to the test name, along with the pixel count and the thresholds in force, so nobody has to reproduce the run locally to see what happened. This matters more in CI than locally because the images that fail are the ones the CI environment produced; a developer re-running on their own machine may not reproduce them at all, and the reviewer often is not the person who wrote the test.

go deeper

for a junior

Be able to say that a failing visual check should hand you the baseline, the actual screenshot and the diff, and that the diff by itself only marks where pixels changed.

for a middle

Explain why the artefacts must come from the run that failed rather than a local re-run, and mention that the upload step has to execute even though the test step failed.

for a senior

Show that you design triage for the person who did not write the test — named files, images grouped per comparison, a link surfaced on the pull request, and a sane retention policy for large binaries.

for a principal

Treat evidence delivery as part of whether the suite works at all: if reviewers cannot see what changed in seconds, they will approve blind, and the checking value of the whole investment goes to zero.

## What each artefact tells you A failing comparison produces three images, and each answers a different question: - **Baseline** — the stored expectation. Answers "what did we agree this should look like?" - **Actual** — the screenshot captured in this run. Answers "what did the code produce this time?" - **Diff** — a derived image, typically the baseline faded out with the differing pixels painted in a high-contrast colour. Answers "where exactly do they disagree?" The diff alone is a heat map of disagreement with no content. A block of red pixels where a button sits could mean the button moved, changed colour, gained a border, lost its label, or is simply anti-aliased differently. You cannot tell which without the two source images, and the difference between those cases is the difference between accepting the change and blocking the merge. Alongside the images, publish the numbers: how many pixels differed, and what the thresholds were. That turns "it failed" into "14,000 pixels differed against a budget of 100", which immediately distinguishes a whole-component change from edge noise. ## Why CI is the place this matters The screenshots that failed were produced by the CI environment, not by a developer's laptop. Re-running the suite locally to "see the failure" often produces a different image, or no failure at all, because the rendering environment is not the same one. That is why baselines are normally generated by the same environment that runs the checks, and why the run's own artefacts are the authoritative evidence. There is also a review dimension: the person deciding whether the change is acceptable is frequently a reviewer or designer, not the author, and may not run the suite at all. If the evidence only exists inside a container that was destroyed when the job finished, the practical outcome is that people accept changes they never looked at. ## Making the artefacts usable A few things separate artefacts people actually use from artefacts nobody opens: - **Name files by the test and screenshot**, not by a run-scoped counter, so ten failures are ten identifiable components rather than `diff-1.png` through `diff-10.png`. - **Keep the three images together** in one directory per failing comparison, so triage does not mean hunting through a flat archive. - **Surface a link in the build output or the pull request**, because an artefact that requires knowing where to look is one nobody looks at. - **Set a retention period.** Screenshots are large and produced in bulk; keeping every run's images forever is a storage bill for data that is worthless a week later. - **Do not gate on a report that only exists locally.** If the reviewer needs the images, the pipeline must publish them, not print a path inside its own container. ## Common shape in practice Most frontend runners already write these three images into an output directory on failure; the CI job's responsibility is to upload that directory as an artefact even when the step failed — which usually means the upload step must be configured to run regardless of the previous step's exit status. Forgetting that is the single most common reason a team has visual tests that fail with no evidence attached.

  • Why is re-running the failing visual test locally a weak way to investigate it?
    The failing image was produced by the CI environment, and a developer's machine is not that environment, so the local run may render differently or not fail at all. The authoritative evidence is the artefact from the run that failed. Local re-runs are useful for iterating on a fix, not for judging whether the recorded difference is real.
  • What should the build publish besides the three images?
    The differing-pixel count and the thresholds that were in force, mapped to the test and screenshot name. Numbers immediately separate a whole-component change from a few hundred edge pixels, and recording the active threshold means a reviewer can see whether the assertion was strict or already generously loosened.

saying these in an interview costs you the question

  • The diff image alone is enough to triage a failure
  • Just re-run the test locally to see what broke
  • Upload artefacts only when the job succeeds
  • Baselines can be regenerated on a laptop for a CI failure

context