A visual regression check fails on a feature branch, the developer regenerates the baseline images, and the build turns green. When is that the right move, when is it a bug being committed, and what workflow separates the two?
answer
- a baseline is an expected value
- regenerating is editing the assertion
- who looked at the diff?
- new baseline ships with the code change
basics
~20 sRegenerating a baseline is correct only when a human has looked at the diff and confirmed the visual change was intended. Done blind to get a build green, it commits the regression as the new truth. The fix is an explicit approve-or-reject step, reviewed like code.
solid answer
~50 sA baseline is not test infrastructure — it is an assertion about what the product should look like, so overwriting one is editing an expected value. That is legitimate when the change is intended: you redesigned the button, you looked at the diff, the new image is what you want. It is a defect when the regeneration happens reflexively because the pipeline was red, because the same command that accepts a deliberate redesign also accepts a broken layout. The structural fix is to make acceptance explicit and reviewable: the failing run publishes baseline, actual and diff images; a person approves or rejects each change; approved baselines land in the same pull request as the code that changed them, so a reviewer sees the new image next to the diff that motivated it. Updating baselines should never be a step someone can perform without seeing what changed.
go deeper
Know that a baseline image is what the component is supposed to look like, and that you never regenerate one without opening the diff and confirming the change was intended.
Explain why the same command serves both an intended redesign and an accidental regression, and describe how shipping the new baseline inside the same pull request makes the difference reviewable.
Demonstrate the process judgment: per-screenshot approval, artefacts published on failure, no self-approval in the dark, and a real answer for the branch-conflict and mass-change cases.
Own the incentive problem — under deadline pressure the cheapest action is bulk acceptance — and design the workflow so the reviewed path is also the fast path, otherwise the suite decays into a rubber stamp.
## A baseline is an expected value In a unit test, nobody would accept "the assertion failed, so I edited the expected number until it passed" without asking what changed. A baseline screenshot is exactly that expected value, stored as an image. The reason blind regeneration is so common is ergonomic: every tool ships a one-command way to rewrite baselines, and that command cannot tell an intended redesign from a regression. Only a human looking at the diff can. ## The two cases **Legitimate update.** The change was intentional — a token changed, spacing was corrected, a component was redesigned. The correct sequence is: read the diff, confirm the new image is the desired appearance, and commit the new baseline *in the same change* as the code that caused it. Now the repository history answers "why does this component look like this?" with a commit. **Illegitimate update.** The pipeline was red and the fastest path to green was to accept everything. This is how a real defect becomes the new expected appearance, and it is worse than having no visual test at all: the suite now actively certifies the bug, and the next person to touch the component sees a green build. The tell between them is not the command used, it is whether a diff was reviewed. ## The approve-or-reject workflow A workable process has four properties: 1. **Every change is presented, not summarised.** The failing run must surface the baseline, the new actual image, and the highlighted diff. "3,412 pixels differ" is not reviewable. 2. **Acceptance is a deliberate act per screenshot.** Accepting all changes at once, in bulk, across a run is where regressions ride along with an intended redesign. If ten screenshots changed because you edited a shared component, you should still be able to see all ten. 3. **The approval is attached to the change.** The new baselines are part of the pull request, so the code reviewer sees the visual delta alongside the code delta and can say "that spacing looks wrong" before it merges. 4. **Nobody can accept their own change invisibly.** Whether that is enforced by code review on the committed images or by a hosted approval UI, the point is the same: a second pair of eyes, or at minimum an auditable record of who accepted what. ## Per-branch baselines and the merge problem When baselines live in the repository, they follow branches naturally: your branch carries its updated images, and merging brings both the code and its new expected appearance. The failure mode is two branches updating the same baseline — a binary file, so git cannot merge it. Whoever merges second must regenerate and re-review rather than picking a side, because "take mine" silently discards the other team's intended change. Hosted diffing services solve the same problem differently: baselines are keyed by branch, a branch inherits the baseline of the branch it was cut from, and an approval on a branch is promoted to the main baseline when it merges. Either model works; what does not work is a single mutable set of baselines with no relationship to the change that produced them. ## What to say about "the build is red and I need to ship" This is the real interview probe. The answer is that the pressure is legitimate and the shortcut is not: look at the diff — it takes seconds — and decide. If the diff is genuinely unreadable noise, the fix belongs in how the screenshot is captured or scoped, not in accepting an image nobody understood. A team that accepts baselines under time pressure once will do it every time, and within a quarter the suite is a rubber stamp.
- Two branches update the same baseline image and both merge. What happens, and how should it be handled?Images are binary, so version control cannot merge them — one side wins and the other's intended change disappears silently. The second merge should regenerate the baseline from the merged code and have someone review that combined image, rather than resolving the conflict by picking a file. It is a signal that the screenshot's scope may be too broad if two teams touch it routinely.
- Why is committing baselines in the same pull request as the code better than updating them on the main branch afterwards?It keeps the expected appearance and the change that caused it atomic and reviewable together: the reviewer sees the diff next to the code, and history explains why the component looks the way it does. Updating on main afterwards means the main branch is briefly certifying an appearance no one approved, and the link between cause and new baseline is lost.
- A shared component changes and forty screenshots go red at once. How do you review that without rubber-stamping?Recognise that forty diffs of the same root cause are one decision, not forty: review a few representative screenshots carefully, confirm the change matches the intent, then accept the set as one reviewed batch with the reason recorded. The thing to avoid is bulk-accepting a mixed run where an unrelated regression is hiding among the expected diffs.
saying these in an interview costs you the question
- Just regenerate the baselines when the build is red
- Baselines are test infrastructure, not product expectations
- Update baselines on main after merging, not in the PR
- Binary baseline conflicts can be resolved by taking one side
- Anyone can accept their own visual changes unseen