A React app's component suite snapshots whole pages. Nearly every pull request turns several snapshots red for reasons unrelated to the change, and the team's routine fix is to re-run the tests with the update flag and commit the regenerated files. What has that suite stopped protecting, and how would you restructure it?
answer
- the loop closes on itself
- regenerating equals no assertion
- breadth causes unreadable diffs
- narrow the capture, state the intent
- stop the inflow, convert on contact
basics
~20 sIt protects nothing: a baseline regenerated without being read is just a copy of current output, so the tests can only fail, never catch anything. Fix the cause — captures too broad to review — by narrowing them to small subtrees or derived data, adding assertions that state intent, and pruning obsolete baselines.
solid answer
~50 sOnce regenerating is the standard response to red, the suite has degraded into a ritual: it costs CI time and review attention and returns no information, because a baseline nobody read is by definition equal to whatever the code currently does — bugs included. I would treat the blind-update habit as a symptom rather than a discipline problem. The cause is that whole-page captures fail on any markup churn and produce diffs too long to judge, so the only economical response is to regenerate. The restructuring is to shrink the blast radius: replace page-level snapshots with explicit assertions naming what each test proves, keep at most a small narrow snapshot of a compact derived value, prune the baselines that are now obsolete, and make the build refuse to write missing snapshots so nothing is baselined silently. Then a red snapshot becomes rare enough that reading it is affordable.
go deeper
Know that updating a failing snapshot without reading the diff makes the expectation equal to the current output, so the test can no longer catch anything.
Explain why whole-page captures fail on unrelated changes, and describe the fix: assert what the test actually claims, and keep any snapshot small enough that its diff can be judged.
Show the diagnosis-to-restructure path: identify capture breadth as the cause of the regeneration habit, preserve the deletion coverage you would lose, prune obsolete baselines deliberately, and make the build refuse to write missing snapshots.
Be ready to fund and sequence the cleanup against feature work: stop the inflow with an enforced standard, migrate on contact rather than in a freeze, and say how you would measure a year later that regressions are being caught rather than merely that churn went down.
## Name the failure first The suite has stopped being evidence. A test's value comes from the moment a human decides whether a failure was intended. Remove that decision and the loop closes on itself: output changes, test goes red, output is recorded as the new expectation, test goes green. The only thing this can ever detect is that something changed — and the team has already agreed in advance not to care. Worse, it is actively harmful in two ways. It burns review attention on diffs nobody reads, which trains reviewers to skim, and it can bury a real regression inside the same batch of "expected" churn. ## Diagnose the cause, not the discipline The tempting answer is "tell the team to stop doing that". It will not hold, because the behaviour is rational under the current design. If a snapshot covers an entire page, then a class rename, a wrapper element added for layout, a copy tweak in a shared footer, or an unrelated child component's refactor all turn it red. Reading a four-hundred-line diff to conclude "nothing that matters changed" is expensive and yields nothing, so people stop. So the question to ask is: *why does a change to X break a test about Y?* The answer is capture breadth. Snapshot scope is the design defect. ## The restructuring **1. Make each test state its claim.** For every page-level snapshot test, write the assertions the test name implies: ```javascript test('shows an empty state when there are no invoices', () => { render(<InvoicesPage invoices={[]} />); expect(screen.getByRole('heading', { name: 'No invoices yet' })).toBeInTheDocument(); expect(screen.queryByRole('table')).not.toBeInTheDocument(); }); ``` Those fail only when the behaviour they name breaks, and their failure message is a bug report. **2. Keep a narrow snapshot only where breadth earns it.** Deletion coverage is the honest argument for snapshots — a targeted test never notices a section that silently vanished. Get it from a compact derived value instead of markup: the list of visible section headings, the table's cell text as an array, the normalized props handed to a chart. Small, readable in a diff, immune to styling churn. **3. Prune the wreckage.** Renamed and deleted tests leave orphaned entries behind in snapshot files; runners report these as obsolete and remove them on a deliberate update run. Do that pass once, deliberately, in its own commit so the deletion is reviewable and not tangled with real changes. **4. Close the silent-write hole.** Configure the build so a missing snapshot fails rather than being recorded — Jest's `--ci` flag does this, and Vitest declines to create new snapshots when it detects a CI environment. Otherwise a snapshot file that was never committed regenerates on the build machine and everything passes. **5. Change what review asks for.** A changed baseline is a claimed behaviour change and belongs in the PR description alongside the code change. If the author cannot say why the output moved, that is the finding. ## Sequencing it in a real codebase You cannot stop feature work to rewrite hundreds of tests. Stop the inflow first: agree that new tests do not snapshot pages, and enforce it in review. Then convert on contact — whenever a snapshot goes red, the person holding it either writes the explicit assertions for that test or deletes it, rather than regenerating. High-churn areas convert themselves within a few sprints, and the stable ones that never fail are, by definition, not costing anything. ## What a strong answer sounds like A weak answer says "snapshots are bad, delete them". A strong one distinguishes the mechanism (breadth causes unreadable diffs), the human consequence (regeneration becomes rational, so the tests stop carrying information), and the structural fix (narrow the capture, state intent explicitly, fail closed in CI), and it has a migration plan that does not require a freeze.
- If you delete the page snapshots, how do you keep coverage for a section that silently disappears?Assert its presence directly where a user depends on it, and keep one compact derived snapshot — for example an array of the visible section headings or the table's cell text. That preserves the deletion signal the broad snapshot gave you, in a form small enough that a reviewer can judge the diff in seconds and stable enough that styling changes do not move it.
- How would you tell whether these snapshots ever caught a real regression?Look at the history of the snapshot files. For each changed baseline, ask whether the PR intended to change output. If nearly every change accompanied an unrelated refactor and none was reverted after a review found the diff wrong, the snapshots have been detecting churn only. That evidence is far more persuasive to a team than an argument about testing philosophy.
- A teammate says the real fix is a review rule that snapshot diffs must be read line by line. Why is that not enough on its own?Because it prices the fix at hours of reviewer time per PR and depends on vigilance that will decay under deadline pressure. A rule that requires people to do an unaffordable thing is a rule that will be broken quietly. Shrink the diff until reading it is cheap, and the rule enforces itself.
- When would you leave a large snapshot in place rather than converting it?When it is stable and generated — a serialized design-token output, a formatted document, a compiled stylesheet — so it rarely fails, and when any change to it genuinely deserves a look. The test is churn: if a baseline has not changed in a year, it is not costing review attention and its breadth is free.
saying these in an interview costs you the question
- Treats blind regeneration as a discipline problem rather than a design one
- Proposes deleting every snapshot without preserving deletion coverage
- Says a stricter review rule alone will fix routine snapshot churn
- Believes a regenerated baseline still verifies the output
- Plans a big-bang rewrite of the whole suite before feature work resumes