skip to content

Handling Nondeterminism

You will learn why the same page renders slightly differently twice, and how to remove every source of that before comparing images. Interviewers ask because unstable baselines are the reason most teams abandon visual testing after a month.

on this pageshow

questions

4

A dashboard screenshot test fails on every run: the page shows "updated 3 minutes ago", a randomly generated order ID, and a third-party map tile. How do you make each of those stable, and when would you mask a region instead?

level: middleimportance: must knowfreq 58%

answer

  1. identical input, identical bitmap
  2. three doors: time, randomness, data
  3. freeze the clock before first render
  4. pin TZ and locale too
  5. masking leaves a permanent blind spot

basics

~20 s

Make the data deterministic at its source: freeze the clock and pin the timezone and locale, seed or stub the ID generator, and serve fixed fixtures. Mask only what you truly cannot control, such as a live third-party tile, because a mask is a permanent blind spot.

solid answer

~60 s

Each of the three has a different owner, so each gets a different fix. The relative timestamp is a function of `Date.now()`, so install a fake clock before the app's first render and pin `TZ` and the locale for the run — otherwise the same fixture reads "3 minutes ago" in one run and "4 minutes ago" in the next, and renders a different date format on a runner with a different locale. The order ID comes from a generator, so stub it: a seeded PRNG in place of `Math.random`, or fixture data that carries the ID, so the same bytes render every time. The map tile is genuinely outside your control — different imagery, different render — so that is where masking earns its place: cover that rectangle with a solid colour in both the baseline and the comparison. The ordering principle is that determinism beats masking, because a masked region is a permanent hole in your coverage: a real regression inside it will never fail a test again.

code

javascript · 13 lines
javascript
// Installed in the page before the app boots.
function pinNondeterminism(fixedIsoTime) {
  const fixed = new Date(fixedIsoTime).getTime();
  Date.now = () => fixed;

  let seed = 42;
  Math.random = () => {
    seed = (seed * 1664525 + 1013904223) % 4294967296;
    return seed / 4294967296;
  };
}

pinNondeterminism('2026-03-07T10:00:00Z');

go deeper

for a junior

Know that a screenshot test needs the same input every run, and name the usual culprits: current time, random IDs, and live API data. Be able to say fixtures and a frozen clock come before any masking.

for a middle

Explain each fix precisely — a fake clock installed before first render, a pinned TZ and locale, a seeded or stubbed generator, checked-in fixture data — and articulate why masking is the last resort rather than the first.

for a senior

Demonstrate the coverage argument: quantify what a mask costs, insist masks trace to something outside the repo, and design the harness so uncontrolled calls fail loudly instead of drifting back into flakiness.

for a principal

Own it as a suite-wide contract: determinism is a property of the test harness, not of individual tests, so the clock, seed, fixtures and state reset live in shared setup, and mask usage is reviewed like any other coverage exemption.

## Three kinds of nondeterminism, three different fixes When a screenshot differs between two runs of unchanged code, the variation entered through one of a small number of doors. Naming the door tells you the fix. **Time.** Anything derived from `Date.now()` or `new Date()` is different on every run: relative timestamps ("3 minutes ago"), absolute dates, countdown timers, date pickers that highlight today, and charts whose x-axis ends at *now*. The fix is a fake clock installed *before* the app first reads the time. Test runners provide this at the module level (Sinon's `useFakeTimers`, Jest's `useFakeTimers().setSystemTime(...)`), and browser automation tools provide a page-level equivalent; whichever you use, the requirement is that it is in place before first render, because a component that captured `Date.now()` at import time cannot be retro-fixed. Two related settings travel with the clock: the **timezone** and the **locale**. A runner in UTC and a laptop in CET render the same instant as different wall-clock text, and `Intl.DateTimeFormat` with a different default locale produces `3/7/2026` versus `07/03/2026`. Pin both for the test run — the `TZ` environment variable for the process, and an explicit locale for the browser context — so the same fixture always produces the same string. **Generated values.** Random IDs, UUIDs, seeded avatars, shuffled orderings, `Math.random()`-driven placeholder content. Replace the source: stub `Math.random` with a seeded pseudo-random generator, or better, remove randomness from the test entirely by having the fixture supply the ID. The point is that the visual test should exercise *rendering*, not the ID generator — so pin the generator's output and let a unit test own its behaviour. ```js // installed before the app boots, in the page context let seed = 42; Math.random = () => { seed = (seed * 1664525 + 1013904223) % 4294967296; return seed / 4294967296; }; ``` **Data.** If the page renders whatever the API returns today, the screenshot is a picture of your staging database. Every visual test should render from fixed, checked-in data so that the only thing that can change the image is a code change. That is also why visual tests belong on stubbed data rather than a live backend — the volatility of production data would swamp the signal. ## Masking: the tool of last resort Masking paints a solid rectangle over a region in both the baseline and the candidate image, so whatever is under it cannot cause a diff. It is the right answer for content you genuinely cannot make deterministic: - third-party embeds you do not control (map tiles, ad slots, social widgets, payment iframes) - live video or camera feeds - content whose rendering is inherently environment-dependent and not worth pinning And it is the wrong answer for content you *could* have controlled. The cost is precise and permanent: a masked rectangle is a region where the test will never fail again. If the timestamp is masked because it was easier than installing a clock, then the day someone breaks its formatting, its colour, or its position, no test notices. Teams that mask liberally end up with a suite whose baselines are green and whose coverage is Swiss cheese. So the ordering is: control the input first; mask only what remains; and keep masks small and individually justified. A useful review rule is that every mask should be traceable to something outside your repository. ## Two supporting habits **Reset state between tests.** If test A leaves a row selected, a sort applied, or a value in localStorage, test B's screenshot depends on execution order — which is a nondeterminism that no clock or seed will fix. Each visual test should start from a known state. **Make it fail loudly if determinism is missing.** Prefer stubbing that *throws* on an uncontrolled call (a fetch to an unmocked URL, a real `Date` read after the clock was expected) over one that silently falls through. Otherwise the suite drifts back into nondeterminism one un-stubbed call at a time, and the symptom you get is an unexplained flake weeks later. ## Why interviewers ask This is the question that separates "has read about visual testing" from "has run a visual suite in CI." Almost every team that abandons visual regression does so because their baselines became untrustworthy, and the majority of that untrustworthiness is exactly this: dynamic data nobody pinned, papered over with masks and tolerances until the tests stopped meaning anything.

  • Why must the fake clock be installed before the app's first render rather than before the assertion?
    Because anything that read the real time earlier has already captured it — a module-level `const startedAt = Date.now()`, a store initialised at import, a component that computed a date on mount. Installing the clock afterwards fixes only future reads, so the page still carries one real timestamp, and it changes on every run.
  • The team masks the timestamp region because it is quicker than wiring a clock. What are you giving up?
    Every regression inside that rectangle, permanently. Wrong format, wrong colour, a truncated string, an element that moved — none of it can fail a test any more, and nothing in the diff report reminds anyone that the region is unwatched. Masks should be reserved for things outside your control, like a third-party embed.
  • Two visual tests pass individually but one fails when the whole suite runs. What kind of nondeterminism is that?
    Shared state leaking across tests — a persisted sort order, a selected row, a value left in localStorage or a cookie, or a fixture mutated in place. The screenshot then depends on execution order. The fix is structural: each visual test starts from a known, freshly reset state rather than inheriting whatever ran before it.

saying these in an interview costs you the question

  • Masking a timestamp instead of freezing the clock
  • Running visual tests against live backend data
  • Installing fake timers after the component has mounted
  • Ignoring timezone and locale differences on the runner
  • Treating a mask as free — it is unwatched forever

context

open as a page

In a visual regression test, a card plays a 300 ms entry animation before it settles. Why is inserting a fixed delay before the screenshot a poor fix, and what would you do instead?

level: middleimportance: must knowfreq 68%

basics

~20 s

A fixed delay is a guess: on a loaded CI machine the shutter still lands mid-animation, and on a fast one it only wastes time. Remove the motion instead — inject test-only CSS that disables animations and transitions and hides the text caret.

open as a page

Screenshot baselines captured on a developer's macOS laptop fail on the Linux CI runner with tiny anti-aliasing differences across every piece of text. Why does that happen, and how would you set the suite up so baselines are portable?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Text rasterization is operating-system specific — hinting, anti-aliasing and the installed fallback fonts all differ — so a macOS baseline is simply not valid on Linux. Produce and compare baselines inside one pinned container image, used identically in CI and on developer machines.

open as a page

A visual regression test screenshots a page, and on some CI runs the headings come out in a fallback font so the image comparison fails. Why does that happen, and what should the test wait for before it captures?

level: juniorimportance: should knowfreq 45%

basics

~20 s

Web fonts load asynchronously, so a screenshot taken too early captures fallback text with different glyph metrics and different line wrapping. Await document.fonts.ready in the page before capturing, and self-host the font files so the test never waits on a third-party CDN.

open as a page