A visual regression suite covers 40 pages across 3 browser engines, 4 viewport widths, 2 colour schemes and 2 device-pixel-ratio settings. How many baseline images is that, and what does each extra axis cost beyond image storage?
answer
- axes multiply, they do not add
- storage is the cheapest cost
- one intentional change, every baseline red
- approve-all is the failure mode
- yield per axis is not equal
basics
~20 sForty-eight configurations per page, so 1,920 baselines — the axes multiply rather than add. The real cost is not disk: it is CI time, a multiplied flake surface, and above all the number of images a human must review and approve on every intentional design change.
solid answer
~40 sThe axes multiply: 3 x 4 x 2 x 2 is 48 configurations, times 40 pages is 1,920 baseline images. Storage is the cheapest thing on that list. Each axis multiplies CI wall-clock, because every configuration is a real render and capture. It multiplies the flake surface, because one unstable element now fails in 48 places. Most importantly it multiplies human review: change the shared header on purpose and you have hundreds of diffs to approve, which is how teams end up rubber-stamping and letting a genuine regression through in the pile. So I size the matrix by defect yield per axis — dark mode usually earns its place because theme tokens are hand-maintained, whereas a second device-pixel ratio rarely finds anything a layout test does not.
code
javascript · 6 linesconst axes = { engines: 3, widths: 4, schemes: 2, dpr: 2 };
const perPage = Object.values(axes).reduce((total, n) => total * n, 1);
const pages = 40;
console.log(perPage, perPage * pages); // 48 1920
console.log(Object.values(axes).reduce((total, n) => total + n, 0) * pages); // 440 if only one axis varies at a timego deeper
Know that the configurations multiply, and be able to compute the baseline count from the axes and the page count without being walked through it.
Explain the four costs an axis multiplies — capture time, flake surface, review queue and baseline churn — and say which axes historically earn their place.
Demonstrate that you would tier the suite or vary one axis at a time rather than run a full Cartesian product, and be explicit that human review throughput is the constraint you are sizing against.
Own the policy: define what the matrix may contain, require evidence of a caught defect before an axis is added, and schedule pruning so the matrix shrinks as well as grows.
## The arithmetic, and why it is the whole point Axes of a visual matrix compose multiplicatively. With three engines, four widths, two colour schemes and two device-pixel-ratio settings, one page has 3 x 4 x 2 x 2 = 48 renderings, and 40 pages produce 1,920 baseline images. Adding a fifth width does not add 40 images, it adds 480. This is the single most useful thing to say in the interview, because a candidate who reasons additively will always propose a matrix nobody can operate. ```javascript const axes = { engines: 3, widths: 4, schemes: 2, dpr: 2 }; const perPage = Object.values(axes).reduce((a, b) => a * b, 1); // 48 const total = perPage * 40; // 1920 ``` ## The four real costs **CI wall-clock and machine-minutes.** Every configuration is an actual page load, render and capture — there is no way to derive the Firefox image from the Chromium one. At two seconds per capture, 1,920 images is over an hour of pure capture time before any setup, and it lands on every pull request if the suite is a required check. **Flake surface.** Any element that is not perfectly deterministic now fails in 48 places instead of one. A single unstable region turns one investigation into a wall of red, and the noise-to-signal ratio of the whole suite drops. **Review throughput — the binding constraint.** Visual testing differs from unit testing in that a failure is not automatically a bug: someone must look at the diff and decide whether it was intended. When a designer legitimately changes the header, every baseline containing the header goes red at once. If the header appears on all 40 pages, that is 1,920 images in the approval queue for one intentional change. Humans do not review 1,920 images; they click approve-all. That is the failure mode that makes a large matrix actively worse than a small one — it converts a detector into a rubber stamp, and a real regression rides through in the batch. **Baseline churn and ownership.** Baselines are versioned artifacts. Multiply them and you multiply merge conflicts on baseline updates, the size of the repository or object store, the time to regenerate after an infrastructure change, and the number of images that quietly go stale because nobody owns that configuration. ## Yield is not equal across axes Because the axes cost the same multiplicatively, the only sane way to choose is by how many real defects each one has historically caught. - **Viewport width** is usually the highest-yield axis: layout branches genuinely differ, and the bugs are structural. - **Colour scheme** is often high-yield in practice, because dark-mode values are hand-maintained per token; a hardcoded colour or a missing dark value produces an unreadable component that no width or engine axis would reveal. - **Browser engine** yields real but rarer defects, and only for genuinely distinct engines — two brands sharing an engine produce near-duplicate images. - **Device pixel ratio** is usually the weakest. It mostly affects which asset a responsive image source set picks and how sub-pixel rendering resolves; it rarely changes layout, and it is a common source of diff noise rather than diff signal. ## Shaping the matrix instead of taking its product The matrix does not have to be a full Cartesian product, and mature suites are not. Two techniques do most of the work. First, **tiering**: run the full matrix over a small set of representative screens or components chosen because between them they exercise every layout branch, and run a single reference configuration over everything else. Second, **axis pinning**: pick one configuration as canonical and vary only one axis away from it at a time, so cost grows with the sum of axis sizes rather than their product — 3 + 4 + 2 + 2 configurations instead of 48. You lose interaction coverage (dark mode at the narrow width in WebKit specifically), which is a real but usually small class of defects. ## What a strong answer sounds like Do the multiplication out loud, name human review as the binding constraint rather than storage, and then say what you would actually cut and why. "I would keep width and colour scheme at full breadth on a tiered set of pages, keep one baseline per distinct engine, and drop the device-pixel-ratio axis unless we have shipped a bug it would have caught" is the answer an interviewer is listening for.
- Which single axis would you drop first from that matrix, and how would you justify it?The device-pixel-ratio axis. It mostly changes which image source set entry is chosen and how sub-pixel edges resolve, so it produces diff noise far more often than a genuine layout defect, and the bugs it does find are usually asset-selection bugs a cheaper check can catch. I would justify it with history: how many real regressions did the 2x baselines catch last quarter?
- How does tiering the suite change the arithmetic without giving up the coverage you care about?Pick a small set of screens that between them exercise every layout branch and theme, and run the full matrix only there; everything else gets one reference configuration. With five representative pages the full-matrix cost falls from 1,920 images to 240 plus 35, and the branches you were actually trying to cover are still covered — what you lose is redundant confirmation that the same header renders the same way on 40 pages.
- Why is a large visual matrix sometimes worse than a small one, rather than just more expensive?Because the output is reviewed by people. Past a few dozen diffs per change, reviewers stop looking and approve in bulk, so a genuine regression is approved along with the intended ones. A small matrix that is actually read catches more bugs than a large one that is rubber-stamped, which makes review throughput — not CI capacity — the number the matrix must be sized against.
saying these in an interview costs you the question
- Adds the axes instead of multiplying them
- Says storage is cheap, so the matrix cost is negligible
- Assumes more configurations always means more real defects found
- Ignores that every red diff needs a human decision
- Treats device-pixel ratio as equal in value to viewport width