You own a product's visual regression suite and its browser and viewport matrix keeps growing. How would you decide what the matrix should contain, and how would you stop it expanding without bound?
answer
- size it against review throughput
- derive, do not accumulate
- full matrix on a representative tier
- adding needs a named defect
- zero yield for two quarters, remove
basics
~20 sTreat the matrix as a fixed budget sized against human review throughput, not CI capacity. Derive it from the published browser support policy, tier it so the full matrix runs over a few representative screens, and require measured defect yield to keep an axis — pruning is a scheduled activity, not an exception.
solid answer
~50 sI would make the matrix an explicit, owned artifact with a budget. The budget is set by review throughput: past a few dozen diffs per change, people approve in bulk and the suite stops detecting anything, so that number is the real constraint, not runner capacity. Within the budget I derive the configurations from a published browser support policy and the design system's declared breakpoints, then tier the suite — the full matrix over a small set of screens chosen because between them they exercise every layout branch and theme, one reference configuration for everything else. To stop growth I make entry and exit symmetric: adding an axis requires naming a real defect it would have caught, and every configuration is reviewed quarterly against how many genuine regressions it actually found. Anything with zero yield over two quarters goes, and I would push some coverage off the matrix entirely — a token lint finds missing dark-mode values far more cheaply than doubling the baselines.
go deeper
Know that a visual matrix has a cost that grows multiplicatively, and that someone has to decide what belongs in it rather than adding configurations on request.
Explain tiering — the full matrix over a few representative screens plus one reference configuration everywhere else — and why engine and theme coverage are properties of components rather than of every page.
Demonstrate evidence-based pruning: attribute genuine failures to the configuration that caught them, and remove configurations with no yield instead of keeping them as free insurance.
Own the policy and the budget: make the support policy and breakpoint list artifacts outside the suite, cap the diffs an ordinary change may produce, and make removal the scheduled default rather than an exception.
## The failure mode you are governing against Matrices grow by well-intentioned increments. Someone ships a Safari bug, so WebKit is added. Someone ships a dark-mode bug, so a colour-scheme axis is added. Someone has a large monitor, so a 1920px width is added. Each decision is locally sensible and none of them is ever reversed, because removing coverage feels like accepting risk. Meanwhile the axes multiply, and the suite crosses the threshold where a human stops reading the diffs. That threshold is the thing to govern. A visual suite is unusual among test suites in that a failure is not self-evidently a bug: every red diff requires a person to decide "intended or not". Once an ordinary design change produces hundreds of diffs, reviewers approve in bulk, and a real regression is approved along with the intended ones. At that point the large matrix is not merely expensive — it detects less than a small one would. So the governing number is review throughput per change, and the matrix is sized to fit inside it. ## Deriving the matrix rather than accumulating it Three inputs should determine the configurations, and all three live outside the test suite: **A published browser support policy.** Which engines the product commits to is a business decision with support and revenue implications. Written down once, it settles the engine axis and takes it out of per-team debate. Adjust it for what informal processes already cover: the engine everyone develops in needs the suite least. **The design system's declared breakpoints.** Capture widths derive from where the layout actually changes. Making the width list a versioned artifact of the design system means one change updates every team's matrix, and no team invents its own device-shaped widths. **Which product surfaces are load-bearing.** Not every page deserves the same coverage. A checkout flow and a component library's controls are worth far more baselines than a rarely-visited settings page. ## Tiering: the main structural lever A full Cartesian product over the whole page set is almost never right. The shape that works is: - **Tier 1** — a small set of screens and components, chosen because collectively they exercise every layout branch, every theme, and every unusual CSS feature the product uses. These get the full matrix. Ten of these across 12 configurations is 120 baselines: reviewable. - **Tier 2** — everything else, at one canonical configuration. Broad shallow coverage that still catches the "someone deleted a stylesheet import" class of regression. The insight is that engine and theme differences are properties of *components*, not of pages. Confirming that a button renders correctly in WebKit does not need to happen on 40 pages; it needs to happen once, on a page that contains the button. Tiering converts a product of large numbers into a small number times a small number. ## Entry and exit rules Growth stops when removal is as routine as addition. **Entry rule.** To add an axis or a configuration, name a specific defect — ideally one that actually shipped — that this configuration would have caught and no existing configuration would. "It would be more thorough" is not an argument; every configuration is more thorough, which is precisely the problem. **Exit rule.** Instrument the suite so every genuine (non-approved, actually-fixed) failure is attributed to the configuration that caught it first. Review that ledger on a fixed cadence. A configuration that has caught nothing real in two quarters is not free insurance; it is consuming review attention that the productive configurations need. Delete it, and note that you can always add it back under the entry rule. **A budget cap.** State the maximum diffs an ordinary change may produce. When a proposed addition would breach it, something else comes out. Caps make the trade explicit instead of letting it happen by drift. ## Push coverage off the matrix where something cheaper exists Several things people buy with baselines have cheaper detectors. Missing dark-mode token values can be found by linting the token set rather than by doubling every image. Responsive-image asset selection can be asserted directly instead of via a second device-pixel-ratio axis. Contrast and semantics have their own tooling. Each of these removes a multiplicative axis and replaces it with an additive check — the single highest-leverage move available, because it changes the exponent rather than the constant. ## What to say Name review throughput as the constraint, derive the configurations from artifacts that live outside the suite, tier so the full matrix runs over a representative few, and make pruning scheduled and evidence-based rather than exceptional. The senior-sounding version optimises the matrix; the principal-sounding version changes who decides what goes in it and makes removal the default outcome of a review that happens on a calendar.
- How would you measure the defect yield of a single configuration in practice?Attribute every genuine failure — one that was rejected in review and led to a fix — to the configuration that caught it first, and record it. Over a quarter you get a ledger of which configurations found real bugs and which only ever produced approved diffs. It is imperfect, since a configuration can be the second to catch something, but the zero-yield entries are unambiguous and those are the ones you act on.
- A team wants to add a 1920px viewport because their designers use large monitors. How do you respond?Ask what the added width renders that 1440px does not. If the layout has a max-width, both widths render the same branch with different margins, and the yield is near zero. If it is unbounded, there is a real case — line lengths and stretched components. Either way it lands against the budget: if it goes in, something with no measured yield comes out.
- Is there a point at which you would shrink the visual matrix to almost nothing?Yes — when the diffs stop being read. A suite whose output is bulk-approved detects nothing while still costing money and attention, so cutting to one configuration over a tier of key screens is strictly better than an unread full matrix. I would rather have a small suite people trust and act on, and grow it back deliberately under the entry rule once the review habit is real.
saying these in an interview costs you the question
- Treats more configurations as strictly better coverage
- Sizes the matrix against CI capacity rather than reviewer attention
- Adds axes on incident and never removes any
- Runs the full matrix over every page instead of a representative tier
- Buys with baselines what a lint rule or targeted assertion would catch