In a design system with several hundred Storybook stories, how do you decide which stories get a visual baseline, and what does each additional baseline cost?
answer
- the budget is reviewer attention
- select for blast radius and contract
- primitives and token surfaces first
- composed page stories attribute poorly
- noisy baseline is a defect, not a cost
basics
~20 sBaseline the stories whose appearance is a contract others depend on — primitives, token surfaces, and states like disabled, invalid and selected. Each baseline costs a render per run and, far more expensively, a change a human must judge.
solid answer
~50 sThe binding constraint is not compute, it is human attention. Every baselined story is a potential change a person has to look at and judge, and a suite that surfaces dozens of changes per pull request gets bulk-approved — at which point it costs money and detects nothing. So I select for blast radius and contract. Primitives and token-driven surfaces earn baselines because one change there propagates everywhere. States that encode a visual contract — disabled, invalid, selected, density variants — earn them because that is what consumers rely on. Documentation-shaped stories, deliberately variable content, and big composed page-like stories usually do not: their changes are frequent, hard to attribute, and teach nobody anything. Then I treat a chronically noisy story as a defect to fix or drop, and make the required story set part of a component's definition of done rather than each team's taste.
go deeper
Know that each baselined story is another capture and another change someone must look at, so teams choose which stories to pin rather than pinning all of them.
Be ready to justify a selection: primitives and contract states first, documentation and randomised-content stories excluded, and say why each choice pays off.
Argue the attention budget explicitly and show how you handle noisy baselines and intentional sweeping changes without training reviewers to approve blindly.
Own the policy end to end — how the set is sized against real review capacity, how the required story set is enforced through scaffolding, and when you contract coverage rather than expand it.
## Name the real budget Engineers instinctively price a visual suite in CI minutes or in a vendor's capture count. Those are real but rarely binding. The binding resource is *reviewer attention*: a visual regression suite only works if a human looks at each reported change and decides whether it was intended. That decision does not parallelise, does not cache, and degrades sharply with volume. The failure curve is well known. A suite reporting two or three changes per pull request gets read carefully. One reporting forty gets a bulk approval, and from then on it is a ritual that consumes CI time and catches nothing — arguably worse than no suite, because it manufactures confidence. So the design question is not "how much can we cover" but "how many judgements per change can this team actually make well". ## Selection criteria Given that budget, baseline for **blast radius** and **contract**. **High value:** - **Primitives and token surfaces.** Button, input, typography scale, elevation, spacing samples. A change here propagates to every product surface, so a pinned appearance has enormous leverage per baseline. - **States that encode the visual contract.** Disabled, invalid, selected, compact density. Consumers build on these looking a specific way; a silent change breaks them everywhere. - **Known-fragile layouts.** The component with the truncation rule everyone gets wrong, the one with the tricky overflow. Past bugs are the best predictor of future ones. **Low value, often negative:** - **Documentation-shaped stories** — prose, usage guidance, kitchen-sink pages. They change for editorial reasons and generate reports with no design meaning. - **Stories with intentionally variable content** — anything randomised or content-driven by design. - **Large composed page-like stories.** They report a change whenever any constituent changes, and the reviewer cannot tell which. One primitive change lights up thirty of them, and the reviewer learns nothing they did not already learn from the primitive's own report. ## The other costs, honestly stated - **Run time.** Each baselined story is a render and a capture; interactive stories add more. This shapes how fast feedback arrives, which shapes whether people wait for it. - **Vendor cost.** Cloud services meter captures, so the baseline set is a line item, not just an engineering choice. - **Maintenance.** Baselines are artefacts that need updating when change is intentional, and every batch update is another chunk of judgement. ## Governance, not per-PR policing Three rules keep this stable over time. **A chronically noisy story is a defect.** If a story reports a change on most runs for reasons nobody caused, the options are to fix the source of variation or to drop it from the baseline set. What is not an option is approving it every run — that is how reviewers learn to approve without looking, and the habit spreads to the stories that matter. **The required story set belongs to the component, not to the team.** Encode it in whatever scaffold or template new components start from, and check for it in review. Otherwise coverage is a function of who wrote the component and how busy they were. **Intentional sweeping change is a planned event.** A token change that legitimately alters three hundred baselines should be its own change with no functional edits mixed in, reviewed as a sample plus the outliers, and accepted as a batch. The value of that run is confirming the change reached exactly the surfaces you expected — not eyeballing three hundred near-identical shifts. Mixing a token change into a feature pull request is what destroys a review culture in one afternoon. ## How I would phrase the decision Start deliberately small — primitives and contract states only — and measure the change rate per pull request for a few weeks. If reviewers are handling it comfortably, expand into the next tier of components. If not, contract before adding anything. That framing matters in a principal-level answer: coverage is not maximised, it is *tuned to the review capacity the organisation actually has*, and it is revisited as that capacity changes.
- A design-token change legitimately alters three hundred baselines. How do you handle that?Ship it as its own change with no functional edits mixed in. Review a representative sample plus anything that shifted unexpectedly, then accept the batch in one pass. The value of that run is confirming the change reached exactly the surfaces you predicted — not inspecting three hundred near-identical shifts one at a time.
- How do you enforce a required story set without policing every pull request?Put it where the work starts: the component scaffold or template ships with the expected states, so writing them is the default rather than an extra task. Add one line to the review checklist for new visual variants, and backfill from real escapes. Enforcement by convention and tooling scales; enforcement by vigilance does not.
- How would you know the baseline set has grown too large?Watch the review behaviour, not the numbers. Rising changes per pull request, approvals landing within seconds of the run finishing, and reviewers accepting in bulk all say the suite has passed the attention budget. That is the signal to contract the set or fix its noisiest members before adding any further coverage.
A baseline set is like an alarm system: adding sensors feels like more safety right up to the point where it goes off nightly and everyone stops getting out of bed.
saying these in an interview costs you the question
- Baselines every story on the principle that more coverage is always better.
- Prices the suite in CI minutes and ignores reviewer attention entirely.
- Keeps a chronically noisy baseline and approves its change every run.
- Leaves the required story set to each team's individual taste.
- Mixes a sweeping token change into an ordinary feature pull request.