In a component library shared by many product teams, which test layers do you run, and what does each catch that the others miss?
answer
- behavior, pixels, rules, real usage
- no single layer sees everything
- variants times themes multiplies states
- a scan cannot press keys
- people still judge meaning
basics
~20 sInteraction tests catch broken behavior and keyboard contracts; visual diffs across variants and themes catch rendering regressions; automated accessibility scans catch rule-detectable defects; consumer-contract tests catch breaks in real usage; manual assistive-technology review covers the rest.
solid answer
~40 sI run four automated layers plus a manual one, because each is blind where another sees. **Interaction tests** drive components by role and name and pin behavior - a menu opens on Enter, a dialog returns focus. **Visual diffs** render each variant and state, across themes, and catch pixel regressions such as a stage label whose color drifted in one brand theme. **Automated accessibility scans** catch rule-detectable defects like an icon-only button with no name. **Consumer-contract tests** run what real consuming screens rely on against a release candidate. Then **manual review with assistive technology** judges what no tool can: whether names make sense and the experience holds together. A green result in one layer says nothing about the others.
go deeper
Recall the layers - interaction, visual, automated accessibility, consumer contract, manual review - and one kind of failure each is built to catch.
Explain each layer's blind spot, such as why a visual diff misses a broken key press and a scan cannot judge a name's meaning.
Show how you place the layers in a pipeline by speed and determinism, and how you keep theme and variant coverage honest without rendering every combination.
Discuss how to fund the expensive layers - manual review and consumer contracts - and which risks you accept when a team cannot afford all of them.
## Why a library needs layers A **component library** is consumed by many product teams who upgrade without reading every diff. Its tests are the only thing standing between a regression and dozens of apps. No single kind of test can guard every promise a component makes - behavior, appearance, accessibility semantics and fitness for real usage are different properties, observed in different ways. A test strategy is therefore a set of **layers**, each chosen for the failures it can see. ## The layers side by side In a recruiting tool's shared library - candidate cards, stage labels, an interview-slot picker, an action menu: | Layer | Catches | Misses | Relative cost | |---|---|---|---| | **Interaction tests** by role and name | Broken behavior, keyboard contracts, lost accessible names, wrong states | Visual drift, contrast, anything not asserted | Low per test, many tests | | **Visual diffs** across variants and themes | Layout shifts, color drift, clipped text, a theme that breaks one variant | Behavior, semantics, states nobody rendered | Moderate; baselines need review | | **Automated accessibility scans** | Rule-detectable defects: missing names, invalid role use, some contrast failures | Meaning of names, keyboard flow, focus order, composition | Low; runs on every render | | **Consumer-contract tests** | Breaks in how real screens use components | Anything consumers do not exercise | Higher; needs coordination | | **Manual assistive-technology review** | Announcement quality, sense of names, reading order, overall experience | Regressions between reviews | Highest; done at milestones | ## Where each layer's ceiling lies - **Interaction tests** check only what they assert. A test that never presses Escape cannot catch a broken Escape. - **Visual diffs** see only the states someone rendered. An error state or an open menu that was never captured is never compared, and a correct-looking screenshot can hide a component that no longer responds to keys. - **Automated scans** evaluate rendered structure against rules. They do not press keys, cannot judge whether 'Button 1' is a useful name, and see one component in isolation rather than the page it lands on. - **Consumer-contract tests** cover what consumers exercise, which is valuable precisely because library authors did not think of it - but it is never complete. - **Manual review** is the only layer that judges meaning, and the only one too expensive to run on every change. ## Variants and themes multiply the visual layer Every **variant** (size, emphasis, state) and every **theme** (brands, contrast modes, densities) multiplies the states a visual layer could render. A single token change can alter every component in one theme while leaving others untouched, so theme coverage cannot be an afterthought. A common approach is to render the full variant and state set in the default theme and a representative subset in each other theme, widening the subset for components that consume the most tokens. ## Putting the layers into the pipeline 1. On **every change**: interaction tests and automated scans for the touched components, since both are fast and deterministic. 2. On **every change that renders differently**: visual diffs for affected components across the chosen theme set, with a human approving intended changes. 3. On **release candidates**: consumer-contract tests against the selected consuming apps. 4. At **milestones** - a new component, a changed interaction model, before marking a component stable: manual review with a screen reader and keyboard on each supported platform. The same structure applies when the library ships to native mobile: interaction tests query by accessibility label and trait, snapshot tests compare rendered views, platform accessibility checkers play the scanner's role, and manual review happens with the platform's own screen reader. The point interviewers look for is not the list of tools but the reasoning: each layer is justified by a class of failure the others cannot see, and none of them is allowed to stand in for another.
- Why not rely on visual diffs alone, since they see everything rendered?They see pixels, not behavior or semantics. A menu that renders perfectly but no longer opens on Enter, or a button that lost its accessible name, produces an identical screenshot. Visual diffs also cover only the states someone captured, so an open or error state never rendered is never compared.
- Where does manual testing fit in a library's strategy?At milestones: before a component is marked stable and whenever its interaction model changes. A person using a screen reader and keyboard on each supported platform checks announcement quality, name meaningfulness, reading order and focus visibility. Automation then guards against regressions between reviews; it cannot judge whether the experience makes sense.
- Which layers should block a merge, and which only a release?Fast, deterministic layers - interaction tests and automated scans - usually block every merge. Visual diffs block too, but with a human approval step for intended changes. Consumer-contract tests are slower and depend on other teams' code, so they typically gate release candidates rather than each merge.
Like inspecting a new building: the structural engineer, the electrician and the fire inspector each sign off on different risks, and a clean electrical report says nothing about whether the stairs hold.
saying these in an interview costs you the question
- Full unit-test coverage of a component makes visual tests redundant.
- A clean automated accessibility scan means the component meets WCAG.
- Visual diffs catch keyboard regressions because they render the component.
- The library's own tests guarantee that no consuming app will break.
- Themes need no visual coverage once the default theme is tested.