skip to content

Component Testing Philosophy

You will learn what a component test should actually assert: user-visible behavior rather than internals, when a snapshot earns its keep, how to test a custom hook honestly, and how far automated accessibility checks get you. This is the reasoning layer beneath any particular testing library.

on this pageshow

questions

19

Your component test suite runs an automated accessibility rule check on every component and reports zero violations. What can you honestly claim from that, and which accessibility problems does it leave undetected?

level: middleimportance: must knowfreq 55%

answer

  1. presence is checkable, meaning is not
  2. a string exists is not a string is good
  3. order, focus and timing go unchecked
  4. no layout means no contrast verdict
  5. roles pass, behaviour still missing

basics

~20 s

Zero violations only means no enabled rule found a machine-decidable markup defect. Automated rules cannot judge whether names and alt text are meaningful, whether focus order and visibility work, or whether announcements arrive usefully — those need manual keyboard and screen-reader checks.

solid answer

~50 s

The honest claim is narrow: for each component, no enabled rule found a defect it can decide from the markup — no missing `alt`, no unlabelled input, no illegal `aria-*` value, no nameless button. What it leaves untested is everything requiring judgment or real rendering: whether the alternative text says something useful, whether a control's name matches its visible label, whether tab order follows the visual layout, whether focus is visible and gets moved and restored around dialogs, whether a status change is announced at a moment the user can use, and whether a widget that has the right roles actually responds to the keyboard. Published estimates put automated coverage somewhere between a third and half of real issues. So I treat the suite as a regression floor and pair it with a manual keyboard pass, a screen-reader spot check on the critical flows, and a real-browser check for anything involving colour or layout.

go deeper

for a junior

Be able to say that automated checks look at markup facts like missing labels or alt text, and that they cannot tell whether the text is meaningful — so a green run is a starting point, not proof.

for a middle

Explain the categories the rules structurally cannot reach: meaning of names, focus order and visibility, timing of announcements, keyboard behaviour behind correct roles, and anything needing real layout. Name the compensating manual checks.

for a senior

Demonstrate that you run a real process — a keyboard pass, screen-reader spot checks on high-value flows, a browser check for rendering-dependent rules — and that you convert each manual finding into a permanent assertion where it is machine-checkable.

for a principal

Own the definition of done and the metric. Be ready to argue why violation count is a misleading target, how much manual audit budget the product warrants, and how you keep the automated floor from crowding out judgment-based review.

## What a rule engine can decide An automated accessibility check answers questions that are decidable by inspecting the DOM. Those questions fall into a few families: - **Presence**: does this image expose alternative text, does this input have an associated label, does this button end up with a non-empty name. - **Validity**: is this ARIA attribute one that exists, is its value legal, is it allowed on this element's role. - **Structural relationships**: does a role that requires particular children have them, do referenced ids resolve. All of these are objective. A machine reads the markup, applies a fixed rule, and returns a verdict that a reasonable human would agree with. That is genuinely valuable — these are also the mistakes people make most often and fix most cheaply, and once a rule is in the suite the defect can never silently return. ## Where the ceiling is The ceiling is not arbitrary; it follows from the same property that makes the rules automatable. A rule can check that a string exists. It cannot check that the string is *good*. **Meaning.** `alt="image"`, `alt="IMG_2043.png"` and `aria-label="button"` all satisfy the presence rules and all fail the user. Same for a label left behind after the copy changed: the control now reads "Save" but announces "Submit form", so a voice-control user saying "click Save" gets nothing, and the check stays green. **Order and focus.** Tab order is a function of DOM order, tabindex and layout. A rule can flag a positive tabindex as suspicious, but it cannot tell you that a visually two-column form tabs down the left column and then jumps back up to the right, which is disorienting. Nor can it tell you that the focus indicator is invisible against the component's background, that a dialog opens without moving focus into it, or that closing the dialog drops focus back to the top of the document instead of the control that opened it. **Announcements over time.** Accessibility is partly temporal. Did the error message reach the user when validation failed? Did the results count announce once, or on every keystroke until it was unusable? Was the loading state communicated? A single-snapshot rule pass over a rendered tree has no notion of what happened when. **Interaction behaviour.** Correct roles do not imply correct keyboard behaviour. A custom widget can carry exactly the right roles and states, pass every rule, and still not respond to arrow keys, Home/End, or Escape — so it announces itself as something the user cannot operate. Roles are markup; behaviour is code. **Anything requiring rendering.** Contrast, target size, overlap, reflow at zoom, and content that becomes unreachable at a narrow viewport all depend on layout and painting. In a simulated DOM there is no layout and no paint, so these rules cannot even produce a verdict — they return undecided results, which do not fail the test. **Cognition and flow.** Whether an interface is understandable, whether an error tells the user how to recover, whether a multi-step task can be completed without sighted assistance — nothing automatable touches these. This is why the widely quoted figures for automated coverage sit between roughly a third and half of real-world issues, depending on the site and who is counting. The exact number matters less than the shape of the gap: automation covers the objective floor, humans cover the judgment. ## What to say you will do instead A good answer does not stop at "automation is limited" — it names the compensating checks: - **A keyboard pass.** Tab through the component or flow with no mouse. Can you reach everything interactive, in an order that matches the visual layout, with a visible focus indicator, and get back out without a trap? - **A screen-reader spot check** on the flows that carry the product's value — sign-up, checkout, search — rather than on every component. You are listening for whether the announced name matches what is on screen and whether state changes are communicated. - **A real-browser check** for the rendering-dependent rules the simulated environment could not decide, so contrast and layout problems are caught somewhere rather than nowhere. - **Turning findings back into tests.** Once a manual pass finds a real defect, encode the machine-checkable part of it: assert the dialog moves focus, assert the error message is associated with the input, assert the control exposes the exact expected name. This is how the manual budget shrinks over time instead of repeating the same discoveries. ## The organisational failure mode The dangerous outcome is not a missed rule; it is a team that adopts the green suite as its definition of done. Violation count then becomes the metric, and effort flows toward the things that move the metric rather than toward the things users hit. Stating the claim precisely — "no enabled rule fired" — is what keeps the rest of the programme honest. ```js // The rule check passes; both of these are defects. <img src="chart.png" alt="image" /> <button aria-label="Submit form">Save</button> ```

  • Give a concrete example of markup that passes every rule and is still broken for users.
    `<button aria-label="Submit form">Save</button>`. Every presence and validity rule is satisfied: the button has a name and the attribute is legal. But a screen-reader user hears a different word from the one on screen, and a voice-control user saying "click Save" gets no match. Only a human comparing the announced name to the visible label catches it.
  • How do you decide which flows get a manual screen-reader pass, given the budget is finite?
    Rank by user impact and irreversibility: the flows where failure costs someone the ability to complete a real task — authentication, payment, search, forms that submit data. Cover those on a schedule, cover new interaction patterns once when they are introduced, and let the automated suite carry the long tail of repeated components.
  • How do you keep manual findings from being rediscovered every quarter?
    Convert the machine-checkable residue of each finding into an assertion. If the audit found focus was not moved into a dialog, add a test asserting focus lands inside on open and returns to the trigger on close. The judgment part stays manual, but each pass permanently retires some of its own scope.
  • Is it worth checking contrast in the component suite at all?
    Not in a simulated DOM — the rule cannot resolve rendered colours, so it returns an undecided result rather than a verdict, and a green run says nothing. Check contrast where pixels exist: in a real browser, or by reviewing the design tokens themselves so the palette is correct by construction rather than per component.

saying these in an interview costs you the question

  • Zero violations means the product is accessible
  • Assumes automation checks whether alt text is meaningful
  • Thinks correct roles guarantee correct keyboard behaviour
  • Believes contrast is covered by the component suite
  • Treats violation count as the accessibility metric

context

open as a page

In a UI component test, what counts as an implementation detail, and what concrete rule do you use to decide whether a given assertion will survive a refactor?

level: middleimportance: must knowfreq 76%

basics

~20 s

An implementation detail is anything a user cannot perceive: internal state, render counts, DOM structure, class names, which child components exist. The working rule is to assert only what a user would notice if it stopped being true.

open as a page

You have a custom React hook `useCart` that several components use. When is it right to test it directly with React Testing Library's `renderHook`, and when should you test it through a component that consumes it?

level: middleimportance: must knowfreq 62%

basics

~20 s

Test a hook directly when its return value is the product: a shared hook whose callers are other developers. Test it through a consuming component when the hook exists to serve one component, because there the rendered output is what people depend on.

open as a page

In a component test written with Testing Library and Jest or Vitest, how does a snapshot assertion such as expect(container).toMatchSnapshot() differ from a targeted assertion such as expect(screen.getByRole('heading')).toHaveTextContent('Invoice'), and when does the snapshot stop earning its place?

level: middleimportance: must knowfreq 65%

basics

~20 s

A snapshot asserts that the whole serialized output is identical to a recording; a targeted assertion states one thing a human decided must be true. Snapshots catch every change including irrelevant ones, so once they fail routinely for reasons unrelated to behaviour, they stop being evidence and become a rubber stamp.

open as a page

In a frontend component test, what does an automated accessibility check such as jest-axe's axe(container) followed by expect(results).toHaveNoViolations() actually do, and what does a passing result prove?

level: juniorimportance: should knowfreq 40%

basics

~20 s

An automated accessibility check evaluates a fixed set of rules against the component's rendered DOM and fails the test if any rule is violated. Passing means no enabled rule fired — not that the component is accessible.

open as a page

A UI component test checks that a Save button is in its primary style by asserting `expect(container.querySelector('.btn-primary')).toBeTruthy()`. Why do reviewers call that assertion brittle, and what would you assert instead?

level: juniorimportance: should knowfreq 52%

basics

~20 s

A CSS class is a styling hook, not behavior: renaming or reorganising it breaks the test even though the button still looks and works the same to a user. Assert what a user perceives — the visible label, the role, the enabled state.

open as a page

In a React Testing Library `renderHook` test, a value destructured out of `result.current` before an update still shows the old number when you assert afterwards. Why, and what do you write instead?

level: juniorimportance: should knowfreq 48%

basics

~20 s

renderHook stores each render's return value in result.current and replaces it after every re-render. Destructuring copies the value at that instant, so the local variable is frozen at the old render. Read result.current fresh inside every assertion.

open as a page

In a Jest or Vitest component test, expect(container).toMatchSnapshot() passes the very first time it runs. Why does it always pass on that first run, and what risk does that create?

level: juniorimportance: should knowfreq 55%

basics

~20 s

On the first run there is nothing to compare against, so the runner serializes the current output, writes it to a snapshot file, and passes. The baseline is whatever the code produced — correct or not — so an unreviewed snapshot can freeze a bug in place.

open as a page

In a component test, what does asserting a control's exact accessible name — for example with jest-dom's toHaveAccessibleName — give you that a rule-based axe run does not?

level: middleimportance: should knowfreq 28%

basics

~20 s

A rule check only asks whether a control has some non-empty name; a name assertion pins which name. That catches wrong, stale or meaningless names — the ones a rule check passes — and locks the string users actually hear and speak.

open as a page

If a component test should assert only what a user can perceive, how do you test a component whose job is to invoke an `onSubmit` prop and send a request — outputs a user never sees on screen?

level: middleimportance: should knowfreq 58%

basics

~20 s

A component's observable contract is wider than its pixels: it includes the callbacks it invokes and the requests it sends. Assert those at the component's outer boundary, driven by a real user interaction — never by reaching into internal handlers.

open as a page

Using React Testing Library's `renderHook`, how do you test that a custom hook such as `useDebouncedValue(value, delay)` behaves correctly when its input argument changes, rather than only on first render?

level: middleimportance: should knowfreq 44%

basics

~20 s

Give renderHook an initialProps object, write the callback to take those props, then call the returned rerender with new props. That re-invokes the hook with different arguments in the same mounted instance, which is the only way to observe behaviour that depends on a change between renders.

open as a page

Jest and Vitest both offer external snapshots stored in a companion file (toMatchSnapshot) and inline snapshots written back into the test file (toMatchInlineSnapshot). What does each form buy you, and how would you choose between them?

level: middleimportance: should knowfreq 40%

basics

~20 s

An inline snapshot puts the expected value in the test file, so a reader sees the assertion and its expectation together and a reviewer cannot miss a change. An external file handles values too large to inline, at the cost of living where nobody reads it. Prefer inline whenever the value is small.

open as a page

You add an automated axe check to component tests running in a simulated DOM. A test that renders a lone button fails with "all page content should be contained by landmarks", and colour-contrast problems are never reported at all. Explain both results and how you would configure the check.

level: seniorimportance: should knowfreq 33%

basics

~20 s

Both come from asserting outside what the environment can decide. Page-level rules like landmark containment are meaningless over a rendered fragment, and a simulated DOM does no layout or painting, so contrast rules return undecided results rather than failures. Scope the rule set; check contrast in a browser.

open as a page

A large form component is split into three smaller components, with no change to what renders on screen or how it behaves — yet about forty of its tests now fail. What kinds of assertions cause that, and how would you rewrite the suite so the next refactor costs nothing?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Those tests were asserting structure rather than outcome: internal state, child components and the props passed to them, DOM nesting, class names. Rewrite each around a user-visible outcome of a real interaction, so the same test passes before and after the split.

open as a page

A custom hook you want to test reads from a context provider that your test does not render, so it throws the moment `renderHook` runs it. What are your options, and what does the amount of setup tell you?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Pass a wrapper component to renderHook that renders the real provider with test-controlled values, or test through a component that already sits inside that provider tree. If the wrapper has to rebuild half the app, the hook is coupled to ambient context and the test is an integration test in disguise.

open as a page

A React app's component suite snapshots whole pages. Nearly every pull request turns several snapshots red for reasons unrelated to the change, and the team's routine fix is to re-run the tests with the update flag and commit the regenerated files. What has that suite stopped protecting, and how would you restructure it?

level: seniorimportance: should knowfreq 50%

basics

~20 s

It protects nothing: a baseline regenerated without being read is just a copy of current output, so the tests can only fail, never catch anything. Fix the cause — captures too broad to review — by narrowing them to small subtrees or derived data, adding assertions that state intent, and pruning obsolete baselines.

open as a page

You own a component library with hundreds of components and no accessibility assertions in its test suite. How would you introduce automated accessibility checks without stalling delivery, and what would you commit to covering outside the suite?

level: principalimportance: should knowfreq 22%

basics

~20 s

Ratchet rather than big-bang: make the check on by default for new and touched components, record existing failures as an explicit accepted backlog, and tighten the rule set in waves. Outside the suite, commit to author-time linting, a browser check for rendering-dependent rules, and manual keyboard and screen-reader review of key flows.

open as a page

You own the testing standards for a large frontend codebase that has accumulated several hundred committed snapshots across component tests. How would you decide where snapshot testing genuinely earns its place, and how would you retire the rest without stopping feature work?

level: principalimportance: should knowfreq 33%

basics

~20 s

Judge each snapshot by evidence: has its failure ever caught something a human wanted to know? Keep the small, stable, hard-to-hand-write captures where any change deserves review; ban new broad ones so the inflow stops, and convert the rest on contact instead of in a dedicated rewrite.

open as a page

You are setting the component-testing standard for a large frontend. Behavior-only assertions survive refactors but produce vaguer failure messages and slower tests than assertions on internals. How do you weigh that, and where would you allow exceptions?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Refactor freedom is worth more than diagnostic precision, so user-visible outcomes are the default. Allow narrow, reviewed exceptions where a state genuinely cannot be reached through the UI, and require each exception to name what it is protecting.

open as a page