You own a component library with hundreds of components and no accessibility assertions in its test suite. How would you introduce automated accessibility checks without stalling delivery, and what would you commit to covering outside the suite?
answer
- ratchet, never big-bang
- default-on at the shared seam
- new and touched must pass
- fix primitives before the long tail
- violation count is a gameable metric
basics
~20 sRatchet rather than big-bang: make the check on by default for new and touched components, record existing failures as an explicit accepted backlog, and tighten the rule set in waves. Outside the suite, commit to author-time linting, a browser check for rendering-dependent rules, and manual keyboard and screen-reader review of key flows.
solid answer
~50 sTurning every rule on at once produces hundreds of red tests, and the predictable outcome is that people disable the check rather than fix the components — so I would ratchet. The check goes into the shared render helper so it is on by default, with a narrow, high-confidence rule set first; existing failures become an explicit, owned backlog rather than silent suppressions, and the gate is "no new violations" from day one. Then I widen the rule set in waves as the backlog drains. Outside the suite I would commit to three things the component tests structurally cannot do: author-time linting so the cheapest mistakes never reach a test, a real-browser pass for contrast and layout-dependent rules — ideally by validating the design tokens once instead of per component — and scheduled manual keyboard and screen-reader review of the flows that carry the product's value. And I would resist making violation count the metric, because it rewards suppression.
go deeper
Understand that adding checks to an existing codebase surfaces many failures at once, and that a gradual rollout exists so teams are not blocked on fixing everything before shipping anything.
Be able to describe the mechanics of a ratchet: the check on by default in the shared helper, a narrow starting rule set, existing failures recorded rather than suppressed, and new or touched code required to pass.
Show that you sequence by leverage — primitives and design tokens before the long tail — and that you name the layers the suite cannot cover, with owners: author-time linting, browser checks for rendering-dependent rules, and scheduled manual review.
Own the incentives. Be ready to argue why violation count is the wrong metric, what you report instead, how you keep suppressions from becoming invisible debt, and how the programme changes what engineers know rather than only what the gate blocks.
## Why the big-bang fails Enabling a full rule set across a mature library produces a wall of red. Under delivery pressure, teams do the rational thing: they disable rules, skip tests, or add a blanket suppression — and once that reflex is learned it is applied to genuine failures too. The rollout has then made the codebase *less* safe than before, because the mechanism exists and is trusted while quietly checking nothing. So the design constraint is not "how do we find all the violations" — that is easy — but "how do we introduce a gate that people will not route around". ## The ratchet The pattern that works is a ratchet: the gate never blocks work that was already in flight, and the situation can only improve. **Default-on at the seam.** Put the check inside the shared render helper the library's tests already use, so new components inherit it without anyone remembering to add it. If accessibility assertions are opt-in, they track the enthusiasm of individual authors, which decays. **Narrow first.** Start with rules that are high-confidence, unambiguous and cheap to fix — missing names, missing labels, invalid ARIA values — and restricted to conformance tags rather than the full best-practice set. A small rule set with zero false alarms builds the credibility you will spend later when you widen it. **An explicit, owned backlog.** Existing failures are recorded as a visible list with owners and dates, not scattered inline suppressions. The distinction matters enormously: a list is an artefact someone can burn down and a reviewer can question; inline suppressions are invisible and immortal. Whatever mechanism you use, the property you need is that adding to the list is a deliberate, reviewable act. **Ratchet the gate.** From day one, a component that is new or being touched must pass. That converts the cleanup into a by-product of ordinary work, which is the only budget that reliably exists. **Widen in waves.** As the backlog drains, add the next tier of rules, announce it, and give teams the fix pattern rather than just the failure. Each wave is a small, absorbable batch instead of a wall. ## Sequencing by leverage Within the library, the order is not arbitrary. Fix the primitives first — the button, the input, the dialog, the menu — because one defect there is present on every screen that composes them, and one fix propagates the same way. The long tail of one-off components can wait; they have a fraction of the reach. The same logic applies to design tokens. Contrast defects usually originate in the palette rather than in any component: validate the permitted foreground/background pairs once, at the source, and every component that uses tokens becomes correct by construction. That is a far better return than auditing components one at a time, and it removes a whole class of defect from the backlog before you start. ## What you commit to outside the suite The component suite covers one layer. A credible programme names the others and who owns them. - **Author time.** Static linting of markup as it is written catches a meaningful share of the same mistakes seconds after they are typed, at the lowest possible cost. It is the cheapest layer and should be first, not last. - **Rendering-dependent rules.** Contrast, target size, overlap and reflow cannot be decided where there is no layout or paint. Commit to checking them where pixels exist — token validation plus a real-browser pass over assembled pages — and state plainly that the component suite makes no claim about them. - **Human review.** A keyboard pass and screen-reader spot checks on the flows that carry the product's value, on a schedule, plus once per genuinely new interaction pattern. This is where meaning, order, timing and operability get judged, and no amount of tooling substitutes for it. - **Feedback into the ratchet.** Each manual finding gets its machine-checkable residue encoded as an assertion, so the manual scope shrinks over time instead of repeating. ## The metric trap The most consequential decision is what you report. "Violations" is a tempting metric and a bad one: it can be driven to zero by suppression, it says nothing about the judgment-based majority of real issues, and it invites gaming precisely because it is easy to measure. Better signals are the share of library primitives that pass the current rule tier, the age and size of the accepted backlog, the number of *new* suppressions added per quarter, and — the only one that really matters — outcomes from manual review and from users. Report the automated number as coverage of a floor, never as accessibility. ## Making it stick Tooling changes what fails; it does not change what people know. Pair the rollout with a definition of done that names accessibility explicitly, a short internal guide showing the fix for each rule you enable, and a review expectation that the accessible name is treated as user-facing copy. Without that, the ratchet holds the line but nothing improves upstream of it — and you will spend forever fixing the same defects one wave at a time.
- Why is an explicit accepted-failure list better than inline suppressions in each test?Because it is an artefact with a size, an age and owners. A list can be burned down, reported on, and questioned in review; inline suppressions are invisible in aggregate, so nobody can tell whether the count is going up. It also makes adding an exception a deliberate act rather than the path of least resistance.
- Which components would you fix first, and why not just go alphabetically?The primitives the rest of the library composes — button, input, dialog, menu. A defect there appears on every screen that uses them, so one fix has hundreds of times the reach of a fix in a one-off component. Sequencing by leverage also produces visible improvement early, which is what buys the political room for the later waves.
- How would you handle contrast, given the component suite cannot decide it?Move it upstream. The colours come from a finite set of design tokens, so validating every permitted foreground/background pairing once makes every token-using component correct by construction. What remains — components composing colours in unforeseen ways — is caught by a real-browser pass over assembled pages, and I would state explicitly that the component suite makes no contrast claim.
- What would you report to leadership instead of a violation count?The share of library primitives passing the current rule tier, the size and age of the accepted backlog, new suppressions added per quarter, and findings from scheduled manual review. Violation count can be driven to zero by suppression and ignores the judgment-based majority of real issues, so reporting it as an accessibility number is actively misleading.
saying these in an interview costs you the question
- Enables every rule at once and expects teams to fix it
- Uses violation count as the accessibility metric
- Suppresses failures inline instead of tracking them
- Treats the component suite as the whole programme
- Audits one-off components before shared primitives