skip to content

Your team's React Testing Library suite is green on every pull request, yet UI regressions keep reaching production. A review shows tests built almost entirely on test ids, mocked-out child components, and assertions against store state. As the lead, how do you decide what to change and in what order?

level: principalimportance: should knowfreq 30%

answer

  1. green but insensitive
  2. coverage already said it was fine
  3. break it on purpose, count the catches
  4. convert by risk, then opportunistically
  5. fix the friction or it returns

basics

~20 s

Start by measuring sensitivity, not coverage: break components deliberately and see which tests stay green. Then rewrite the highest-risk flows first, as they are touched, and change the defaults so new tests do not repeat the pattern.

solid answer

~50 s

First establish the problem is real and quantified. Coverage numbers will look fine, so measure differently: deliberately break a handful of components — remove a label, render the wrong value, drop a child — and record how many tests notice. That gives a sensitivity number leadership can act on and tells you which pattern is costing the most. Then sequence by risk, not by alphabet: rewrite the tests covering checkout, auth and the flows that generated the actual escaped defects, and convert the rest opportunistically as files are touched, rather than freezing feature work for a big-bang migration. In parallel, change the default so the debt stops growing — a custom render helper that supplies the real providers, fixtures that make rendering real children cheap, one exemplary test per pattern to copy, and review attention on new tests. Finally, re-run the deliberate-breakage exercise after a quarter to show the number moved.

go deeper

for a junior

Know the individual patterns well enough to avoid writing them: query the way a user finds things, render the real children, assert what is rendered rather than what is stored.

for a middle

Be able to explain why a green suite can still miss regressions — each pattern removes a specific kind of sensitivity — and to convert one bad test into a meaningful one on request.

for a senior

Show how you would prove the gap rather than assert it: reintroduce real regressions, count what the suite catches, and use that to pick which tests to rewrite first.

for a principal

Own the sequencing and the economics — risk-ordered and opportunistic conversion over a big-bang rewrite, changing the defaults so the debt stops growing, refusing the coverage-target reflex, and closing with a measured delta.

## First, name what is actually broken The suite is green and regressions still ship, so the suite is insensitive, not incomplete. That distinction drives everything. Adding more tests of the same kind makes the numbers better and the product no safer. The three patterns found in review each remove sensitivity in a specific way: - **test ids everywhere** — the test finds the node no matter what the node says, so label, copy and reachability regressions pass; - **mocked children** — the composed UI never renders, so every parent-to-child integration bug passes; - **state assertions** — the render half of the component's job is never checked, so stale views and formatting bugs pass. Stack them and you have a suite that verifies the code executes, which is roughly what a coverage tool already told you. ## Measure sensitivity before you spend anyone's quarter The most persuasive move is cheap: take five to ten recent production regressions, reintroduce each one on a branch, and run the suite. Then do the same with synthetic breakages — remove a button's accessible name, render the wrong field, make a child render nothing, format a number wrong. Count how many the suite catches. The output is a single number — "our tests catch 3 of 10 real regressions" — that survives contact with a planning meeting in a way "our tests are low quality" does not. It also localises the damage: if the misses cluster in pages whose children are stubbed, that is where the first fixes go. This is a poor engineer's mutation testing, and if the codebase can afford a real mutation-testing run, use that instead; the reasoning is identical. ## Sequence by risk, and prefer opportunistic conversion A big-bang migration of an old suite is the classic way to burn a quarter and get halfway. Better ordering: 1. **Flows where an escape actually hurt** — checkout, payment, auth, anything that generated an incident. Rewrite those tests properly now, on purpose, with time allocated. 2. **Files under active change.** Attach conversion to work already touching the file, so the cost lands where somebody is already paying attention. 3. **Everything else: leave it.** An insensitive test that costs nothing to keep is not worth a rewrite budget on its own. It becomes worth rewriting the day it obstructs a change. An important corollary: delete rather than convert where the test asserts nothing meaningful. A test whose only assertion is that a stub received a prop can go, and the honest coverage number that results is more useful than the fiction it replaces. ## Remove the reason the pattern happened Every one of these mistakes is usually a workaround for friction, and if the friction stays, the pattern comes back into every new test. - People stub children because setup is painful → give them a custom render that wires the real providers, plus realistic fixture builders, so rendering the true tree is the path of least resistance. - People use test ids because role queries are awkward → some of that is missing labels in the product, which is a real bug worth fixing; the rest is unfamiliarity, which one worked example fixes. - People assert on store state because nothing visible is easy to assert → that often means the component's responsibilities are split badly, and it is worth looking at. Then make the new default stick: a short written convention, lint rules where the ecosystem provides them, and — most effective in practice — reviewing new tests with the same seriousness as production code. One senior engineer consistently asking "what would this test catch?" changes a team's habits faster than any document. ## What you deliberately do not do - Do not chase a coverage target; it is the metric that produced this suite. - Do not respond by pushing everything to end-to-end tests. They are slower, flakier and more expensive per assertion; the right fix is component tests that actually assert. - Do not rewrite the suite silently as a side project. The number from the sensitivity exercise is what buys the time. ## Close the loop Re-run the breakage exercise a quarter later on the converted areas and report the delta, alongside escaped-defect counts for those flows. That is what makes it a completed piece of engineering leadership rather than an opinion about test style — and it is what an interviewer at this level is listening for: a diagnosis, a sequencing rule tied to risk, a change to the defaults so the debt stops growing, and evidence at the end.

  • How do you keep new tests from regressing to the old patterns once the migration is underway?
    Make the good path the easy one: a shared render helper that supplies real providers, fixture builders, and one exemplary test per pattern to copy from. Add lint rules where the ecosystem offers them, and put new tests under real code review — a reviewer asking "what would this catch?" on every PR moves habits faster than a style document.
  • The team proposes covering the gap with more end-to-end tests instead. How do you respond?
    Partly agree, then bound it. A handful of end-to-end tests over the critical flows is worth having regardless. But they cost minutes per run, flake under real network and timing, and are the most expensive place to assert a label or a formatted total. Fixing the component tests is cheaper per assertion and gives faster feedback; e2e should cover what only it can — real navigation, real integration across services.
  • How do you justify the time to a product owner who sees a green suite and no obvious problem?
    With the escaped defects and the sensitivity number together: these five incidents shipped through a green suite, and when we reintroduced them the tests caught two of ten. That reframes it from test-style preference to a measured gap in the release gate, with a cost the incidents already demonstrated.

saying these in an interview costs you the question

  • Raise the coverage threshold and the quality follows
  • Freeze features and rewrite the whole suite at once
  • Add end-to-end tests to cover everything the unit tests miss
  • The tests are green, so the process is working
  • Delete the old suite and start over from scratch

context