A large form component is split into three smaller components, with no change to what renders on screen or how it behaves — yet about forty of its tests now fail. What kinds of assertions cause that, and how would you rewrite the suite so the next refactor costs nothing?
answer
- mass failure with zero user-visible change
- child stubs pin the composition
- state, structure, snapshots, render counts
- recover the outcome under the mechanism
- fix on touch, not big bang
basics
~20 sThose tests were asserting structure rather than outcome: internal state, child components and the props passed to them, DOM nesting, class names. Rewrite each around a user-visible outcome of a real interaction, so the same test passes before and after the split.
solid answer
~60 sForty failures from a change with no user-visible effect is a diagnosis in itself: the suite is testing the component's shape, not its behavior. In practice the culprits are a small set — assertions on internal state, mocked child components with checks on the props they received, queries that walk the DOM structure or class names, and often snapshots of the whole tree. I'd start by triaging: for each failing test, ask what user-facing outcome it was actually trying to protect. Usually there is one buried under the mechanism. Then rewrite it as *interact, then assert the outcome*: fill the field the user fills, click the control they click, assert the message, the state of the button, or the callback payload. I'd expect the forty to collapse into far fewer, better tests, because several were asserting different internal routes to the same outcome. I would not do this as one big-bang rewrite — I'd fix the tests as each area gets touched, and hold the new standard on everything written from that day.
go deeper
Recognise the symptom: if nothing a user sees changed but many tests failed, the tests were checking internals. Be able to name one such assertion, like reading internal state.
Walk through the specific categories — stubbed children, state reads, structural queries, snapshots — and show the rewrite as interaction-then-outcome on a concrete failing example.
Diagnose from the failure pattern, expect the test count to shrink as redundancy collapses, and verify the rewritten tests can still fail. Sequence the work so no feature freeze is needed.
Frame it as a standard and an enforcement path across teams: what review rejects from today, how remediation is funded by ordinary work, and which hot spots justify a dedicated investment.
## Reading the symptom A refactor that changes no observable behavior should change no test. When forty go red, the tests were not describing behavior — they were describing the previous internal structure. That is worth saying out loud in an interview, because it reframes the task from "fix forty tests" to "the suite has the wrong contract". ## The usual culprits, in order of frequency **Mocked children with prop assertions.** The old component rendered its fields inline; the tests replaced a child with a stub and asserted it received `{ value: 'x', onChange: fn }`. After the split, the props move to a different child, or the intermediate component now passes a differently-shaped object. Nothing the user sees changed, but every one of those assertions is about an internal wiring diagram. **Internal state assertions.** Tests that reached in for `errors`, `touched` or `isSubmitting`. Splitting the component moves that state — perhaps into one of the children, perhaps into a shared parent — and every such assertion breaks. **Structure-dependent queries.** Anything that walked the tree: `container.firstChild.children[2]`, a CSS selector describing nesting, an assertion on how many `<div>` wrappers exist. A split almost always adds or removes a wrapper. **Whole-tree snapshots.** One extra wrapper element and every snapshot in the file is a diff. These usually get bulk-accepted, which is its own problem — but the immediate symptom is the mass failure. **Render-count and lifecycle assertions.** Splitting a component genuinely changes how many components render and when. If a test pinned those numbers, it fails, and it was never protecting anything a user cares about. ## The rewrite recipe For each failing test, the question is: *what user-facing outcome was this test trying to protect?* Almost always there is one, obscured by the mechanism. Recover it and express it directly. The shape is invariant: 1. Render the **whole** component under test — with the real children, not stubs. The internal composition is exactly the thing you want the test to be indifferent to, and stubbing children hard-codes it. 2. Drive it through a **user-level interaction**: type into the field found by its label, click the control found by its visible name. 3. Assert a **perceivable outcome**: the message that appears, the control that becomes disabled, the row that is removed — or an outward effect such as the callback payload. ```javascript // before: pinned to the old internal shape expect(EmailField).toHaveBeenCalledWith( expect.objectContaining({ error: 'Email is required' }), undefined, ); // after: the outcome a user experiences await userEvent.click(screen.getByRole('button', { name: 'Sign up' })); expect(await screen.findByText('Email is required')).toBeInTheDocument(); ``` The second version is indifferent to whether the error lives in one component or three, and it catches a defect the first one cannot: an error that is computed correctly but never displayed. ## Expect the count to shrink A coupled suite is usually redundant as well as brittle, because each internal route to the same outcome got its own test. Forty structure-coupled tests commonly collapse to fifteen behavior tests with *better* coverage — plus a handful of cases that turn out never to have been tested at all, because everyone was busy asserting wiring. The pruning rule: two tests that assert the same user-visible outcome by different internal routes are one test. ## Sequencing the work A big-bang rewrite of a test suite is risky — you lose your safety net exactly while you are changing things — and it is hard to get prioritised. A better sequence: - **Stop the bleeding first.** Agree the standard for new tests and enforce it in review, so the coupled set stops growing. - **Fix on touch.** When a test breaks during a refactor, rewrite it behaviorally instead of patching it back to green. The refactor pays for the fix, and the worst-coupled areas are exactly the ones being touched. - **Do the hot spots deliberately.** If one file breaks on every unrelated change, it earns a dedicated rewrite. One caution: while rewriting, verify the new test can actually **fail**. Rewriting a coupled assertion into a vaguer one that passes unconditionally is a real risk — break the behavior deliberately and confirm the new test goes red before you delete the old one. ## What the interviewer is checking That you can name the specific assertion categories rather than saying "they were badly written"; that your rewrite is expressed as a repeatable shape rather than case-by-case cleverness; that you understand why stubbing children caused this; and that you have a rollout plan that does not require freezing feature work. The candidate who says "I'd just update the forty tests" has missed that the update will be needed again on the next refactor.
- Why does stubbing child components in particular cause this kind of mass failure?Because a stub freezes the composition. The moment you replace a child with a fake and assert on the props it received, the test's contract becomes "this component is built from these children, wired this way" — which is precisely what a refactor rearranges. Rendering the real children instead makes the test indifferent to how the work is divided up, and it catches integration bugs between parent and child that stubs hide.
- During the rewrite, how do you make sure the new behavioral test is not simply weaker than the one it replaced?Mutate the code and confirm the test fails. Break the behavior the test claims to protect — remove the error message, leave the button enabled, drop a field from the payload — and check it goes red before deleting the old test. Rewriting a specific structural assertion into a vague behavioral one that passes unconditionally is the real hazard, and only a deliberate failure proves you avoided it.
- The team says rewriting forty tests is too expensive right now. What do you actually commit to?Stop the growth and fix on touch. New tests must follow the behavioral standard, enforced in review — that costs nothing today and caps the problem. Then, whenever a refactor breaks one of the forty, rewrite it rather than patching it green, so the work is paid for by the change that needed it. Reserve a dedicated rewrite only for the one or two files that break on every unrelated change.
saying these in an interview costs you the question
- "Just update the forty tests and move on"
- Blaming the refactor rather than the assertions
- Assuming stubbing children makes a component test more focused
- Bulk-accepting the failing snapshots as the fix
- Rewriting the whole suite at once with no safety net
- Replacing a structural assertion with one that can never fail