skip to content

Behavior Over Implementation

You will learn to draw the line between behavior and implementation detail, so a refactor that changes nothing a user sees does not break fifty tests. Interviewers pose it as 'what makes a test brittle?' and expect a concrete rule, not a slogan.

on this pageshow

questions

5

In a UI component test, what counts as an implementation detail, and what concrete rule do you use to decide whether a given assertion will survive a refactor?

level: middleimportance: must knowfreq 76%

answer

  1. two surfaces: contract versus mechanism
  2. would a user notice this change?
  3. only use handles a real caller has
  4. refactor freely, suite stays green
  5. callbacks count as outward effects

basics

~20 s

An implementation detail is anything a user cannot perceive: internal state, render counts, DOM structure, class names, which child components exist. The working rule is to assert only what a user would notice if it stopped being true.

solid answer

~60 s

I think of a component as having two surfaces. The **public surface** is what the outside world experiences: the props you hand it, what it renders, and what it calls or sends back out. The **private surface** is everything it does to get there — internal state variables, how many times it re-rendered, how deeply the DOM nests, what the wrapper's class is, which child components it happens to be built from. Behavior lives on the public surface; everything on the private surface is an implementation detail. The rule I apply per assertion is: *if this stopped being true, would a user notice?* If yes, it's behavior worth asserting. If no, I'm nailing the test to an implementation I might legitimately want to change tomorrow. The practical test of a good suite follows from that: I should be able to restructure a component's internals — rename state, split it in two, swap the styling approach — without touching a single test, and I should be able to break a user-facing outcome and watch a test go red.

go deeper

for a junior

Recall the two lists — what a user perceives versus what the component does internally — and be able to give one example from each side without hesitating.

for a middle

Explain the rule and apply it live: given an assertion, say which side of the line it falls on and why, and rewrite an internal-state assertion into a rendered-output one on the spot.

for a senior

Demonstrate the two-way check on a real suite — internals change and tests stay green, outcomes break and tests go red — and explain what a coupled suite costs a team that needs to refactor.

for a principal

Own where the line sits as a written standard, including the grey areas: what escape hatches are permitted, who verifies appearance and performance, and how you keep the exception list from growing.

## The two surfaces of a component Every UI component has an outward-facing contract and an inward-facing mechanism, and the whole discipline reduces to testing the first and ignoring the second. The **outward-facing contract** is: - **Inputs** — the props (and context, and stored data) the component is given. - **Rendered output** — the text, controls, images and structure a user actually perceives, including the accessible roles and names that decide how the control is reachable. - **Outward effects** — the callbacks it invokes, the requests it sends, the navigation it triggers. These are not visible, but they are still observed by someone outside the component. The **inward-facing mechanism** is everything else: which state variables exist and what they are called, how many times a render function ran, whether the markup is two nested `<div>`s or one, what the class attribute holds, whether the widget is built from three child components or one, in what order internal helpers were called. Behavior is the first list. Implementation detail is the second. ## The concrete rule Before writing an assertion, ask: **"if this stopped being true, would a user notice?"** - "There is a button labelled *Save*" — a user notices if it vanishes. **Behavior.** - "The error text reads *Email is required*" — a user reads it. **Behavior.** - "The submit button is disabled while the request is in flight" — a user is blocked by it. **Behavior.** - "The component re-rendered exactly twice" — no user has ever noticed a render count. **Detail.** - "`isOpen` is `true`" — the user notices the *menu appearing*, which is a different, better assertion. **Detail.** - "The root element has class `card--elevated`" — the user notices the shadow, which a class string does not prove. **Detail.** A useful second formulation for the same rule: **only touch the component through the seams a real caller has.** A caller passes props and reads what renders; a caller never reaches inside for a state variable. If your test needs a handle the real application does not have, you are testing mechanism. ## Why refactor-resilience is the metric The reason this matters is not aesthetic. Tests exist to let you change code with confidence. A suite that fails whenever internals change inverts that: it makes the refactor expensive, so the refactor does not happen, so the code rots. Worse, a developer who has watched forty tests break on a rename learns that red tests mean "the tests need updating", which is exactly the reflex you cannot afford when a red test is real. So the quality bar for an assertion is a two-way check: 1. **Restructure the internals with no user-visible change — do the tests stay green?** If they go red, they were asserting mechanism. 2. **Break a user-facing outcome — do the tests go red?** If they stay green, they were asserting mechanism that happened to survive. A suite that passes both is doing its job. Class-name assertions, state assertions and render-count assertions fail one or both. ## The tempting exceptions Two assertions look internal but are not: **Callback and request assertions.** Checking that `onSubmit` was called with the right payload, or that a `POST` went out with the right body, is not an implementation detail — it is the component's outward effect, observed at its boundary. A user does notice when saving silently does nothing. **Performance-shaped assertions.** "This must not re-render on every keystroke" is occasionally a real requirement, but a render-count assertion is a fragile way to state it and it will fight every legitimate internal change. Prefer measuring the user-perceivable consequence, or accept that this belongs to a performance check rather than a behavior test. And one genuine grey area: an element with no visible text and no meaningful role — a decorative container, a chart canvas — has no user-facing handle. A dedicated test attribute is the pragmatic escape hatch there. It is not user-perceivable, so it is a compromise, but at least it is a stable one that nobody renames during a styling pass. Treat it as the exception, not the default. ## What an interviewer is listening for The weak answer is a slogan: "test behavior, not implementation." The strong answer names the categories on both sides of the line, gives the one-sentence rule for deciding, and states the falsifiable consequence — *I can refactor the internals and the suite stays green*. That last part is what turns the slogan into an engineering standard.

  • Is asserting that a callback prop was called with certain arguments an implementation detail?
    No — it's the component's outward effect, observed at its boundary. The caller supplied that prop, so checking it was invoked correctly is testing the contract, not the mechanism. What would be an implementation detail is asserting *how* the component arrived there: which internal handler ran, in what order, or how many times state updated on the way.
  • How would you convince a team their suite is coupled to implementation, without arguing about individual tests?
    Run the experiment. Pick a refactor with zero user-visible effect — rename internal state, split a component in two, swap the styling approach — and count how many tests go red. That number is the coupling, and it's not debatable. Then run it the other way: break a real outcome and count how many tests notice. The two numbers make the case better than any review comment.
  • Does this rule mean a component test should never know anything about the DOM structure?
    It means the test shouldn't depend on structure that a user can't perceive — nesting depth, wrapper elements, generated ids. Structure that carries meaning is fair game: that an error message is associated with its input, that a control sits inside the dialog rather than behind it, that a list contains three items. Those are perceivable, and they survive refactors that only rearrange markup.

saying these in an interview costs you the question

  • "Behavior means whatever the component's code does"
  • Asserting state variables because they're easier to reach than rendered output
  • Treating render counts as a correctness assertion
  • "The tests break on every refactor, that's just testing"
  • Assuming any assertion about the DOM is automatically implementation-coupled
  • Calling callback assertions implementation details and skipping them entirely

context

open as a page

A UI component test checks that a Save button is in its primary style by asserting `expect(container.querySelector('.btn-primary')).toBeTruthy()`. Why do reviewers call that assertion brittle, and what would you assert instead?

level: juniorimportance: should knowfreq 52%

basics

~20 s

A CSS class is a styling hook, not behavior: renaming or reorganising it breaks the test even though the button still looks and works the same to a user. Assert what a user perceives — the visible label, the role, the enabled state.

open as a page

If a component test should assert only what a user can perceive, how do you test a component whose job is to invoke an `onSubmit` prop and send a request — outputs a user never sees on screen?

level: middleimportance: should knowfreq 58%

basics

~20 s

A component's observable contract is wider than its pixels: it includes the callbacks it invokes and the requests it sends. Assert those at the component's outer boundary, driven by a real user interaction — never by reaching into internal handlers.

open as a page

A large form component is split into three smaller components, with no change to what renders on screen or how it behaves — yet about forty of its tests now fail. What kinds of assertions cause that, and how would you rewrite the suite so the next refactor costs nothing?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Those tests were asserting structure rather than outcome: internal state, child components and the props passed to them, DOM nesting, class names. Rewrite each around a user-visible outcome of a real interaction, so the same test passes before and after the split.

open as a page

You are setting the component-testing standard for a large frontend. Behavior-only assertions survive refactors but produce vaguer failure messages and slower tests than assertions on internals. How do you weigh that, and where would you allow exceptions?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Refactor freedom is worth more than diagnostic precision, so user-visible outcomes are the default. Allow narrow, reviewed exceptions where a state genuinely cannot be reached through the UI, and require each exception to name what it is protecting.

open as a page