skip to content

Why do unit tests that assert on a class's internals turn red during a behaviour-preserving restructuring?

level: middleimportance: must knowfreq 68%

answer

  1. Ask what the assertion is really describing
  2. How it works versus what it does
  3. Behaviour unchanged, suite still red
  4. Assert through the public surface
  5. It became a change detector

basics

~20 s

Because those assertions describe how the code is built, not what it does. Rename an internal helper or replace a private data structure and the observable behaviour is unchanged, but the assertions no longer match, so the suite fails.

solid answer

~50 s

An assertion picks a subject. If it names a private field, an internal collection, a helper's name or the number of intermediate steps, it is describing the implementation, and a restructuring is allowed to change all of that while preserving behaviour — so the suite goes red without any defect existing. The practical test when a suite fails is: could a caller outside this unit observe a difference? If not, the test was structural and should be rewritten at the public surface rather than re-tuned to the new internals. Over-specification costs twice: the suite fires on safe changes, so the team stops restructuring, and it crowds out the assertion that states the rule the code exists to uphold, so real defects still slip through. Assert the observable rule; couple to an internal only when that internal genuinely is the contract, and name the test so the coupling is visibly deliberate.

code

pseudocode · 7 lines
pseudocode
test "fee split across signers":
    splitter = FeeSplitter(rounding = "half_up")
    splitter.allocate(total = 12.47, shares = 3)

    assert splitter.remainderBuffer == [0.01, 0.00, 0.00]
    assert splitter.stepsTaken == 3
    assert splitter.roundingMode == "half_up"

go deeper

for a junior

Recall the core distinction: a test should describe what the code does for a caller, not how it is built inside. Be ready to say why renaming an internal helper should never break a test.

for a middle

Explain the mechanism. Show which assertions are structural, why a behaviour-preserving change is allowed to break them, and how you would restate the same intent as an assertion on returned values or observable state.

for a senior

Demonstrate the diagnosis on a real suite: deciding in minutes whether a red run is a regression or over-specification, and leading the rewrite in green batches instead of re-tuning assertions to match new internals.

for a principal

Own the consequence at the codebase scale: a change-detector suite prices restructuring out, so the design ossifies. Be ready to argue how you measure that, and how you set an assertion standard a team can apply without you.

### The two things a test can describe Every assertion you write picks a side. It either describes **what the code does** — a value handed back to a caller, a state a caller can observe through the same public surface, a message that crosses a boundary the system is contractually required to cross — or it describes **how the code is built**: the name of an internal helper, the contents of a private working collection, the number of intermediate objects created, the exact order in which internal steps ran, the shape of a data structure that never leaves the object. The second kind of assertion is not "stronger". It is bound to a different subject. A restructuring that preserves behaviour is, by definition, allowed to change everything the second kind of assertion is looking at. So the suite goes red while nothing a user or a calling system could observe has changed at all. ### Why test-driven work falls into this so easily The pitfall is specific to writing tests alongside code. While you drive an implementation you are holding its shape in your head — the intermediate list, the helper you just extracted, the flag you set on the way through. That shape is the most available thing to assert on, and asserting on it feels like thoroughness: more assertions, more lines exercised, more confidence. Test-after work falls into the same trap for a different reason — the code already exists, so the easiest cases to write are the ones that read its internals back. ### A worked example A four-person team builds a document e-signing flow. An envelope carries a per-envelope fee that must be split across the signers so the parts sum exactly to the fee. Driving it out, they write an allocator that walks the signers, keeps a remainder buffer, and pushes leftover units onto the last part. The tests they write alongside it look like this: ``` assert splitter.remainderBuffer == [0.01, 0.00, 0.00] assert splitter.stepsTaken == 3 assert splitter.roundingMode == "half_up" ``` Six weeks later someone replaces the buffer with a running-remainder allocation that produces identical outputs for every input. Thirty-seven tests turn red. Nobody can say, from the failures alone, whether behaviour broke — so the change is reverted, and the team quietly learns that this area is not safe to touch. The second half of the story is worse. Because the tests were watching the buffer rather than the result, none of them ever asserted the property that actually mattered: **the parts must sum to the fee.** A currency-rounding drift of two hundredths per envelope survived the whole suite and was found in reconciliation after 1,912 envelopes had been processed. Over-specified tests are not just fragile; they crowd out the assertion that would have caught the defect, because the writer feels already covered. ### The diagnostic When the suite goes red after a change, ask one question: **could any caller outside this unit observe a difference?** If yes, the test is doing its job and you have a real regression. If no, the test was structural, and the correct fix is to rewrite the test, not to re-tune it to the new internals. Re-tuning is the move that keeps the pitfall alive — the team pays the maintenance cost every time and never gets the safety in exchange. A second diagnostic runs at write time: before adding an assertion, say out loud what it would mean if it failed. "The fee split no longer sums to the fee" is a sentence a product owner understands. "The remainder buffer holds different values" is not a defect statement at all. ### What to assert instead Drive from the caller's side of the boundary. State the rule the code exists to uphold — the parts sum to the total, no part differs from another by more than the smallest currency unit, an envelope with no signers is rejected — and assert exactly those. Where an internal genuinely *is* the contract (an invariant a data structure exists to maintain, a documented ordering guarantee) it is fair to assert it, but say so in the test name so the next reader knows the coupling is deliberate rather than accidental. ### The compounding cost A suite of structural assertions inverts the value of the safety net. It stops being the thing that lets you restructure freely and becomes the thing that prices restructuring out — a change detector that fires on every edit and reports nothing about correctness. Teams respond rationally: they stop restructuring, the design ossifies, and the same suite that was supposed to enable change is the reason change stopped. Recovering means rewriting assertions at the public surface in small batches while the suite is green, deleting the structural ones outright rather than keeping both, and resisting the urge to preserve an assertion just because someone once wrote it.

  • When a suite goes red, how do you tell an over-specified test from a genuine regression?
    Ask whether a caller outside the unit could observe any difference: a different returned value, a different observable state, a different message crossing a boundary. If the answer is no, the change preserved behaviour and the assertion was structural. Reverting the change and reading the diff usually settles it in seconds. The distinction matters because the two failures have opposite fixes: fix the code, or delete and rewrite the test.
  • Is asserting on an internal ever legitimate?
    Yes, when the internal genuinely is the contract — an invariant a data structure exists to maintain, or an ordering guarantee the unit publishes and callers rely on. Then it is behaviour that merely looks internal. Keep it rare, and name the test so the coupling reads as deliberate. The anti-pattern is asserting on incidental structure that nobody promised and nobody depends on.
  • You inherit a suite full of structural assertions. How do you get out?
    Rewrite in small batches while the suite is green, one area at a time. For each unit, state the rule it exists to uphold, write that assertion at the public surface, confirm it fails when you break the rule, then delete the structural assertions rather than keeping both. Resist preserving an assertion just because it exists — a structural assertion you keep is a restructuring you have priced out.

A test bound to internals is a building inspection that measures the scaffolding: take the scaffolding down and the report fails, though nothing about the building changed.

saying these in an interview costs you the question

  • Says more assertions always makes a stronger test
  • Asserts on private fields and calls that thorough coverage
  • Blames the restructuring when only structural assertions broke
  • Re-tunes the assertion to the new internals and moves on
  • Believes a red suite after a rename proves behaviour changed

context