skip to content

Your refactoring left every test green. Why is that not proof that behaviour is unchanged?

level: seniorimportance: should knowfreq 45%

answer

  1. Evidence, not proof
  2. Two gaps: reach and aspect
  3. The tests stopped before the boundary
  4. Nothing asserted on timing or payload shape
  5. Check the region's coverage before you start

basics

~20 s

A suite pins only what it exercises and only what it asserts on. Paths it never reaches, wire-level shapes it never inspects, and timing, ordering or output it never measures can all drift while every test passes.

solid answer

~50 s

A green suite is evidence, not proof, and its strength is exactly the strength of its coverage and its assertions. Two gaps do the damage. First, reach: any branch, input class or integration path the tests never execute is unguarded, so a restructuring that changes it is invisible. Second, aspect: even on covered paths the tests assert on a slice — usually a returned value — and say nothing about the serialized shape crossing a boundary, iteration order, error text, emitted telemetry, allocation, latency or thread-safety, all of which restructuring routinely disturbs. So before restructuring a region I check what actually covers it and add behaviour-level tests where it is thin, prefer tool-assisted moves for anything that crosses a boundary, keep steps small, and watch the non-functional signals — a green suite plus an unchanged wire contract plus a stable latency profile is a much stronger claim than green alone.

code

pseudocode · 7 lines
pseudocode
# passes before and after the rename
view = cabin_view(flight_id = "SM-4417")
assert view.seats[0].price == 118.50

# only this one can fail on schema drift
payload = serialize(cabin_view(flight_id = "SM-4417"))
assert field_names(payload.seats[0]) == ["row", "class", "available", "price"]

go deeper

for a junior

Take away the core idea: tests only check the paths they run and the things they assert on, so a green run means no test objected, not that nothing changed.

for a middle

Be able to separate the two gaps — code the suite never reaches, and behaviour the suite reaches but never asserts on — and give a concrete example of each.

for a senior

Demonstrate the production instinct: check what guards the region before you touch it, pin boundary-observable behaviour, watch call counts and latency, and state plainly what the change was not verified against.

for a principal

Own the risk framing: decide which boundaries deserve permanent contract-level and budget-level checks so that restructuring anywhere behind them is cheap, and where the team accepts unverified change deliberately.

## Two independent gaps Candidates usually name one and stop. There are two, and they fail differently. **Reach.** A test can only object to code it runs. Whatever the suite does not execute — a rarely-taken branch, a null-ish edge case, a second locale, a retry path — is simply not in the conversation. Restructure it and the run is green because nothing looked. This is the gap a coverage report speaks to, and it is worth reading *before* a restructuring as an input ("is this region actually guarded?") rather than chased afterwards as a target. **Aspect.** Even where the suite has full reach, it asserts on a chosen slice of behaviour. A test on a returned object says nothing about the shape of the payload that object becomes at a boundary, the order in which a collection comes back, the exact wording of an error a client parses, the events or metrics emitted, the number of round trips made downstream, how long any of it takes, or whether it is still safe under concurrent calls. Consumers depend on all of those, and restructuring is precisely the activity that disturbs them: you moved a computation, merged two loops, dissolved a cache, or changed which type carries a field. ## Worked example An airline seat-map service returns a cabin view for a flight — a few hundred seats with class, availability and price. The suite has 47 tests over the region, all asserting on the returned in-memory view. Two restructurings, both green, both harmful. **A schema-drift mismatch.** Extracting the pricing part gave the price field a clearer internal name. The payload written at the service boundary is produced by a mapper that derives its field names from the code's own names, so the outgoing document quietly renamed a field. Every one of the 47 tests still passes, because every one of them stops at the in-memory object and never inspects what is written on the wire. The downstream seat-selection client, which reads that field by name, starts showing blank prices. The suite never had reach *past the boundary*, so the aspect that broke was never observed. **A budget breach.** Inlining a lookup that had been hoisted out of a loop was, in behaviour terms, exactly equivalent: same values, same results, same assertions green. It also turned one availability lookup into one lookup per seat row. On a 340-seat cabin the endpoint's 92nd-percentile response time went from 118 ms to 402 ms, against a budget of 214 ms. No test measured time, so no test objected. Neither of these is a failure of discipline in the small — the moves were small, the suite ran after each one. They are failures of the net's shape. ## How to strengthen the claim - **Read the coverage of the region before you start**, and if the branches you are about to disturb are unguarded, add behaviour-level tests that pin them first, while the code still does the old thing. Adding them afterwards pins whatever you just produced. - **Test at the boundary that consumers actually observe** when a restructuring can reach it. If a downstream party reads a payload, something in your build should assert on that payload's shape, not only on the object behind it. A contract-style check at the boundary catches the entire schema-drift family that in-memory assertions structurally cannot. - **Prefer tool-assisted moves**, and remember their guarantee stops where names are resolved outside the analysed source — which is exactly where boundary drift lives. - **Keep non-functional signals in the loop.** A latency or allocation profile before and after, or a pipeline check against a budget, turns an invisible regression into a red result. You do not need this on every move; you need it on moves that change how often work is done. - **Say what you did not verify.** The honest summary of a refactoring is "behaviour is unchanged on everything the suite exercises and asserts on, and here is what that excludes". That sentence is what separates a senior answer from a confident one. ## The framing that scores Green means "no counter-evidence from the checks I ran". It never means "equivalent". The engineering question is not whether the net has holes — it always does — but whether you know where they are before you start cutting, and whether the ones over the code you are about to restructure are small enough to accept.

  • Does raising the coverage percentage before a restructuring fix this?
    It helps with reach and does nothing for aspect. Executing a line proves only that it ran; a test that touches the boundary path but asserts on an in-memory object still cannot see a payload rename. Use coverage to find unguarded branches to pin, then choose assertions at the level the consumer actually observes — that second choice is where the real protection comes from.
  • How do you catch a restructuring that preserves results but multiplies the work done?
    Assert on the work, not only the result: a check on downstream call counts or query counts at the boundary catches the loop-hoisting family cheaply and deterministically. Keep a latency or throughput check against an explicit budget in the pipeline for the endpoints that have one, and compare profiles before and after any move that changes how often a computation runs.
  • The region you must restructure genuinely cannot be covered cheaply. What do you do?
    Shrink the exposure rather than pretend it is covered: keep steps very small, prefer mechanical tool-assisted moves over hand edits, restructure behind a boundary you can monitor, and stage the rollout so a regression is observable and reversible quickly. Then record explicitly what the change was not verified against, so the risk is a decision rather than an assumption.

A metal detector that finds nothing tells you there is no metal where you swept it, at the depth it reaches — not that the field is empty.

saying these in an interview costs you the question

  • Treats a green suite as proof of equivalence
  • Names coverage gaps but never assertion gaps
  • Ignores payload shape, ordering and error text
  • Assumes performance is not behaviour
  • Adds the missing tests after restructuring, not before
  • Chases a coverage percentage instead of pinning behaviour

context