Why run the test suite after each small refactoring move instead of once at the end?
answer
- Green states are checkpoints, not milestones
- Which move broke it?
- Undo is cheaper than debugging forward
- Batch size sets the search space
- Fast tier per move, full tier per commit
basics
~20 sBecause a green run after every move tells you exactly which move changed behaviour. Run once at the end and a red suite only says something in the last hour was wrong, so you debug instead of undoing.
solid answer
~50 sRefactoring is a sequence of small structure-only edits, and the suite is the only thing that says behaviour survived each one. Running after every move keeps the distance between two known-green states down to a single edit, so a red run points straight at the culprit and the cheapest fix is to undo that one move and retry it in smaller steps. Batch ten moves and the same red run gives you a ten-move search space plus interacting edits, which turns a mechanical exercise into a debugging session — exactly what the safety net was supposed to prevent. When the whole suite is too slow to run every minute, run a fast subset that covers the region you are touching after each move and the whole suite before you commit, and treat any gap in that fast subset as a risk you accepted knowingly.
code
pseudocode · 13 linesstate = last_green_commit
for move in [extract, rename, inline, delete_dead_branch]:
apply(move)
result = run(fast_tests_covering(region))
if result == RED:
undo(move) # not: debug forward
split(move) # retry in smaller pieces
continue
state = checkpoint() # cheap, local, throwaway
run(full_fast_suite())
commit(state)go deeper
Remember the loop: one small move, run the covering tests, checkpoint, repeat. If a run goes red, undo the last move rather than trying to fix it forward.
Explain why batch size drives the cost: the search space for a red run is the number of edits since the last green state, and later edits sit on top of earlier ones.
Show how you keep the rhythm on a slow suite — tiering fast and slow tests, checking that the region has fast coverage before you start, and naming what the fast tier leaves unprotected.
Own the economics: suite speed is what sets the team's affordable step size, so investment in a trustworthy fast tier is investment in everyone's ability to restructure safely.
## The rhythm The refactor step of the cycle is not one edit; it is a chain of them. Pull a fragment out, give the new unit a name, move a call, drop a now-unused parameter. Each link changes structure and is supposed to leave observable behaviour identical. The suite is the only mechanised opinion you have about whether that is actually true, and its verdict is only useful in proportion to how narrowly it can be aimed. Running after every move keeps the interval between two known-green states at exactly one edit. That gives you three things: a **precise locator** (the move you just made is the only suspect), a **cheap undo** (throwing away one edit costs nothing), and a **decision rule that needs no thought** (red means undo, not debug). Running once at the end gives up all three at once. The suite still tells you something is wrong, but it cannot tell you which of the twelve edits did it, and because later edits sit on top of earlier ones, unpicking them is no longer mechanical. ## Why undo beats debug here This is the part candidates most often miss. A red test during a *behaviour-preserving* edit is not a puzzle to solve; it is evidence that the edit was not behaviour-preserving. The fastest correct response is to return to the last green state and take a smaller step — perhaps splitting the move into two, perhaps discovering on the second attempt that the fragment you wanted to extract was not as independent as it looked. Debugging forward from a red state is how a refactoring quietly turns into a rewrite: you start changing production logic to satisfy the suite, and the property you were trying to preserve is gone. That rule only pays off if the last green state is genuinely reachable. In practice that means committing (or at least stashing a checkpoint) at green states you care about, so "undo" is a command rather than an act of memory. ## Worked example An airline seat-map service builds a cabin view: rows, seat classes, blocked seats, and a per-seat price. The build method has grown to a few hundred lines and you want to pull the pricing part into its own unit. Eight moves: introduce a variable, extract the price fragment, move a lookup up, inline a redundant temporary, rename two parameters, delete a now-dead branch, and narrow a visibility. Run the covering tests after each and move six goes red on the two tests that check blocked-seat pricing: deleting the branch you thought was dead was not dead for seats marked unavailable after check-in opened. You undo move six, keep the five green moves, and either write a test that names the branch's real purpose or leave the branch alone. Elapsed cost: under a minute. Batch the eight moves and you get the same two red tests over an eight-edit diff in which the pricing code no longer lives where it used to. Now you are bisecting your own uncommitted work, and the tempting shortcut — loosening the two assertions until everything is green — destroys the evidence that anything happened at all. ## When the suite is too slow A suite that takes minutes cannot be run every ninety seconds, and pretending otherwise just means it gets run never. The standard compromise is tiered: after each move, run the fast, in-process tests that cover the region you are editing; before each commit, run the full fast tier; leave the slow, environment-dependent tiers to the pipeline. Two consequences follow. First, the fast tier's coverage of the region you are about to restructure is a **precondition** for refactoring there, worth checking before you start. Second, anything the fast tier does not exercise is not protected by this rhythm, and you should know which parts of the change those are rather than discover it later. ## What good looks like Small moves, a run between each, a checkpoint at each green state, and an undo reflex rather than a debug reflex. Tool-assisted moves do not exempt you from the run: the tool's guarantee covers the references it can see, and the run is what covers the rest.
- The full suite takes eleven minutes. How do you keep this rhythm without waiting eleven minutes per move?Tier it. Run the fast in-process tests covering the region after every move, the whole fast tier before each commit, and leave slow environment-dependent tests to the pipeline. The trade is explicit: whatever the fast tier does not exercise is unprotected during the moves, so check that the region has fast coverage before you start restructuring it.
- Should each green step be committed, or only the finished refactoring?Commit at green states while you work — they are the undo points that make the rhythm cheap, and a commit is a more reliable checkpoint than remembering what you changed. Tidy the history afterwards if the team wants one commit per logical change; that is a presentation decision, and it should never be the reason to work without checkpoints.
- A move goes red and you are certain the test is simply wrong. What now?Undo the move first, so production code is back at green. Then treat the test as its own piece of work: fix or replace it with production code held still, confirm it passes, and only then reattempt the move. Fixing the test while the restructuring is half-applied means neither side is checking the other.
It is the difference between saving a document every paragraph and saving it once at midnight: both keep the text, but only one lets you lose a single paragraph.
saying these in an interview costs you the question
- Claims running once at the end is equivalent
- Debugs forward from a red instead of undoing the move
- Assumes a red during refactoring means the test is stale
- Adjusts assertions to get back to green
- Thinks tool-assisted moves need no test run
- Refactors for an hour with no checkpoint to return to