skip to content

While refactoring production code under a green suite, may you edit the tests in the same step?

level: middleimportance: must knowfreq 58%

answer

  1. One side must stay still
  2. The suite is the oracle here
  3. Two edits can cancel each other out
  4. Separate move, production code frozen
  5. A refactoring commit changes no assertion

basics

~20 s

No. While structure changes, the tests are the fixed oracle that says behaviour did not. Change both at once and neither is checking the other, so any drift you introduce can be absorbed by the edited assertion.

solid answer

~50 s

The safety net works because one side is held still. Restructuring production code with the tests untouched means every red run is real evidence; editing the assertions in the same step means a behaviour change and a matching assertion change cancel out and the suite goes green on a defect. So the rule is: production code moves while test code is frozen, and when test code genuinely must change — it named something you renamed, or it asserted on a structure that no longer exists — you stop the refactoring, make that a separate change with production code held still, confirm the suite is green again, and only then resume. Two exceptions are not really exceptions: a mechanical rename propagated by a tool into test files changes no assertion, and adding a brand-new test is a different piece of work, not part of the restructuring.

code

pseudocode · 13 lines
pseudocode
# WRONG: both sides move in one step
extract(pricing_part)
rewrite_assertion(test_counts_helper_calls)   # oracle edited mid-move
run()  # green, and it proves nothing

# RIGHT: alternate which side is frozen
undo(extract)
# step A - tests are the variable, production code frozen
rewrite(test_counts_helper_calls -> asserts_on_returned_view)
run(); commit("test: assert on behaviour, not call structure")
# step B - production code is the variable, tests frozen
extract(pricing_part)
run(); commit("refactor: split pricing out of cabin view builder")

go deeper

for a junior

Learn the rule and the reason: while you are changing structure, leave the expectations alone, because a test you edit to agree with you has stopped checking anything.

for a middle

Be ready to walk through the alternation — freeze tests while production code moves, freeze production code while a test is repaired — and to name which case a given red run falls into.

for a senior

Show how you keep this visible in review: a refactoring change set that alters assertions is a behaviour change mislabelled, and you should be able to say what you do when you find one.

for a principal

Own the consequence for the team: a suite people feel free to edit whenever it disagrees with them stops being an oracle, and restructuring across the codebase then carries risk nobody can quantify.

## Why one side has to be frozen A refactoring makes a claim: structure changed, observable behaviour did not. The suite is what tests that claim, and a test only has authority over code it was not written against. The moment you edit an assertion in the same step as the production code it guards, the two changes can cancel: you alter behaviour by accident, alter the expectation to match, and the run comes back green. Nothing lied — you simply removed the observer and the observed at the same time. So the discipline is a simple alternation. During the refactor step, **production code is the variable and test code is the constant**. If the test code has to move, invert it: test code becomes the variable and production code the constant, and you verify the tests still pass against the unchanged implementation before touching structure again. ## The three cases people conflate **1. A test broke because the refactoring changed behaviour.** This is the net doing its job. Undo the move; do not touch the test. **2. A test broke because it was coupled to structure rather than behaviour** — it named a helper you renamed, counted calls to a collaborator you merged, or reached into an intermediate object you dissolved. Nothing about the system's observable behaviour changed, but the test cannot compile or cannot pass. This is where discipline earns its keep. The tempting move is to fix the assertion in place, mid-refactoring, which puts you in the cancelling state above. The disciplined move is: return to green, change the test alone so that it expresses the behaviour rather than the shape, confirm it passes against the untouched implementation, commit that, then resume the restructuring. **3. A test needs to exist that does not yet.** Adding a test is a different hat entirely. If you discover mid-refactor that the region is thinly covered, stop, add the test against current behaviour, watch it pass, and then continue. Writing that test after the restructuring means you are pinning whatever the code does now — including any drift you just introduced. ## The rename exception, and why it is not one A tool-driven rename that propagates a new name into test files does touch test source. It is still safe, because it changes no expectation: the same call is made, the same value is compared, only the identifier differs, and the change is applied mechanically rather than judged case by case. The distinguishing question is always *did any assertion's meaning change?* If a human decided what an expected value should now be, that is an edit to the oracle. If a tool substituted a token everywhere at once, it is not. ## Worked example An airline seat-map service exposes `cabinView(flightId)`. You want to split the sprawling builder into an availability part and a pricing part. Forty-seven tests cover the region; forty-four assert on the returned cabin view, three assert that the builder called an internal `resolveBlockedSeats` helper a specific number of times. The extraction is behaviour-preserving, and the forty-four stay green. The three break — not because behaviour drifted, but because they were written against the call structure. If you rewrite those three assertions while the extraction is half-applied, you no longer know whether the forty-four went green because behaviour survived or because you were lucky. Instead: undo the extraction, rewrite the three tests to assert on blocked seats appearing in the returned view, run them against the original implementation, commit, and only then redo the extraction. Now the extraction is guarded by forty-seven behaviour-level tests, and the fact that they all stay green means something. ## The smell of the rule being broken A change set where production and test files move together in a restructuring commit is the visible symptom, and the review question that catches it is simply "which assertions changed, and why?" A refactoring commit should be able to answer "none". If the answer is a list, the commit is a behaviour change wearing a refactoring's label, and its green suite proves nothing about the behaviour that was preserved. The deeper point is that discipline here is not ceremony: it is what makes the green run *informative*. A suite you are willing to edit whenever it disagrees with you is not an oracle; it is a transcript of your intentions.

  • How would you spot this rule being broken in someone else's change set?
    Ask which assertions changed in a commit labelled as a refactoring. The honest answer is none. If expected values, counts or comparisons moved alongside the production edits, the commit changed behaviour and its own oracle together, and the green run says nothing about preservation. Tool-propagated identifier renames in test files are the one benign case, because no expectation's meaning changed.
  • Mid-refactoring you find the region has almost no coverage. Do you press on?
    No. Stop, put production code back at green, add tests that pin the region's current observable behaviour, watch them pass, and commit them. Then restructure. Adding the tests afterwards pins whatever the code does after the change, including drift you introduced, which is precisely the evidence you needed and no longer have.
  • Is deleting a test ever acceptable during a refactoring?
    Not during it. A test that is genuinely redundant or was pinning a structure that no longer exists can be removed, but as its own change, with production code held still and the reason stated. Deleting a test in the same step as the restructuring it is failing is indistinguishable from suppressing the evidence, whatever the intent.

You cannot weigh yourself while adjusting the scale; one of the two has to be the reference for the reading to mean anything.

saying these in an interview costs you the question

  • Adjusts assertions until the suite goes green again
  • Deletes a failing test as obsolete mid-refactoring
  • Ships production and assertion changes in one refactoring commit
  • Assumes a red test during refactoring is always stale
  • Believes test code may never be touched at all
  • Writes the missing tests after restructuring rather than before

context