skip to content

Explain Kent Beck's "two hats" rule for refactoring, and describe what you should do when a refactoring reveals a genuine bug in the code you are restructuring.

level: middleimportance: should knowfreq 44%

answer

  1. Two hats: adding function vs refactoring
  2. Never both at once; swap consciously
  3. Refactor hat = no new tests, no new behavior
  4. Bug found → write it down, commit refactor first
  5. Separate commits = revertable, bisectable, reviewable

basics

~20 s

You wear one of two hats: adding functionality (new tests, new behavior) or refactoring (structure only, no behavior change, no new tests). Swap deliberately, never wear both at once. If a refactoring uncovers a bug, note it, finish or revert the structural step, then switch to the functionality hat to fix it.

solid answer

~60 s

Kent Beck's *two hats* metaphor says that when changing software you are always doing exactly one of two jobs. Wearing the **adding-function** hat, you add tests and write code to make them pass; you do not restructure existing code. Wearing the **refactoring** hat, you change structure only — no new tests, no new behavior, existing tests must pass untouched. You may swap hats many times an hour, but you should always know which one is on. The payoff is diagnostic clarity. If a test fails while refactoring, the structural move is wrong — no second explanation is possible. Mix the hats and every failure has two candidate causes, which is exactly the situation that turns a five-minute step into an hour of debugging. When refactoring exposes a real bug, do not silently fix it: a "refactoring" that changes behavior breaks the promise the label makes to reviewers and to release risk analysis. Record it, complete or revert the structural step, commit it, then switch hats and fix the bug as its own commit with a failing test first.

go deeper

for a junior

State the two activities and that you do one at a time; if you spot a bug while tidying, note it and fix it as a separate change with its own test.

for a middle

Add why: one failing test then has exactly one possible cause, and commits stay revertable and reviewable. Give the concrete bug-found sequence.

for a senior

Discuss commit hygiene, bisectability, release risk signalling, the case where tests encode the bug, and the case where unifying duplicated paths forces a deliberate behavior decision.

for a principal

Frame it as a team policy and a property of the change stream, not personal habit: mixed commits erode the trustworthiness of the "refactor" label across the org, break incident triage and automated risk scoring, and make large mechanical migrations unreviewable. Pair it with tooling — IDE/codemod-driven structural commits, PR conventions, and characterization tests as the precondition for the refactoring hat on legacy code.

## The metaphor Kent Beck, in the *Refactoring* book's opening chapter, writes that when you use refactoring you divide your time between two distinct activities: - **Adding function** — you add new capabilities. You may add tests; you get them to pass. Progress is measured in new working behavior. - **Refactoring** — you change structure only. You add no new functionality and no new tests (except tests you needed to make the refactoring safe in the first place). You only rearrange, rename, extract, inline, move. Progress is measured in clarity. "You should be aware of which hat you're wearing" — and swapping is fine and frequent, as long as it is **conscious** and the hats are never worn simultaneously. ## Why single-hat discipline is not pedantry ### 1. One failure, one explanation Under the refactoring hat, all existing tests were green before the step. If one goes red after it, the structural transformation is wrong — full stop. That single-cause property is what lets you fix it in seconds or revert with no thought. As soon as you also changed behavior in the same step, a red test could mean "my extraction was wrong" *or* "my new logic is wrong" *or* "the test encoded the old, buggy behavior" — a three-way ambiguity you now have to resolve by debugging. ### 2. Review and risk A commit labelled "refactor" tells reviewers: judge the design, the behavior is unchanged, and the blast radius on release is low. If refactorings quietly carry behavior changes, that label becomes noise, reviewers stop trusting it, and incident triage loses a valuable signal ("which of these five commits could have changed what the system does?"). ### 3. Revertability Pure structural commits can be reverted freely — reverting loses only clarity, never a feature or a fix. Mixed commits cannot: reverting a mixed commit re-introduces a bug or removes a feature, so people avoid reverting and debug forward under pressure instead. ### 4. Bisectability A history alternating clean behavior commits with clean structural commits keeps `git bisect` meaningful. Mixed commits make "the first bad commit" an unhelpful answer. ## The awkward case: refactoring reveals a bug This happens constantly — clarifying code is one of the most reliable bug-finding techniques. Extracting a conditional makes it obvious that a branch is unreachable; naming a variable properly reveals that the wrong one is used; consolidating duplicated code reveals the copies had drifted apart. The wrong move is to fix it inline "while I'm here." The disciplined sequence: 1. **Write it down.** Keep a running to-do list (Beck's habit). Do not lose the finding, but do not chase it now. 2. **Finish or revert the current structural step.** Prefer finishing if you are seconds away; revert if the step is large and the bug makes it doubtful. 3. **Commit the refactoring**, preserving the buggy behavior exactly. Yes — you deliberately preserve the bug for one commit. That is what makes the commit a refactoring. 4. **Swap hats.** Write a failing test that demonstrates the bug (red), fix it (green), refactor if the fix leaves mess. 5. **Commit the fix separately**, referencing the defect. A subtlety: sometimes the existing tests *encode* the bug — a test asserts the wrong output because it was written from the implementation. Then fixing the bug means changing that test, and the change should be obvious and justified in the fix commit, not hidden inside a structural one. ### When behavior must change to enable the refactoring Occasionally a structure move is impossible without a behavior change — e.g. two duplicated code paths differ subtly, and unifying them must pick one behavior. Handle it as two commits: first change the behavior of one path deliberately (with a test, under the function hat, and with stakeholder agreement if it is user-visible), then unify structurally. Never let "unify" silently choose a winner. ## Related discipline - **Preparatory refactoring** — Beck: "for each desired change, make the change easy (warning: this may be hard), then make the easy change." That is literally hat-swapping: refactor first (structure), then implement (function). It also makes review easy: the structural commit is large but mechanical, the behavior commit is small and interesting. - **Opportunistic refactoring / Boy Scout rule** — leave code cleaner than you found it. Still hat-disciplined: the cleanup goes in its own commit, not smuggled into the feature diff. - **The two-hat rule at team scale** — codified as PR policy: "no mixed PRs", or at minimum "mixed PRs must separate the commits". Reviewers can then read the mechanical commit quickly and spend attention on the behavioral one. ## Trade-offs and pushback - *"Two commits is bureaucratic overhead."* For a one-line rename plus a one-line fix, splitting costs seconds and buys revertability; for anything larger the split pays for itself the first time a release goes wrong. - *"I can't refactor without touching behavior because there are no tests."* Then the missing precondition is the safety net, not the hat rule — write characterization tests first (a third, preparatory activity), then refactor. - *"Big mechanical refactorings create huge diffs that hide things."* True; that is an argument for automated/IDE-driven refactorings and for clear commit messages stating the transformation applied, not for mixing.

  • Is it ever acceptable to ship a refactoring and a behavior change in one pull request?
    In practice yes, if they are separate commits within it, so a reviewer can read the mechanical change quickly and scrutinise the behavioral one, and so either can be reverted independently. What is not acceptable is interleaving them in a single commit, because that destroys the one-failure-one-cause property and makes revert an all-or-nothing choice.
  • How does preparatory refactoring relate to the two-hats rule?
    It is the rule applied in sequence: put on the refactoring hat and reshape the code until the desired feature becomes a small, obvious addition; then swap to the function hat and add it. Beck's phrasing is "make the change easy, then make the easy change." It produces a large-but-mechanical commit followed by a small-but-meaningful one, which is the easiest pair to review.
  • You are refactoring and a previously passing test goes red. What is your first hypothesis?
    That your structural transformation is not behavior-preserving — because under the refactoring hat that is the only thing that changed. Read the diff of the last small step and revert it rather than debugging forward. The secondary hypothesis is that the test was coupled to the structure you moved (mocking a deleted collaborator), which is a signal about the test, not the code.

A surgeon does not remodel the operating theatre mid-operation. Either you are treating the patient or you are rebuilding the room — doing both at once means you cannot tell which one caused the alarm.

saying these in an interview costs you the question

  • "While I was in there I also fixed…" — behavior changes smuggled into refactoring commits
  • Believing the rule forbids swapping hats, rather than forbidding wearing both at once
  • Adding new tests during a refactoring step and treating the refactoring as verified by them
  • Refactoring on a red suite and planning to sort it out at the end
  • Treating separate commits as bureaucracy rather than as revert/bisect/review capability
  • Ignoring that some tests encode the current bug, so "tests pass" does not always mean "behavior correct"

context