What does it mean that refactoring is "behavior-preserving", and which kinds of code changes are therefore NOT refactorings?
answer
- Structure changes, behavior doesn't
- Observable = returns, errors, side effects, contracts
- Tests pass untouched
- Feature/bugfix ≠ refactoring — separate commits
- Small step → run tests → commit
basics
~20 sRefactoring changes how code is structured, never what it does. Callers see the same observable results before and after. Adding a feature, fixing a bug, or changing output is not refactoring — it is a behavior change, done separately.
solid answer
~50 sRefactoring is restructuring existing code so its externally observable behavior stays the same while its internal design improves (clearer names, smaller units, less duplication, better boundaries). "Observable behavior" means what clients can detect: return values, thrown errors, persisted state, messages emitted, and any performance or ordering guarantee people actually depend on. Renaming a variable, extracting a function, inlining a redundant indirection, or moving a method to the class that owns its data are refactorings. Adding a feature, fixing a defect, changing an error message a client parses, tightening validation, or altering a public contract are not — they are behavior changes and belong in separate commits. The distinction matters for review and risk: a pure refactoring can be judged on design alone and, if the test suite is trusted, is verified by the suite passing unchanged. If your tests had to change to make a "refactoring" pass, you probably changed behavior.
code
pseudocode · 16 lines// BEFORE
function total(order) {
var t = 0
for (line in order.lines) { t = t + line.qty * line.price }
if (order.customer.isVip) { t = t * 0.9 }
return t
}
// REFACTORING (behavior preserved): extract + name the concepts
function total(order) {
return applyVipDiscount(subtotal(order), order.customer)
}
function subtotal(order) { return sum(order.lines, line -> line.qty * line.price) }
function applyVipDiscount(amount, customer) { return customer.isVip ? amount * 0.9 : amount }
// NOT a refactoring: 0.9 -> 0.85 changes results for VIP customers.go deeper
Say refactoring improves the code's structure without changing what it does, give one concrete example (extract a function, rename a variable), and state that adding features or fixing bugs are separate kinds of change.
Define observable behavior concretely — return values, errors, side effects, published contracts — and note that a true refactoring leaves the existing tests passing untouched, so behavior work goes in its own commit.
Discuss verification: IDE-automated refactorings, a trusted suite, characterization tests when none exists. Acknowledge the grey zone where log wording, iteration order or latency have become de-facto contracts, and explain why commit separation buys revertability, bisectability and readable release risk.
Frame it as a property of the change stream and a team promise: the "refactor" label carries a risk contract that incident triage and review allocation depend on, so mixing behavior into it has organisational cost. Cover how "observable" is defined by real consumers rather than access modifiers, how to establish equivalence evidence at scale (parallel run, contract tests, staged rollout), and why small-step discipline must survive scale-up rather than being replaced by long-lived branches.
## The definition **Refactoring** (Martin Fowler's term, popularized in his book *Refactoring*) is a change to the internal structure of software that makes it easier to understand and cheaper to modify **without changing its observable behavior**. Fowler uses the word both as a noun (a single named transformation, e.g. *Extract Function*) and as a verb (the activity of applying a series of them). The key phrase is *observable behavior*. It does **not** mean "identical machine instructions" or "identical internal structure" — the whole point is that internals change. It means: **nobody outside the code being changed can tell the difference by legitimate means.** ## What counts as "observable" For a unit of code, observable behavior is everything a client can legitimately depend on: - **Return values** for the same inputs. - **Errors / exceptions**: same type and same conditions that trigger them. - **Side effects**: rows written to a database, files created, messages published, emails sent. - **Contracts stated publicly**: API shapes, wire formats, log lines that a monitoring system parses. - Sometimes **non-functional characteristics** people actually rely on: ordering guarantees, thread-safety, latency budgets, memory bounds. Everything else is fair game to change: variable and function names, the number and size of functions, which class holds which method, control-flow shape, use of a loop vs. a mapping construct, internal data structures, private helpers. The grey zone is real. Log wording is usually internal — until someone builds an alert on it. Iteration order of a collection is usually internal — until a caller has silently come to depend on it. A senior answer acknowledges this: "observable" is defined by the actual contract and the actual consumers, not by a language keyword like `public`. ## What is NOT refactoring | Change | Why it isn't a refactoring | |---|---| | Adding a feature | New behavior appears | | Fixing a bug | Output for some input deliberately changes | | Tightening validation | Inputs that used to be accepted now fail | | Changing an HTTP status code or error code | Clients can observe it | | Optimizing in a way that alters results (e.g. changing rounding) | Output changes | | Deleting genuinely dead code | Borderline — usually treated as safe cleanup; if truly unreachable, no observable behavior changes | | Adding tests | Not a refactoring of production code, but it is a normal, encouraged preparatory step | A **pure performance optimization** that keeps the same results is an interesting case: it preserves functional behavior but deliberately changes a non-functional property. Fowler treats optimization and refactoring as different activities with different goals — refactoring optimizes for *understandability*, optimization for *speed* — even though both are behavior-preserving in the functional sense. ## Why the distinction has practical value 1. **Risk isolation.** If a deployment breaks and the release contained one commit labelled "refactor" and one labelled "add discount rule", you know which to suspect. Mixing them destroys that signal. 2. **Review economics.** A reviewer of a pure refactoring asks only "is the new structure better, and is it really equivalent?" — they don't have to re-derive the requirements. 3. **Test semantics.** For a true refactoring, existing tests should pass **untouched**. Tests are the safety net; if you edit the net at the same time as walking the wire, the net proves nothing. (Exception: tests that assert on private structure — e.g. mocks of a class you just deleted — legitimately change. That is itself a signal your tests were coupled to implementation rather than behavior.) 4. **Communication.** Saying "this PR is a refactoring" is a promise to your team about the risk profile. Breaking that promise erodes trust in the label. ## The verification question How do you *know* behavior is preserved? Options, from strongest to weakest: - **Automated refactorings in an IDE** (rename, extract method, move) — mechanically derived, the strongest guarantee available, though not perfect (reflection, dynamic dispatch by string name, and text-based templates can defeat them). - **A trusted test suite** run before and after. - **Characterization / golden-master tests** written specifically to pin current behavior when no tests exist. - **Type checking and compiler feedback** — catches shape errors, not semantic ones. - **Careful reading** — weakest; fine for a rename, not for restructuring conditionals. ## Small-step discipline Behavior preservation is much easier to guarantee if each step is tiny and independently verified. The classic rhythm is: make one small structural change, run the tests, commit; repeat. If a step goes wrong, you revert seconds of work rather than days. A long-running "refactoring branch" that touches hundreds of files and cannot be compiled for two days has lost the property that makes refactoring safe.
- Your existing tests had to be edited for your "refactoring" to pass. What does that tell you?Either you changed behavior (so it isn't a refactoring), or the tests were coupled to implementation details — asserting on private methods, internal call sequences, or mocks of collaborators you just removed. Both are worth naming explicitly in the PR: the first means splitting the commit, the second means the tests were testing structure rather than behavior.
- Is deleting unused code a refactoring?Usually treated as one, because truly unreachable code contributes no observable behavior. The risk is proving unreachability — reflection, dynamic dispatch, feature flags, scheduled jobs, or external callers of a published library can make "dead" code live. Verify with usage search plus runtime telemetry before deleting, and keep the deletion in its own commit so it is trivially revertible.
- Does a pure performance optimization count as refactoring?It preserves functional behavior but its goal is a non-functional property, so Fowler classifies it as a separate activity. Practically: keep it in its own commit with a benchmark as evidence, because the risk profile and the review criteria differ from a readability refactoring.
Rearranging the furniture and rewiring the cupboards in a kitchen: from outside, the same meals come out at the same times. If the menu changes, that's not rearranging — that's a new restaurant.
saying these in an interview costs you the question
- "Refactoring" used as a synonym for any large code change, including rewrites and feature work
- Bundling a bugfix into a "refactoring" commit, so the risk profile of the release becomes unreadable
- Editing tests until they pass and calling the result behavior-preserving
- Assuming only `public` members are observable — logs, ordering, timing and side effects can also be depended on
- Long-lived refactoring branches that cannot compile for days, losing the small-step safety property
- Claiming a change is behavior-preserving with no test suite and no characterization tests to back it