You must restructure an important module that has no automated tests. How do you make that refactoring safe? Explain characterization tests and seams.
answer
- Legacy = code without tests (Feathers)
- Characterization test pins actual behavior, bugs included
- Assert wrong value → read actual → paste it in
- Seam = change behavior without editing there; enabling point
- Golden master for wide outputs; parallel run for real traffic
basics
~20 sDon't refactor blind. First pin the current behavior with characterization tests — tests that record what the code actually does today, bugs included. To get the code under test, introduce seams: small, low-risk points where you can substitute a dependency. Then refactor in small steps, running the tests each time.
solid answer
~50 sRefactoring without a safety net is just editing and hoping. The standard approach (Michael Feathers, *Working Effectively with Legacy Code*) is: 1. **Characterization tests** — write tests that document existing behavior rather than intended behavior. Call the code, see what it returns, and assert exactly that, including quirks and bugs. They are not correctness tests; they are change detectors. 2. **Seams** — a seam is a place where you can change behavior without editing the code there: an injectable parameter, an overridable method, a link/config-level substitution. Introduce the minimum seam needed to instantiate the class or intercept a hard dependency (clock, network, database, global singleton), using only the safest possible mechanical edits. 3. **Golden-master / approval testing** when outputs are large or complex: capture the current output for many inputs, then assert byte equality after each refactoring step. 4. Refactor in small steps under green, prefer automated IDE refactorings, commit often. If a characterization test fails later, you consciously decide: bug fixed, or regression introduced.
code
pseudocode · 23 lines// 1. No seam: cannot test — real clock + real gateway inside
class Invoicer {
function issue(order) {
var now = SystemClock.now() // hard dependency
var rate = TaxGateway.lookup(order) // hard dependency
...
}
}
// 2. Introduce a seam with the safest edit available (extract + override)
class Invoicer {
protected function now() { return SystemClock.now() } // seam
protected function taxRate(order) { return TaxGateway.lookup(order) } // seam
function issue(order) { var now = this.now(); var rate = this.taxRate(order); ... }
}
// 3. Characterization test: assert a deliberately wrong value, read the actual, pin it
test("issue() current behavior for a 2-line domestic order") {
var invoicer = new TestInvoicer(fixedNow = "2020-01-01", fixedRate = 0.2)
assertEquals("PLACEHOLDER", invoicer.issue(sampleOrder()).toString())
// run -> actual is "TOTAL 240.00 (rounded down)"; paste it in.
// NOTE: rounding looks wrong -> pinned as-is, tracked in TICKET-123
}go deeper
Say you would write tests that capture what the code does today before changing anything, then refactor in small steps and rerun them each time.
Name characterization tests and how to write them (assert a wrong value, read the actual, pin it), and mention breaking hard dependencies like clocks and network calls so the code can be instantiated in a test.
Add seams and their enabling points, the safest-edit-first ordering, golden-master/approval testing for wide outputs, pinning at a stable boundary, committing tests separately, and the classify-on-failure rule (regression vs deliberate fix).
Frame it as risk management with a portfolio of verification techniques matched to consequence: targeted characterization plus IDE-verified refactorings for ordinary modules; parallel run/shadow traffic, feature toggles and staged rollout for high-consequence engines whose real input distribution cannot be reproduced. Address cost and sequencing — pinning is the price of incrementalism, and it is the honest comparison point against a rewrite proposal — plus the organisational issue that pinned bugs must be tracked, not forgotten.
## The problem Refactoring is defined as behavior-preserving. Preservation has to be *verified*, and the ordinary verifier is a test suite. Legacy code — Feathers' definition: **code without tests** — offers no verifier. The chicken-and-egg problem is famous: *to change the code safely you need tests; to add tests you often need to change the code* (because the class cannot be instantiated without a database, or the logic is buried in a 900-line method reachable only through a UI). ## Characterization tests (a.k.a. pinning tests) A **characterization test** records what the system *currently does*, not what it *should do*. Procedure: 1. Write a test that calls the code with some realistic input and asserts something you know is wrong, e.g. `assertEquals("PLACEHOLDER", result)`. 2. Run it. The failure message tells you the actual value. 3. Change the assertion to that actual value. Now the test passes. 4. Repeat for inputs that exercise different branches — use coverage tooling to find untested paths and add cases until the branches you intend to touch are covered. Properties to internalise: - **They can encode bugs, and that is correct.** Their job is to detect *change*, not to certify correctness. If a wrong-looking value is pinned, add a comment: "pinned; suspected defect, see TICKET-123". - **They are temporary scaffolding turning permanent.** Many become genuine regression tests. Some, especially brittle golden masters, should be replaced by intention-revealing tests once the design allows. - **They must be fast and deterministic**, or the refactoring loop dies. Freeze time, seed randomness, fix locales/time zones, sort non-deterministic collections. - **Coverage before change is the metric that matters** — not overall coverage, but coverage of the exact lines and branches you are about to move. ### Golden master / approval testing When output is large (a rendered document, a report, a JSON payload, a log of side effects), don't hand-write assertions. Capture the output to a file (the *approved* file), then assert that each run produces an identical file. Combine with input generation: run thousands of generated or sampled-from-production inputs, store the outputs, and treat any diff after a refactoring step as a red flag. This is extremely effective for gnarly calculation engines, and its weakness is equally clear: a diff tells you *something changed* but not *what is conceptually wrong*, and the files rot as behavior legitimately evolves. ## Seams A **seam** (Feathers) is a place where you can alter program behavior **without editing in that place**. Every seam has an **enabling point** — the place where you choose which behavior is used. Common kinds: | Seam type | Mechanism | Enabling point | |---|---|---| | **Object seam** | Call goes through an interface / overridable method | Where the object is constructed or injected | | **Parameter seam** | Dependency passed as an argument (clock, id generator, gateway) | The caller | | **Subclass-and-override** | Extract the awkward call into a protected method; test subclass overrides it | Which subclass the test constructs | | **Link seam** | Substitute a different implementation at link/classpath/module-resolution time | Build or test configuration | | **Preprocessing / config seam** | Compile-time switch, dependency-injection container binding, environment variable | Build or runtime configuration | The discipline is to create the seam using the *least risky* edit possible, ideally an automated IDE refactoring, because at that moment you still have no tests. Feathers' named techniques include **Extract and Override Call**, **Parameterize Constructor**, **Extract Interface**, **Introduce Instance Delegator**, and **Sprout Method/Class** (put the new behavior in a fresh, tested unit and call it from the untested code, deferring the untangling). ## The full playbook 1. **Define the blast radius** — which behavior must be preserved, and for whom. Not the whole system; the module and its consumers. 2. **Find the pinch point** — a narrow interface through which most of the behavior flows; test there rather than at ten scattered internal methods, so tests survive the refactoring. 3. **Break the smallest dependency that blocks instantiation** using a seam, with the safest available edit. 4. **Write characterization tests** until the code paths you will touch are covered. Add golden-master coverage for wide outputs. 5. **Commit the tests separately** — they are valuable on their own and provably change no production behavior. 6. **Refactor in small steps**, running the suite each time, committing at green. Prefer automated refactorings; hand edits are where behavior slips. 7. **When a characterization test fails, stop and classify**: regression (revert), or a bug you have deliberately fixed (update the pin, in a separate commit, under the function hat). 8. **Grow real tests** as the structure improves; delete or rewrite the ugliest pins once expressive tests replace them. ## Complements when unit tests are impossible - **Higher-level tests first**: an end-to-end or contract test around the module gives coarse but real protection while you carve out testable units. - **Production comparison (dark launching / parallel run)**: run old and new implementations side by side on real traffic, serve the old result, log differences. This is the strongest evidence available for behavior equivalence on a system whose real input distribution you cannot reproduce — at the cost of infrastructure and doubled compute. - **Feature toggles** to switch between old and new paths, enabling instant rollback. - **Static/mechanical safety**: strong typing, compiler checks, and IDE-verified refactorings reduce the classes of error you need tests to catch. ## Trade-offs and honest limits - Characterization tests **cement current behavior**, including behavior you may later want to change; expect to update them intentionally. - They are typically **structure-agnostic only if written at the right level**. Pin at a stable boundary, or your tests will fight the very refactoring they were written to enable. - They cost real time up front, which is exactly the cost a rewrite proposal tries to dodge — and exactly the cost that makes incremental change possible at all. - Full coverage is not the goal; coverage of the code you intend to move is.
- A characterization test pins behavior you are fairly sure is a bug. Do you fix it while refactoring?No. Pin it as-is with a comment and a ticket, complete and commit the refactoring with behavior preserved, then switch to the behavior-changing hat: write a test for the correct behavior, change the pin in that same commit, and ship the fix separately so it can be reviewed, communicated to consumers, and reverted independently.
- How do you decide at what level to write the characterization tests?At the narrowest boundary that both (a) is stable across the refactoring you plan and (b) captures the behavior consumers depend on. Pinning internal private methods means your tests break as soon as you move them — they fight the refactoring. Pinning at the module's public entry point, a service interface, or an HTTP contract keeps them valid while everything inside changes.
- What can you do when the module's real behavior depends on production data you cannot reproduce in tests?Run a parallel/shadow comparison: execute old and new implementations on live traffic, serve the old result, and log or aggregate differences. It gives evidence against the real input distribution, which no synthetic fixture can. Costs are doubled compute, care with side effects (the new path must be side-effect-free or fully sandboxed), and privacy handling on logged payloads.
- Is high overall test coverage the goal before refactoring?No — targeted coverage is. What matters is that the specific lines and branches you are about to move are exercised, and that the tests sit at a boundary that survives the move. A codebase at 80% coverage with 0% on the class you are restructuring is unprotected for this work.
Before demolishing an interior wall you don't know is load-bearing, you install temporary props and put tell-tales — thin glass strips — across the cracks. The tell-tales don't say the building is well-designed; they say "something moved," which is exactly the signal you need while you work.
saying these in an interview costs you the question
- Refactoring untested code and calling it safe because "it compiles" or "I read it carefully"
- Writing characterization tests that assert intended behavior, so they fail before you start
- Refusing to pin behavior that looks buggy, leaving the riskiest paths unprotected
- Pinning private methods and internal call order, so the tests block the refactoring they were meant to enable
- Chasing a global coverage number instead of covering the exact code being moved
- Non-deterministic pins (real clock, real network, unsorted collections) that flake and destroy trust in the suite
- Treating "add tests first" as unaffordable while treating a full rewrite as affordable