When is a stream of small, continuous in-place improvements the right strategy for a decaying system, and when do you need a larger planned effort such as a Strangler Fig migration or a full rewrite? How do you decide?
answer
- Local pain vs structural pain
- Spectrum: tidy → campaign → branch-by-abstraction → strangler → rewrite
- Rewrite kills: parity treadmill, dual maintenance, lost edge cases
- Second-system effect; same forces rot the new one
- Don't polish what you're about to strangle
basics
~20 sIncremental cleanup works when the architecture is sound and the pain is spread across code you touch often. You need a bigger planned effort when the problem is structural — wrong boundaries, dead platform — so no local edit can reach it.
solid answer
~50 sDecide by asking whether the pain is *local* or *structural*. If bad names, duplication and long functions inside otherwise-reasonable boundaries are the problem, opportunistic cleanup wins: it's continuously funded, self-targeting toward high-churn files, and carries almost no risk. If the problem is the boundaries themselves — a shared mutable database everything reaches into, an unsupported runtime, a coupling topology that makes every feature cross ten modules — no sequence of local tidies converges, and you need a planned intervention. Prefer incremental *architectural* techniques over big-bang rewrites: Strangler Fig (route traffic progressively from old to new behind a façade), branch by abstraction (introduce an abstraction, build the replacement behind it, flip, delete), and parallel run for verification. Reserve full rewrites for small scope, a genuinely dead platform, or when feature parity is explicitly not required. Big-bang rewrites fail on the parity treadmill, dual maintenance and lost undocumented behaviour.
code
text · 7 linesBranch by abstraction — incremental replacement, always releasable
1. callers -> LegacyStore (start)
2. callers -> Store (interface) -> LegacyStore
3. callers -> Store -> {LegacyStore | NewStore} # flag/router, parallel run + diff
4. callers -> Store -> NewStore # flip when diffs are clean
5. delete LegacyStore, optionally inline Storego deeper
Say small continuous improvements are usually safer, and that rewrites are risky because you can lose behaviour the old code handled. Naming the strangler idea at a high level is a bonus.
Distinguish local from structural problems, name Strangler Fig and feature flags, and list rewrite failure modes: parity treadmill, dual maintenance, no value until the end.
Present the full spectrum with branch-by-abstraction and parallel run, argue from evidence (hotspots, defect density, lead time), give the honest cases where a rewrite wins, and mention not polishing code you plan to delete.
Add organizational reasoning: funding stamina and why slice-wise value delivery de-risks sponsorship, the second-system effect, addressing the forces that caused the decay (ownership, tests, capacity), and how you'd sequence migrations against business roadmap and compliance deadlines.
## Framing the decision The real question is: **can a sequence of local, low-risk edits reach the problem?** - **Local pain** — bad names, duplicated blocks, long functions, missing tests, magic values. Each instance is independently fixable and each fix is independently valuable. → Boy Scout Rule / opportunistic refactoring. - **Structural pain** — the module boundaries are wrong, everything shares one mutable data store, the framework is end-of-life, the deployment unit forces the whole org to release together. No local edit improves this; you can polish every function in the system and still have the same architecture. → planned, funded intervention. A useful test: *if every file were individually beautiful, would the problem be gone?* If no, it's structural. ## The spectrum of interventions There are more than two options, and knowing the middle of the spectrum is what distinguishes a senior answer. 1. **Opportunistic cleanup (Boy Scout Rule)** — continuous, unfunded, zero-ceremony, targets whatever you touch. 2. **Scheduled/campaign refactoring** — a named, ticketed refactoring of one area, still behaviour-preserving, still incremental, but planned and estimated. 3. **Branch by abstraction** — introduce an abstraction over the thing you want to replace, build the new implementation behind it, migrate callers gradually (often behind a feature flag), then delete the old one. Keeps the system releasable throughout; avoids long-lived branches. 4. **Strangler Fig application** (named by Martin Fowler after the strangler fig, which grows around a host tree until the host dies) — put a façade/router in front of the legacy system, implement slices of functionality in a new system, redirect traffic slice by slice, and retire the old system when nothing routes to it. Each slice ships and earns value; you can stop or pause at any point. 5. **Big-bang rewrite** — build a full replacement in parallel, cut over once. Supporting technique: **parallel run** (dark launching) — send production traffic to both old and new implementations, serve the old result, and diff the outputs to verify equivalence before flipping. This is how you get confidence in a replacement when the specification is "whatever the old system does". ## Why big-bang rewrites usually fail Joel Spolsky's "Things You Should Never Do" (2000, on the Netscape rewrite) is the canonical reference; the mechanisms are worth naming individually: - **Feature parity treadmill.** The old system keeps getting features while you rebuild, so the target moves. You can be perpetually 90% done. - **Dual maintenance.** Every bug must be fixed twice, doubling cost during the longest, most stressful period. - **Encoded knowledge loss.** Ugly conditionals often *are* the requirements — accumulated fixes for real edge cases, regulatory quirks, a customer's odd data. "Nobody knows why this branch exists" means you're about to rediscover it in production. - **No incremental value.** Nothing ships until everything ships, so the effort is politically fragile: a leadership change, a budget cut, or one missed milestone kills it, and all invested work is lost. - **Second-system effect** (Fred Brooks, *The Mythical Man-Month*) — the replacement accumulates every feature the team wished they'd had, so it over-generalizes and is late. - **Same team, same forces.** If the pressures that produced the mess (deadlines, no tests, unclear ownership) persist, the new system rots on the same trajectory, just later. ## When a rewrite *is* the right call Be able to name real cases, otherwise the answer sounds dogmatic: - The **platform is dead**: unsupported runtime/framework with unpatched security holes, or hardware/OS no longer available. - **Small scope**: the component is a few thousand lines with a narrow, well-understood interface — rewriting is cheaper than archaeology. - **Parity is explicitly not required**: the business has decided the product should behave differently, so preserving old behaviour has negative value. - **The core model is wrong** in a way abstraction can't hide (e.g. a fundamentally single-tenant design that must become multi-tenant), *and* the system is small enough to replace. - **Economics dominate**: licensing or hosting costs of the old stack exceed the rebuild cost. Even then, prefer to execute the rewrite *incrementally* via Strangler Fig rather than as a cutover, unless the system is genuinely tiny. ## How to decide — a practical checklist 1. **Measure, don't feel.** Where is the cost actually incurred? Use hotspot analysis (change frequency × complexity), defect density per area, lead time and change-failure rate per component, and time-to-onboard. Decay you never touch costs little. 2. **Classify the pain**: local vs structural, per hotspot. 3. **Check reachability**: is there a sequence of behaviour-preserving steps from here to the target? If yes, prefer it. The **Mikado Method** is the explicit technique — attempt the change, note what breaks, revert, fix prerequisites first, building a dependency graph of small safe steps. 4. **Check the safety net.** Without tests, *no* strategy is safe; characterization tests and a parallel-run harness are prerequisites for either path. 5. **Check the ceiling.** If incremental work would take years to reach an acceptable state, or is blocked by an external constraint (EOL platform, compliance deadline), escalate. 6. **Check organizational stamina.** Long efforts need sustained sponsorship. If you can't guarantee 18 months of funding, choose the strategy that delivers value in slices — which is an argument *for* Strangler Fig and *against* big-bang. 7. **Fix the forces, not just the code.** Whatever you choose, if the root cause was no tests, no ownership, or no time, address that too — otherwise you'll be here again. ## Where the Boy Scout Rule sits Even during a large migration, the rule still applies to the code you touch — but with a caveat: **don't polish code you're about to delete.** Cleanup in a module scheduled for strangulation in six weeks is wasted effort, and worse, it invests people in keeping it. Direct opportunistic cleanup at the code that will survive. ## How to answer Open with the local-vs-structural test, list the *spectrum* (not a binary) with Strangler Fig and branch-by-abstraction in the middle, give the concrete failure mechanisms of big-bang rewrites, name the honest cases where a rewrite wins, and finish with how you'd gather evidence (hotspots, defect density, lead time) rather than deciding on taste.
- How would you verify a Strangler Fig slice behaves identically to the legacy path before switching traffic?Parallel run: send real production requests to both implementations, serve the legacy response, and record and diff the outputs. Investigate every divergence — many will turn out to be undocumented legacy behaviour you must preserve. Flip traffic gradually by percentage, keeping the ability to route back instantly.
- A team says they must rewrite because the code is 'unmaintainable'. What evidence would change your mind either way?Ask for numbers per component: change frequency and complexity hotspots, defect density, lead time for a typical change, and how much of the cost concentrates in a few modules. If 80% of pain is in 5% of files with sound boundaries, that's a targeted refactoring, not a rewrite. If the pain is spread and rooted in the boundaries or a dead platform, that supports escalation.
- What is the Mikado Method and where does it fit?A technique for large changes in legacy code: attempt the change naively, observe what breaks, revert, and record the breakages as prerequisites; recurse until you find changes that can be made safely, then unwind the graph. It turns a scary structural change into an ordered sequence of small, always-green steps — bridging opportunistic cleanup and full migration.
Repainting rooms and fixing taps keeps a sound house pleasant, but no amount of painting fixes a cracked foundation. And even for the foundation you underpin it section by section while people live there — you rarely demolish and rebuild the house from scratch.
saying these in an interview costs you the question
- Treating the choice as a binary between tidying and a full rewrite, missing branch-by-abstraction and Strangler Fig
- Proposing a big-bang rewrite without a plan for feature parity, dual maintenance, or preserving undocumented behaviour
- Assuming ugly code encodes no requirements — the weird conditionals are often bug fixes for real edge cases
- Deciding from taste rather than evidence (churn/complexity hotspots, defect density, lead time)
- Rewriting with the same team under the same pressures and no change to tests or ownership, guaranteeing a repeat
- Spending careful cleanup effort on modules that are scheduled for deletion or strangulation