skip to content

When is a stream of small, continuous in-place improvements the right strategy for a decaying system, and when do you need a larger planned effort such as a Strangler Fig migration or a full rewrite? How do you decide?

level: seniorimportance: must knowfreq 45%

answer

  1. Local pain vs structural pain
  2. Spectrum: tidy → campaign → branch-by-abstraction → strangler → rewrite
  3. Rewrite kills: parity treadmill, dual maintenance, lost edge cases
  4. Second-system effect; same forces rot the new one
  5. Don't polish what you're about to strangle

basics

~20 s

Incremental cleanup works when the architecture is sound and the pain is spread across code you touch often. You need a bigger planned effort when the problem is structural — wrong boundaries, dead platform — so no local edit can reach it.

solid answer

~50 s

Decide by asking whether the pain is *local* or *structural*. If bad names, duplication and long functions inside otherwise-reasonable boundaries are the problem, opportunistic cleanup wins: it's continuously funded, self-targeting toward high-churn files, and carries almost no risk. If the problem is the boundaries themselves — a shared mutable database everything reaches into, an unsupported runtime, a coupling topology that makes every feature cross ten modules — no sequence of local tidies converges, and you need a planned intervention. Prefer incremental *architectural* techniques over big-bang rewrites: Strangler Fig (route traffic progressively from old to new behind a façade), branch by abstraction (introduce an abstraction, build the replacement behind it, flip, delete), and parallel run for verification. Reserve full rewrites for small scope, a genuinely dead platform, or when feature parity is explicitly not required. Big-bang rewrites fail on the parity treadmill, dual maintenance and lost undocumented behaviour.

code

text · 7 lines
text
Branch by abstraction — incremental replacement, always releasable

1. callers -> LegacyStore                (start)
2. callers -> Store (interface) -> LegacyStore
3. callers -> Store -> {LegacyStore | NewStore}   # flag/router, parallel run + diff
4. callers -> Store -> NewStore           # flip when diffs are clean
5. delete LegacyStore, optionally inline Store

go deeper

for a junior

Say small continuous improvements are usually safer, and that rewrites are risky because you can lose behaviour the old code handled. Naming the strangler idea at a high level is a bonus.

for a middle

Distinguish local from structural problems, name Strangler Fig and feature flags, and list rewrite failure modes: parity treadmill, dual maintenance, no value until the end.

for a senior

Present the full spectrum with branch-by-abstraction and parallel run, argue from evidence (hotspots, defect density, lead time), give the honest cases where a rewrite wins, and mention not polishing code you plan to delete.

for a principal

Add organizational reasoning: funding stamina and why slice-wise value delivery de-risks sponsorship, the second-system effect, addressing the forces that caused the decay (ownership, tests, capacity), and how you'd sequence migrations against business roadmap and compliance deadlines.

## Framing the decision The real question is: **can a sequence of local, low-risk edits reach the problem?** - **Local pain** — bad names, duplicated blocks, long functions, missing tests, magic values. Each instance is independently fixable and each fix is independently valuable. → Boy Scout Rule / opportunistic refactoring. - **Structural pain** — the module boundaries are wrong, everything shares one mutable data store, the framework is end-of-life, the deployment unit forces the whole org to release together. No local edit improves this; you can polish every function in the system and still have the same architecture. → planned, funded intervention. A useful test: *if every file were individually beautiful, would the problem be gone?* If no, it's structural. ## The spectrum of interventions There are more than two options, and knowing the middle of the spectrum is what distinguishes a senior answer. 1. **Opportunistic cleanup (Boy Scout Rule)** — continuous, unfunded, zero-ceremony, targets whatever you touch. 2. **Scheduled/campaign refactoring** — a named, ticketed refactoring of one area, still behaviour-preserving, still incremental, but planned and estimated. 3. **Branch by abstraction** — introduce an abstraction over the thing you want to replace, build the new implementation behind it, migrate callers gradually (often behind a feature flag), then delete the old one. Keeps the system releasable throughout; avoids long-lived branches. 4. **Strangler Fig application** (named by Martin Fowler after the strangler fig, which grows around a host tree until the host dies) — put a façade/router in front of the legacy system, implement slices of functionality in a new system, redirect traffic slice by slice, and retire the old system when nothing routes to it. Each slice ships and earns value; you can stop or pause at any point. 5. **Big-bang rewrite** — build a full replacement in parallel, cut over once. Supporting technique: **parallel run** (dark launching) — send production traffic to both old and new implementations, serve the old result, and diff the outputs to verify equivalence before flipping. This is how you get confidence in a replacement when the specification is "whatever the old system does". ## Why big-bang rewrites usually fail Joel Spolsky's "Things You Should Never Do" (2000, on the Netscape rewrite) is the canonical reference; the mechanisms are worth naming individually: - **Feature parity treadmill.** The old system keeps getting features while you rebuild, so the target moves. You can be perpetually 90% done. - **Dual maintenance.** Every bug must be fixed twice, doubling cost during the longest, most stressful period. - **Encoded knowledge loss.** Ugly conditionals often *are* the requirements — accumulated fixes for real edge cases, regulatory quirks, a customer's odd data. "Nobody knows why this branch exists" means you're about to rediscover it in production. - **No incremental value.** Nothing ships until everything ships, so the effort is politically fragile: a leadership change, a budget cut, or one missed milestone kills it, and all invested work is lost. - **Second-system effect** (Fred Brooks, *The Mythical Man-Month*) — the replacement accumulates every feature the team wished they'd had, so it over-generalizes and is late. - **Same team, same forces.** If the pressures that produced the mess (deadlines, no tests, unclear ownership) persist, the new system rots on the same trajectory, just later. ## When a rewrite *is* the right call Be able to name real cases, otherwise the answer sounds dogmatic: - The **platform is dead**: unsupported runtime/framework with unpatched security holes, or hardware/OS no longer available. - **Small scope**: the component is a few thousand lines with a narrow, well-understood interface — rewriting is cheaper than archaeology. - **Parity is explicitly not required**: the business has decided the product should behave differently, so preserving old behaviour has negative value. - **The core model is wrong** in a way abstraction can't hide (e.g. a fundamentally single-tenant design that must become multi-tenant), *and* the system is small enough to replace. - **Economics dominate**: licensing or hosting costs of the old stack exceed the rebuild cost. Even then, prefer to execute the rewrite *incrementally* via Strangler Fig rather than as a cutover, unless the system is genuinely tiny. ## How to decide — a practical checklist 1. **Measure, don't feel.** Where is the cost actually incurred? Use hotspot analysis (change frequency × complexity), defect density per area, lead time and change-failure rate per component, and time-to-onboard. Decay you never touch costs little. 2. **Classify the pain**: local vs structural, per hotspot. 3. **Check reachability**: is there a sequence of behaviour-preserving steps from here to the target? If yes, prefer it. The **Mikado Method** is the explicit technique — attempt the change, note what breaks, revert, fix prerequisites first, building a dependency graph of small safe steps. 4. **Check the safety net.** Without tests, *no* strategy is safe; characterization tests and a parallel-run harness are prerequisites for either path. 5. **Check the ceiling.** If incremental work would take years to reach an acceptable state, or is blocked by an external constraint (EOL platform, compliance deadline), escalate. 6. **Check organizational stamina.** Long efforts need sustained sponsorship. If you can't guarantee 18 months of funding, choose the strategy that delivers value in slices — which is an argument *for* Strangler Fig and *against* big-bang. 7. **Fix the forces, not just the code.** Whatever you choose, if the root cause was no tests, no ownership, or no time, address that too — otherwise you'll be here again. ## Where the Boy Scout Rule sits Even during a large migration, the rule still applies to the code you touch — but with a caveat: **don't polish code you're about to delete.** Cleanup in a module scheduled for strangulation in six weeks is wasted effort, and worse, it invests people in keeping it. Direct opportunistic cleanup at the code that will survive. ## How to answer Open with the local-vs-structural test, list the *spectrum* (not a binary) with Strangler Fig and branch-by-abstraction in the middle, give the concrete failure mechanisms of big-bang rewrites, name the honest cases where a rewrite wins, and finish with how you'd gather evidence (hotspots, defect density, lead time) rather than deciding on taste.

  • How would you verify a Strangler Fig slice behaves identically to the legacy path before switching traffic?
    Parallel run: send real production requests to both implementations, serve the legacy response, and record and diff the outputs. Investigate every divergence — many will turn out to be undocumented legacy behaviour you must preserve. Flip traffic gradually by percentage, keeping the ability to route back instantly.
  • A team says they must rewrite because the code is 'unmaintainable'. What evidence would change your mind either way?
    Ask for numbers per component: change frequency and complexity hotspots, defect density, lead time for a typical change, and how much of the cost concentrates in a few modules. If 80% of pain is in 5% of files with sound boundaries, that's a targeted refactoring, not a rewrite. If the pain is spread and rooted in the boundaries or a dead platform, that supports escalation.
  • What is the Mikado Method and where does it fit?
    A technique for large changes in legacy code: attempt the change naively, observe what breaks, revert, and record the breakages as prerequisites; recurse until you find changes that can be made safely, then unwind the graph. It turns a scary structural change into an ordered sequence of small, always-green steps — bridging opportunistic cleanup and full migration.

Repainting rooms and fixing taps keeps a sound house pleasant, but no amount of painting fixes a cracked foundation. And even for the foundation you underpin it section by section while people live there — you rarely demolish and rebuild the house from scratch.

saying these in an interview costs you the question

  • Treating the choice as a binary between tidying and a full rewrite, missing branch-by-abstraction and Strangler Fig
  • Proposing a big-bang rewrite without a plan for feature parity, dual maintenance, or preserving undocumented behaviour
  • Assuming ugly code encodes no requirements — the weird conditionals are often bug fixes for real edge cases
  • Deciding from taste rather than evidence (churn/complexity hotspots, defect density, lead time)
  • Rewriting with the same team under the same pressures and no change to tests or ownership, guaranteeing a repeat
  • Spending careful cleanup effort on modules that are scheduled for deletion or strangulation

context