skip to content

When is incremental refactoring the right choice versus a full rewrite, and how do the strangler fig pattern and branch by abstraction let you restructure a running system without a big-bang cutover?

level: seniorimportance: should knowfreq 56%

answer

  1. Default = refactor; rewrite only if substrate is untenable
  2. Rewrite loses undocumented behavior + freezes delivery
  3. Strangler fig: router → slice by slice → delete legacy
  4. Branch by abstraction: interface → both impls → toggle → delete
  5. Failure mode: stall at 70%, two systems forever

basics

~20 s

Prefer incremental refactoring: it keeps the system shippable and delivers value continuously. Rewrite only when the platform itself is untenable. Strangler fig routes traffic feature-by-feature from old to new until the old is unused; branch by abstraction puts an interface in front of a component so old and new implementations can coexist behind a toggle.

solid answer

~60 s

Default to refactoring. A rewrite restarts the clock on every undocumented behavior the old system accumulated, halts feature delivery, and creates a long period where two systems must be maintained — the classic reason "second system" projects overrun or die. Rewrites are justified when the constraint is not structure but substrate: an unsupported runtime or platform, a data model that cannot be migrated in place, a licence or compliance wall, or a codebase so small that rewriting is genuinely cheaper. In between sits *incremental replacement*. **Strangler fig** (Fowler): put a facade/router in front of the legacy system, implement one capability at a time in the new system, route that slice's traffic there, and repeat until the old system carries no traffic and is deleted. **Branch by abstraction** (Hodgson/Fowler): introduce an abstraction over the component you want to replace, make all callers use it, build a second implementation behind it, switch via configuration, then delete the old one — all on mainline, no long-lived branch. Both give incremental value, continuous shippability and cheap rollback; both cost temporary duplication plus routing/toggle machinery.

code

pseudocode · 19 lines
pseudocode
// BRANCH BY ABSTRACTION — all on mainline, always releasable

// 1. Abstraction over the component being replaced
interface PaymentStore { save(p); findById(id) }

// 2. Existing behavior wrapped, all callers migrated to the interface
class LegacyPaymentStore implements PaymentStore { /* old SQL */ }

// 3. New implementation built incrementally behind the interface (dark)
class NewPaymentStore implements PaymentStore { /* new schema */ }

// 4. Switch at the wiring point — per-tenant, instantly revertible
function paymentStore(tenant) {
  return toggles.on("new-payment-store", tenant) ? new NewPaymentStore()
                                                 : new LegacyPaymentStore()
}

// 5. When 100% is on the new path and stable: delete LegacyPaymentStore,
//    delete the toggle, and delete the abstraction if it was only scaffolding.

go deeper

for a junior

Say incremental refactoring is usually safer because the system keeps working, and that rewrites are risky because behavior nobody wrote down gets lost.

for a middle

Add concrete rewrite triggers (unsupported platform, unmigratable data model), and describe strangler fig as routing feature-by-feature to the new system until the old one can be deleted.

for a senior

Compare honestly — refactoring cost including characterization tests versus rewrite cost including behavior archaeology, dual maintenance and cutover — and explain branch by abstraction and why it beats a long-lived branch. Mention shared data as the hard part.

for a principal

Frame it as a portfolio/risk decision with organisational constraints: who funds a delivery freeze, how dual maintenance is staffed, how you prove equivalence at scale (parallel run, contract tests, per-tenant rollout), and how you prevent the 70% stall — explicit deletion milestones, a legacy-traffic metric, and treating removal as the definition of done. Note that a rewrite delivered incrementally is largely a strangler fig, so the real dichotomy is big-bang versus incremental, not new-code versus old-code.

## The decision ### Why refactoring is the default - **The system keeps working and shipping.** Business value continues to flow while the structure improves. A rewrite freezes it — and organisations rarely tolerate a feature freeze long enough to finish. - **Behavior knowledge lives in the code.** A mature system encodes years of edge cases, regulatory quirks, customer-specific hacks and bug-compatibility that exist in nobody's head or document. Rewriting means rediscovering them, usually via production incidents. - **Risk is spread.** Every refactoring step is small, verified and revertible. A rewrite concentrates all risk into one cutover event. - **Feedback arrives early.** You learn whether the new design is actually better after a week, not after a year. - **The "two systems" tax is bounded.** During a rewrite, every urgent change must be made twice, or the old system diverges and the target keeps moving. ### When a rewrite (or incremental replacement) is genuinely right | Trigger | Why refactoring cannot fix it | |---|---| | Runtime/platform end-of-life or unsupported, with no upgrade path | The constraint is outside the code | | A language/framework nobody can hire for or that has no security patches | Sustainability, not structure | | A fundamentally wrong data model (e.g. one table encoding five entities) that cannot be migrated in place | The migration *is* the project | | Regulatory/licensing wall (e.g. a dependency that can no longer be used) | External | | The system is small — weeks of work — and well understood | Rewrite is simply cheaper | | Requirements have changed so fundamentally that <20% of behavior survives | Preservation has little value | Note the honest framing: a rewrite that ships as a **big bang** is the risky version. A rewrite delivered **incrementally** — strangler fig — retains most of refactoring's risk profile. ### The dishonest comparison to watch for Rewrite proposals usually compare *the messy known system* to *the imagined clean one*. The fair comparison is: cost of refactoring including characterization tests, versus cost of rewriting **including** behavior archaeology, dual maintenance, migration of data and integrations, retraining, and the cutover. Also beware Chesterton's Fence: the ugly conditional you plan to drop probably encodes a real requirement. ## Strangler fig pattern Named after the fig that grows around a host tree until the host dies and only the fig's structure remains. **Mechanics** 1. Put an **interception layer** in front of the legacy system: an HTTP router/reverse proxy, an API gateway, an event-stream splitter, or a facade class inside a monolith. 2. Pick a **slice** — one capability, one endpoint, one bounded context — preferably one with clean boundaries and meaningful value. 3. Implement it in the new system. 4. **Route that slice's traffic** to the new implementation. Optionally run both and compare first (shadow/parallel run). 5. Verify, keep the ability to route back instantly. 6. Repeat. When no traffic reaches the legacy path, **delete it** — the deletion is part of the pattern, not an optional afterthought. **Hard parts** - **Shared data.** The classic blocker: both systems need the same tables. Options are shared database (fast, couples them), synchronisation/CDC (eventual consistency, ordering issues), or migrating data ownership with the slice (cleanest, hardest). - **Cross-cutting concerns.** Authentication, sessions, feature flags, audit trails, reporting must work across both halves during the transition. - **Reporting and analytics** often read the legacy schema directly and are forgotten until they break. - **Discipline to finish.** The most common failure is stalling at 70%: two systems forever, double the maintenance. Track "legacy traffic remaining" as an explicit metric with a deletion milestone. ## Branch by abstraction For replacing a component *inside* a codebase without a long-lived branch (Paul Hammant / Steve Smith; described by Fowler). 1. **Introduce an abstraction** (interface/port) over the part to be replaced. 2. **Migrate all clients** to call through it, incrementally, keeping mainline green. This is often the bulk of the work. 3. **Add a second implementation** behind the abstraction, built incrementally on mainline; it can be incomplete because nothing routes to it yet. 4. **Switch** — by configuration, feature toggle, or dependency wiring, per-environment and ideally per-request or per-tenant for gradual rollout. 5. **Remove** the old implementation, and often the abstraction too if it was scaffolding. **Why not a long-lived branch?** Because a branch that diverges for months creates a merge event with the same all-at-once risk you were trying to avoid, blocks continuous integration, and hides work from the team. Branch by abstraction keeps everything on mainline, integrated continuously, and always releasable — the toggle simply keeps the new path dark. **Costs**: the abstraction may leak or be shaped by the old implementation; toggles proliferate and must be cleaned up; two implementations coexist and both may need bug fixes for a while. ## Choosing between them - **Strangler fig** operates at the *system/deployment* boundary — replacing a service, a monolith, or a whole application; the router is infrastructure. - **Branch by abstraction** operates at the *code* boundary — replacing a persistence layer, a template engine, an internal component; the switch is a toggle in the app. - They compose: strangler at the edges, branch by abstraction within each slice. ## Supporting practices for either - **Parallel run / dark launch**: execute both implementations on real traffic, serve the old, compare and log differences. The strongest equivalence evidence available. - **Feature toggles with per-tenant/percentage rollout** and a tested kill switch. - **Contract tests** at the boundary so both implementations are held to the same spec. - **Explicit deletion milestones** with dates, plus a metric for remaining legacy traffic. - **Preparatory refactoring first** — often the old code must be reshaped before an abstraction can be drawn at all.

  • What most often blocks a strangler fig migration in practice?
    Shared data. The new slice and the legacy system need the same tables, so you must choose between a shared database (fast but couples the systems and prevents schema evolution), synchronisation/change-data-capture (works, but introduces eventual consistency, ordering and conflict problems), or moving data ownership with the slice (cleanest end state, most work). Second most common is reporting or analytics reading the legacy schema directly and breaking unnoticed.
  • Why is branch by abstraction preferred over a long-lived refactoring branch?
    A long-lived branch reintroduces big-bang risk at merge time, diverges from mainline as others keep committing, and prevents continuous integration of the new work. Branch by abstraction keeps every change on mainline, integrated and releasable, with the incomplete new implementation simply not routed to. The cost is temporary duplication and toggle management, which is far cheaper than a months-long merge.
  • How do you know when a strangler fig migration is actually finished?
    When no traffic reaches the legacy path and the legacy code and its infrastructure are deleted. Until deletion, you are paying for two systems, and dual maintenance is the cost that kills these programmes. Make it measurable: track remaining legacy request share, set a dated deletion milestone, and treat removal of the routing layer and toggles as part of the work, not cleanup for later.
  • A team argues the code is beyond refactoring and only a rewrite will do. What questions do you ask?
    What exactly is untenable — the structure (refactorable) or the substrate (runtime, data model, licensing)? What behavior exists that nobody has documented, and how will you rediscover it? Who maintains the old system during the rewrite and how do urgent changes get made twice? What is the cutover plan and rollback? Can the same end state be reached slice by slice with a strangler fig, so value ships continuously and risk stays bounded?

Renovating a bridge that must stay open: you close one lane, rebuild it, move traffic across, rebuild the other. A rewrite is demolishing the bridge and telling the city to wait — defensible only if the foundations are condemned.

saying these in an interview costs you the question

  • Comparing a messy real system to an imagined perfect one and ignoring behavior archaeology, dual maintenance and cutover cost
  • Assuming a rewrite means big-bang — incremental replacement retains most of refactoring's risk profile
  • Treating strangler fig as done when the new system is feature-complete, rather than when the legacy code is deleted
  • Long-lived refactoring branches, which recreate the big-bang merge risk they were meant to avoid
  • Ignoring shared data ownership, the usual blocker for incremental replacement
  • Leaving toggles and routing layers in place permanently, so complexity is added but never removed
  • Dismissing odd legacy conditionals as cruft without checking what requirement they encode (Chesterton's Fence)

context