What strategies exist for paying down technical debt, and how do you choose between opportunistic refactoring, incremental replacement, and a big-bang rewrite?
answer
- Camping rule for everyday debt
- Make the change easy, then make the easy change
- Branch by abstraction keeps trunk releasable
- Strangler fig for whole legacy systems
- Rewrite last; tests first or it is not refactoring
basics
~20 sSmall debt: clean up as you pass through the code, leaving it better than you found it. Medium debt: replace it piece by piece behind a stable interface while the system keeps running. Full rewrites are a last resort - slow, risky, and they freeze delivery.
solid answer
~50 sMatch the strategy to the size and interest of the debt. Opportunistic (camping-rule) refactoring handles everyday debt: whoever touches a file improves it slightly, financed inside normal work; it needs tests and small commits and cannot fix structural problems. Preparatory refactoring - reshape the code so the upcoming feature becomes easy, then add it - has the highest return because the change is justified by imminent work. Branch by abstraction introduces an abstraction over the old implementation, builds the new one behind it and switches with a flag, keeping trunk releasable during multi-week structural change. Strangler fig does the same at system scale: route traffic capability by capability from the old system to the new until the old one can be retired. Dedicated capacity - a fixed share of each iteration - covers debt no feature will touch. A big-bang rewrite is last resort: it must reproduce every undocumented behaviour and pauses delivery. Choose the smallest strategy that fits.
code
text · 9 linesBranch by abstraction, in five steps:
1. callers -> LegacyStore (start)
2. callers -> Store (abstraction) -> LegacyStore
3. callers -> Store -> LegacyStore | NewStore (new built behind a flag)
4. callers -> Store -> NewStore (flag flipped, legacy idle)
5. callers -> NewStore (legacy + abstraction deleted)
Trunk is releasable after every step; each step is reversible.go deeper
Mention cleaning up as you go, always backed by tests, and that rewrites are risky.
Name preparatory refactoring, dedicated capacity, and why rewrites usually overrun.
Lay out the ladder from opportunistic to strangler fig, the safety-net prerequisites, and match the strategy to the debt's shape.
Add the organizational dimension: keeping trunk releasable during multi-quarter structural change, reversible incremental migration, funding models, and the narrow conditions under which a rewrite is defensible.
## Strategy ladder, smallest first **1. Opportunistic / camping-rule refactoring** - always leave the code better than you found it. Every change carries a small cleanup: a better name, an extracted function, a removed duplicate. Financed inside feature work, no approval needed. Requires a test safety net and a review culture that accepts mixed diffs (or separate refactor-only commits). Limit: it never crosses module boundaries and cannot repay structural debt. **2. Preparatory refactoring** - Kent Beck's "make the change easy, then make the easy change". You reshape only what the imminent feature needs. Highest return per hour, because the payoff is realized immediately and the scope is bounded by the feature. It is also easiest to justify: part of the estimate, not a separate request. **3. Comprehension refactoring** - restructuring while reading unfamiliar code so your understanding ends up in the code rather than in your head. Cheap and self-limiting. **4. Branch by abstraction** - for structural change too large for one commit where long-lived branches are unacceptable. Introduce an abstraction in front of the current implementation; migrate callers to it; build the replacement behind it; switch consumers (feature flag or gradual percentage); delete the old implementation and, if no longer needed, the abstraction. Trunk stays releasable throughout and each step is reversible. **5. Strangler fig** - the system-level analogue, named after the strangler fig vine. A facade or router sits in front of the legacy system; capabilities are reimplemented one at a time and traffic is redirected per capability, often running both and comparing outputs for a period. The legacy system shrinks until it can be switched off. Slower in total, but value ships continuously. **6. Dedicated capacity** - an explicit slice of each iteration (commonly quoted around 10-20%), a maintenance rotation, or a hardening period after a release. Needed for debt no feature will naturally touch. Risk: if it is the *only* mechanism, quality becomes someone else's job and everyday habits decay. **7. Big-bang rewrite** - build the replacement, then cut over. Justifiable only narrowly: the platform is unsupportable, the domain has fundamentally changed, or the system is small enough that a rewrite is genuinely cheap. Failure modes: undocumented behaviour is discovered only by breaking it, new value stops during the build, two systems must be maintained in parallel, and the replacement accrues its own debt before it is finished. ## Prerequisites that decide feasibility All incremental strategies depend on a **safety net**: automated tests (adding characterization tests around legacy behaviour first if none exist), fast builds, small commits, and ideally feature flags with quick rollback. Without them, refactoring is indistinguishable from risky rewriting - which is precisely why teams with no tests drift toward big-bang rewrites. ## Choosing | Debt shape | Fit | |---|---| | Naming, duplication, long functions in code you are already editing | Opportunistic / preparatory | | A module that resists an upcoming feature | Preparatory, possibly branch by abstraction | | Cross-cutting structural change (persistence layer, framework swap) | Branch by abstraction | | A whole legacy application that must keep running | Strangler fig | | Debt no feature will touch, but interest is real | Dedicated capacity | | Unsupportable platform or a trivially small system | Rewrite, with eyes open |
- What must be in place before large-scale refactoring is safe?A test safety net covering the behaviour you intend to preserve - adding characterization tests around legacy code if none exist - plus fast builds, small reversible commits, and preferably feature flags with quick rollback. Without these you are rewriting, not refactoring.
- Why do big-bang rewrites so often fail?The legacy system encodes years of undocumented behaviour and edge cases rediscovered only when they break; delivery of new value stops while the rewrite runs; both systems must be maintained in parallel; and estimates are based on the visible feature set rather than the hidden one.
- How does branch by abstraction differ from a long-lived refactoring branch?Branch by abstraction keeps all work on trunk: the abstraction lets old and new implementations coexist and be switched at runtime, so the system is releasable at every commit. A long-lived branch defers integration, accumulating merge conflicts and drift.
Renovating a house you still live in: repaint a wall while you are in the room, rebuild one room at a time behind plastic sheeting, or move out and demolish. The last is fastest on paper and the most likely to leave you homeless for a year.
saying these in an interview costs you the question
- Assuming a rewrite is the default cure for a messy system
- Calling behaviour-changing work 'refactoring' - refactoring preserves observable behaviour
- Attempting large refactoring with no tests and no rollback path
- Believing a fixed debt allocation alone solves debt without everyday cleanup habits
- Treating strangler fig as free - it requires running and maintaining two systems for a period
- Doing big refactoring on a long-lived branch that then cannot be merged