How do you decide what architectural debt to remediate and when, so that it competes fairly with feature work rather than being permanently deferred?
answer
- Interest ≈ change frequency × pain per change
- Hotspots: version-control churn × complexity/defects
- One backlog, business framing, cost of delay ÷ duration
- "Make the change easy, then make the easy change"
- Stop the bleeding before repaying; no big-bang rewrite
basics
~20 sPrioritise by how much the debt actually costs: fix the places you change often and that hurt every time, not the ugliest code. Estimate the fix cost and the ongoing cost of leaving it, tie remediation to upcoming features, and put it on the same backlog as features so it is compared, not deferred.
solid answer
~50 sRank debt by its **interest rate**, not its ugliness. Interest ≈ how often the area changes × how much pain each change causes; hotspots are found by overlaying version-control change frequency with complexity, defect density and incident history. Compare that recurring cost against the principal (fix effort) and against the value of competing features — cost of delay divided by duration is a workable common currency. Then time the work: remediate **just ahead of** the roadmap that will touch that area ("make the change easy, then make the easy change"), so repayment is justified by imminent work rather than aesthetics. Tactically, mix continuous small improvement inside feature work (boy-scout rule), explicitly scoped enabler items on the same backlog with business-language justification, and dedicated increments only for structural moves too large to slice. Stop the bleeding first with fitness functions so new debt is not created faster than you repay. Avoid unfunded "debt sprints" and big-bang rewrites.
go deeper
Say that not all debt is worth fixing, that the code you change often matters most, and that debt work should be on the same backlog as features so it gets prioritised rather than forgotten.
Introduce principal versus interest, hotspot analysis from version-control history crossed with complexity or defects, and the boy-scout rule plus explicitly scoped enabler items ahead of related features.
Argue the economics — cost of delay versus principal, evidence from lead time, change failure rate and incidents — and the timing rule of repaying just ahead of roadmap work. Insist on stopping the inflow with fitness functions, defined exit criteria, and strangulation over rewrite.
Operate a portfolio: a debt register with principal, evidence of interest, triggers and owners; formal acceptance decisions with review dates; alignment of boundaries with team ownership so debt is not re-created; and a narrative to executives that frames remediation as risk and throughput management rather than engineering preference.
## The core question: which debt is actually expensive? Most debt backlogs are sorted by how offensive the code looks. That is the wrong ordering, because interest is only paid where work happens. Two inputs matter: - **Principal** — effort to remediate. Estimate it like any other work, ideally after a timeboxed spike, because architectural fixes are notoriously underestimated. - **Interest** — recurring cost of *not* fixing it, roughly **change frequency × pain per change**. Concretely: additional hours per change, defects escaping from that area, incidents traced to it, review and coordination overhead, onboarding time. **Hotspot analysis** operationalises this: take change frequency from version-control history (how many commits touched each file or module over the last N months) and cross it with a structural proxy (size, complexity, defect count). Frequently-changed *and* structurally bad areas are where the interest is being paid. This is the core of behavioural code analysis, popularised by Adam Tornhill. Two useful companions: **change coupling** (files that repeatedly change together despite no logical relationship — a boundary in the wrong place) and defect/incident maps. This analysis routinely overturns intuition: the module everyone complains about may be stable and rarely touched (low interest — leave it), while an unremarkable-looking module in every second commit is quietly taxing the team. ## Making the comparison fair Debt loses to features when the two are argued in different languages — features in revenue, debt in aesthetics. Fixes: 1. **One backlog.** Remediation items sit alongside features and are prioritised by the same people using the same criteria. A separate "tech backlog" is a graveyard. 2. **Business framing.** State the item as an outcome: "payment changes currently take three weeks and fail one deploy in four; this reduces both" — not "refactor the payment package". 3. **A common currency.** Cost of delay divided by duration (from Donald Reinertsen's work on flow economics) ranks heterogeneous items: a remediation whose payoff is a permanent 30% cut in lead time for a high-traffic area can legitimately beat a mid-value feature. Coarse estimates are fine; the discipline of stating the value is what matters. 4. **Name the trigger, not the ideal.** "We must repay this before the multi-region work starts next quarter" is actionable; "this should be cleaner" is not. 5. **Show the trend.** Lead time, change failure rate, and defect density per area over time make interest visible to non-engineers without asking them to read code. ## Timing: remediate just in time, not just in case Kent Beck's "make the change easy, then make the easy change" is the timing rule. Repay debt in the area the roadmap is about to hit, because: - the fix is justified by imminent, funded work; - you now know which direction the design must flex, so you refactor toward a real requirement rather than a guessed one; - the payoff is realised immediately instead of being speculative. Corollary: debt in areas with no roadmap presence is usually *correctly* deferred — while remaining recorded so the position is known if plans change. ## Tactics, and when each fits - **Boy-scout rule / opportunistic refactoring** — leave each touched area better. Free and continuous, but only reaches code you happen to open and cannot move a module boundary. - **Enabler / preparatory items on the backlog** — explicitly scoped, estimated, visible work preceding a feature. The workhorse for medium debt. - **Fixed capacity allocation** — reserve a standing share of capacity (teams commonly use something in the 10–20% range) for structural work. Sustains attention, but risks becoming an unexamined tax if items are not justified individually. - **Dedicated remediation increment** — for structural moves that cannot be sliced (splitting a shared database, extracting a bounded context). Needs an explicit business case and a defined end state, otherwise it drifts. - **Strangler fig** — incrementally route functionality to a new implementation while the old shrinks. Preferred for large-scale replacement because value keeps flowing and risk stays incremental. - **Stop-the-bleeding first** — before repaying, add fitness functions (automated dependency, boundary and cycle checks, budgets) so new debt of the same class cannot be added. Repaying while the inflow is uncontrolled is bailing without plugging the hole. ## What to avoid - **The mythical debt sprint** scheduled "after this release" that never arrives; it also batches structural change into a single high-risk drop. - **Big-bang rewrite** — freezes feature delivery, carries unknown parity risk, and typically reproduces the same decay because the forces that created it were never addressed. - **Refactoring without a target design** — motion without direction produces churn and merge pain. - **No exit criteria** — remediation without a defined end state expands indefinitely and destroys trust in the next request. - **Repaying only what is easy** — cheap cosmetic wins while the expensive structural item causing the actual interest stays untouched. ## Governance at scale Maintain an architectural debt register with, per item: location, origin (deliberate/inadvertent, prudent/reckless), estimated principal, evidence of interest (hotspot data, incidents, lead-time impact), repayment trigger, and owner. Review it on a regular cadence with the same stakeholders who prioritise features, and be explicit about items formally **accepted** — accepted debt with a named owner and review date is a legitimate position; forgotten debt is not.
- A stakeholder asks why they should fund a refactor that delivers no visible feature. What do you say?Express it as cost and risk, with evidence: this area appears in 40% of commits and 60% of production incidents; changes there take three times longer than comparable work elsewhere; next quarter's roadmap has four items in it. The refactor costs roughly X weeks and is expected to cut lead time and change failure rate for that area, so the roadmap work lands sooner and more safely. Then offer the honest alternative — defer, accept the slower delivery, and revisit at a named date.
- How do you stop a remediation effort from expanding without end?Define the end state and the measure before starting (for example: no cycles between these modules; this component no longer reads that schema directly; lead time for changes here below N days), timebox the work, slice it so each slice is independently shippable and reversible, and encode the achieved state as a fitness function so it cannot regress. If the timebox expires without the criteria being met, that is a decision point, not an automatic extension.
- Is a fixed percentage of capacity for technical work a good policy?It is a reasonable default for sustaining attention and avoids relitigating every small improvement, but it degrades if the allocation stops being justified item by item — it becomes an untracked tax that funds whatever engineers find interesting. It also cannot fund large structural moves, which need their own business case. Best used alongside evidence-based prioritisation, not instead of it.
A landlord does not repair the prettiest apartment first; they repair the one that is rented out every week and generates the most complaint calls. The ugly, empty unit in the basement can wait — until someone signs a lease on it, which is exactly when the repair becomes urgent.
saying these in an interview costs you the question
- Prioritising the ugliest code rather than the code where interest is actually paid
- Keeping a separate 'technical backlog' that product never prioritises
- Arguing for remediation in aesthetic or moral terms instead of cost, risk and lead time
- Proposing a big-bang rewrite as the default remedy for accumulated debt
- Starting remediation without defined exit criteria or a target design
- Repaying debt without first preventing new debt of the same class from being added
- Assuming all identified debt must eventually be repaid, even in code that is never touched