skip to content

Architecture Risk and Technical Debt

Architectures erode: boundaries blur, shortcuts accumulate, and a coherent design turns into a big ball of mud. You will learn risk-storming, architecture smells, the debt quadrant, and how to argue for remediation against a roadmap full of features.

part ofSoftware design & architectureoverview, primer and where to startread it →
on this pageshow

questions

6

What is technical debt, and what do the terms "principal" and "interest" mean when applied to it?

level: juniorimportance: must knowfreq 82%

answer

  1. Cunningham 1992 — code vs current understanding
  2. Principal = fix cost, interest = per-change tax
  3. Debt ≠ bug: correct behaviour, poor structure
  4. Interest only paid where you actually work
  5. Compounds: mess invites more mess

basics

~20 s

Technical debt is a shortcut in design or code that makes today faster but future changes slower. The principal is the work needed to fix the shortcut; the interest is the extra effort every change costs while it stays unfixed.

solid answer

~50 s

Technical debt (coined by Ward Cunningham in 1992) is a metaphor for structural shortcomings — missing abstractions, duplicated logic, tangled dependencies, outdated libraries, weak tests — that let you ship sooner but make every later change more expensive. The principal is the one-off remediation cost: refactoring, splitting a module, writing the missing tests. The interest is the recurring tax paid on each subsequent change: longer lead times, more defects, more coordination, slower onboarding. Debt is not the same as a bug — a system can behave perfectly and still be deeply indebted. Nor is debt automatically bad: taking it deliberately to hit a market window can be sound, provided it is recorded and repayment is planned. The metaphor's operational insight is that interest is only paid where you actually work: debt in frequently changed code is expensive, debt in code nobody touches costs almost nothing.

go deeper

for a junior

Define the metaphor, give one concrete example (e.g. copy-pasted logic in three services), and distinguish debt from a bug.

for a middle

Add the principal/interest split, name categories (test, architectural, dependency), and explain why interest depends on how often the area changes.

for a senior

Discuss where the metaphor breaks (unknown rate, no lender, no clean bankruptcy), argue for a debt register with repayment triggers, and connect debt to measurable delivery signals like lead time and change failure rate.

for a principal

Frame debt as an economic instrument: deliberate leverage with an explicit repayment plan, governed portfolio-wide. Talk about who bears the interest across teams, how to stop new architectural debt at source (automated boundary checks), and how to communicate the position to executives in cost-of-delay terms rather than moral ones.

## Where the metaphor comes from Ward Cunningham introduced "technical debt" in 1992 to explain to non-technical stakeholders why shipping fast can be rational and still costly. His original framing was about **understanding**: you ship code reflecting your current, incomplete grasp of the domain; the gap between the code and what you have since learned is the debt, and you repay it by refactoring the code to match your improved understanding. The metaphor was later broadened to cover any structural shortcoming that trades future ease for present speed. ## The two components - **Principal** — the one-time cost to put the structure right: extract the abstraction, break the dependency cycle, split the god module, add the missing test harness, upgrade the framework. Measured in engineer-days. - **Interest** — the recurring extra cost imposed on *every* piece of work that touches the indebted area: extra time to understand it, extra care to avoid breaking it, extra defects that escape, extra people who must be consulted, extra manual testing because automated tests are absent. Interest **compounds** in two ways. First, debt makes it harder to do the *next* change cleanly, so people take further shortcuts — debt breeds debt. Second, indebted code attracts more debt because a mess raises the perceived cost of doing it properly (the *broken windows* effect). ## What technical debt is not - **Not a bug.** A bug is incorrect behaviour visible to users. Debt is invisible to users; it is a property of the *internal* structure. A system can pass every acceptance test and still be indebted. - **Not "code I dislike".** Style preferences, unfamiliar-but-valid idioms, and old-but-stable code are not debt unless they demonstrably raise the cost of change. - **Not free to ignore forever.** But it *is* cheap to ignore in code with no upcoming change. That distinction is what makes prioritisation possible. ## Where the metaphor breaks down - **The interest rate is unknown and non-uniform.** Unlike a loan, you cannot read the rate off a contract; you infer it from how often the area changes and how painful those changes are. - **There is no lender and no schedule.** Nothing forces repayment, so debt can be carried indefinitely — and often is, until the team can no longer deliver. - **You cannot declare bankruptcy cleanly.** The nearest equivalent — a full rewrite — is usually the most expensive and riskiest option, not a reset button. - **Some "debt" is really an unmade design decision**, not a shortcut. Calling everything debt can excuse carelessness; that is why classification (deliberate vs inadvertent, prudent vs reckless) matters. ## Common categories | Category | Example | Typical interest | |---|---|---| | Code-level | duplication, long functions, poor naming | moderate, local | | Test | no automated coverage for a critical path | high — slows every change | | Architectural | cyclic module dependencies, god component, wrong boundaries | highest — global, hard to repay | | Dependency / platform | unsupported framework version, end-of-life runtime | spiky — low until a vulnerability or forced upgrade | | Documentation / knowledge | tribal knowledge, no decision records | rises with team turnover | Architectural debt is the dangerous kind: the principal is large, repayment is slow and risky, and the interest is paid by everyone, not just the team that took it on. ## Making it visible Useful signals that interest is being paid: rising lead time for a change of comparable size; change failure rate creeping up; estimates for "small" features in one area consistently exceeding others; the same area appearing in most incident post-mortems; new joiners taking unusually long to make their first change in a given module. Practical hygiene: keep a **debt register** (what, where, why taken, estimated principal, observed interest, trigger for repayment) rather than a wall of ad-hoc TODO comments, and link each item to the code area so it can be prioritised against real change traffic.

  • If debt is only expensive where you keep working, how would you find the areas worth repaying first?
    Overlay change frequency (from version-control history) with structural complexity or defect density. Areas that are both frequently modified and structurally poor are the hotspots paying the most interest; stable, ugly-but-untouched code can safely be left alone.
  • Is taking on technical debt ever the right call?
    Yes — when the value of shipping earlier genuinely exceeds the discounted future interest, for example validating a hypothesis or hitting a contractual date. The condition is that the debt is deliberate, recorded, bounded in scope, and has an agreed repayment trigger, rather than being accidental and forgotten.

A loan lets you buy a house now rather than in twenty years, but you pay interest every month. Technical debt buys you a release date now; you pay interest in slower delivery on every change until you repay the principal. The difference: nobody sends you a statement, so the interest is easy to pay without noticing.

saying these in an interview costs you the question

  • Calling every bug or outage "technical debt"
  • Claiming all debt must be repaid regardless of whether the code is ever touched
  • Treating debt as a purely technical matter that stakeholders need not hear about
  • Equating debt with "old code" or with code written in an unfamiliar style
  • Assuming a full rewrite is the standard way to repay architectural debt

context

open as a page

Distinguish architectural erosion from architectural drift, explain how a system becomes a "big ball of mud", and describe concrete mechanisms for preventing both.

level: seniorimportance: must knowfreq 55%

basics

~20 s

Erosion is when code violates the intended architecture — for example a layer calling something it is not allowed to call. Drift is when things are added that the architecture never covered, so the design loses coherence without any explicit rule being broken. Both, left unchecked, end in a structureless "big ball of mud".

open as a page

How do you decide what architectural debt to remediate and when, so that it competes fairly with feature work rather than being permanently deferred?

level: seniorimportance: must knowfreq 66%

basics

~20 s

Prioritise by how much the debt actually costs: fix the places you change often and that hurt every time, not the ugliest code. Estimate the fix cost and the ongoing cost of leaving it, tie remediation to upcoming features, and put it on the same backlog as features so it is compared, not deferred.

open as a page

Explain Martin Fowler's technical debt quadrant (deliberate vs inadvertent, prudent vs reckless) and how the classification changes your response.

level: middleimportance: should knowfreq 58%

basics

~20 s

Fowler classifies debt on two axes: was it taken on purpose (deliberate) or by accident (inadvertent), and was the decision sensible (prudent) or careless (reckless). The four combinations call for different responses — from scheduled repayment to coaching or changing how the team works.

open as a page

What is risk-storming as an architecture practice, how is a session actually run, and why is the first step done silently and individually?

level: middleimportance: should knowfreq 42%

basics

~20 s

Risk-storming is a group technique where people mark risks directly on the architecture diagrams. Everyone first writes risks alone and silently, then all notes are placed on the diagram, discussed, and ranked by probability and impact so the riskiest spots are visible and get mitigation owners.

open as a page

Name several architecture smells and the structural metrics used to detect them, and explain the limits of managing architecture by such metrics.

level: principalimportance: should knowfreq 38%

basics

~20 s

Architecture smells are structural warning signs above the code level: dependency cycles between modules, a god component everything depends on, a shared 'utils' hub, and features scattered across many modules. Metrics such as coupling counts, instability and change coupling help spot them, but they are indicators, not verdicts.

open as a page