skip to content

What is the "broken windows" theory applied to software quality, and what concrete signals would tell you a codebase is suffering from it?

level: middleimportance: should knowfreq 32%

answer

  1. Wilson & Kelling 1982 → Pragmatic Programmer
  2. Neglect signals that neglect is OK
  3. Decay is non-linear once visible
  4. Ignored warnings, skipped tests, growing baseline
  5. Zero-tolerance beats a numeric goal; ratchet baselines

basics

~20 s

The idea that visible neglect invites more neglect: one unfixed mess signals that mess is acceptable, so people stop caring and decay accelerates. Signals include ignored warnings, disabled tests, copy-pasted hacks and TODOs nobody ever removes.

solid answer

~50 s

Borrowed by Hunt and Thomas in *The Pragmatic Programmer* from Wilson and Kelling's 1982 criminology essay: an unrepaired broken window signals that nobody is watching, so more windows get broken. In code, the mechanism is social, not technical — a developer who sees a file full of hacks concludes that the local standard is "hacks are fine", and matches it. Decay is therefore non-linear: quality collapses fastest once it visibly slips. Diagnostic signals are things people have stopped reacting to: a build that emits hundreds of warnings, permanently red or skipped tests, a lint baseline that only grows, dead code nobody dares delete, ancient TODOs with no ticket, duplicated "copy this file and edit" modules, and commit messages like "fix". The countermeasure is the Boy Scout Rule plus zero-tolerance signals: keep the warning count at zero, keep the build green, delete the disabled test or fix it. Repairing the window matters more than which window.

go deeper

for a junior

Explain the analogy — visible mess makes people care less, so mess spreads — and give a couple of signals like ignored warnings or disabled tests.

for a middle

Add the social mechanism (people match the local standard they observe) and a concrete signal list, plus the fix: keep the build green, zero warnings, delete or ticket skipped tests.

for a senior

Emphasize non-linear decay and normalization of deviance, contrast zero-tolerance thresholds with numeric goals, and describe baselines/ratchets and clean-as-you-code gates as mechanisms rather than exhortations.

for a principal

Add measurement and targeting — hotspot analysis, onboarding time, change-failure rate — the caveat that the underlying criminology is contested, and the counter-pathology where over-strict gates get routed around.

## Origin **Wilson & Kelling, "Broken Windows" (*The Atlantic*, 1982)** argued that visible signs of disorder in a neighbourhood — an unrepaired broken window, graffiti, litter — signal that nobody is monitoring or cares, which invites further disorder. (The criminology claim is contested; that debate is irrelevant to the software analogy, and it's fine to say so.) **Andrew Hunt & David Thomas, *The Pragmatic Programmer* (1999)** carried it into software: *"Don't live with broken windows."* Fix bad designs, wrong decisions and poor code **as soon as you see them**; if you truly can't fix it now, board it up — comment it out, mark it, put up a `NOT IMPLEMENTED` — so it's visible that the mess is known and abnormal, not endorsed. ## The mechanism is social, not technical This is the part interviewers want. The code does not rot on its own; **people calibrate their standard to what they observe.** 1. A developer opens an unfamiliar file to make a change. 2. They read the surrounding code to infer "how we do things here" — this is a legitimate and necessary heuristic, since the local conventions usually *are* the authority. 3. If what they observe is duplication, magic numbers and dead code, they infer that the effort bar is low. 4. They match the local bar (also rationally — writing pristine code in a swamp looks inconsistent and costs more). 5. The mess is now slightly worse, and the signal to the next person slightly stronger. That feedback loop makes decay **non-linear**: near-clean codebases stay clean cheaply; once neglect is visible, quality falls off a cliff. It also explains why the Boy Scout Rule pays off asymmetrically — reversing the signal early is far cheaper than reversing it late. A second mechanism is **normalization of deviance** (Diane Vaughan's term from the *Challenger* accident analysis): a warning that never corresponds to a real failure gradually stops being treated as a warning at all. That is precisely what happens to a build emitting 800 compiler warnings. ## Concrete diagnostic signals The honest test is: *what has the team stopped reacting to?* | Signal | What it tells you | |---|---| | Build emits hundreds of warnings and nobody reads them | Warning channel is dead; a genuinely important warning will be missed | | Tests permanently disabled/skipped/quarantined, or a "known-flaky" list that only grows | The test suite no longer gates anything; green means nothing | | Static-analysis baseline (grandfathered existing issues) grows every release | The ratchet has been reversed; the gate is decorative | | `TODO` / `FIXME` / `HACK` markers years old with no ticket or owner | Debt is recorded but never scheduled; markers have become noise | | Commented-out code kept "just in case" | Team doesn't trust version control; deletion feels dangerous | | "Copy this module and edit it" as the standard way to add a feature | No shared abstraction; every bug must be fixed N times | | Wide dead zones nobody dares delete | Loss of system knowledge; fear-driven development | | Comments contradicting the code | Nobody is reading either; documentation channel is dead | | Commit messages that are all `fix`, `wip`, `.` | History has stopped being a communication medium | | Long-lived branches that can't be merged | Integration pain has exceeded the pain threshold | | Onboarding time trending up; "only Dana can touch billing" | Comprehension cost, the real interest payment on debt | Note that several of these are **process** windows, not code windows — a permanently red CI pipeline is the loudest broken window a team can have. ## Countermeasures - **Zero-tolerance thresholds beat quality goals.** "Zero warnings" is enforceable and self-sustaining because any new warning is instantly visible. "Fewer than 200 warnings" is unenforceable — no individual warning is ever the problem. The same argument favours treating warnings as errors in CI. - **Ratchet, don't reset.** Baselines/suppression files for legacy issues are fine, but the count must be allowed only to shrink. "Clean as you code" gates — fail the build only on newly changed lines — mechanize the Boy Scout Rule. - **Fix or delete, never disable-and-forget.** A skipped test must carry a ticket and an expiry. Better still, delete it: a test that doesn't run is worse than no test, because it advertises coverage that doesn't exist. - **Board it up if you can't fix it.** Explicitly mark known-bad regions (a `HACK(TICKET-123):` note, an architecture decision record) so the next reader knows the state is recognized and abnormal. - **Repair visibly and early.** Because the mechanism is signalling, a small conspicuous repair — deleting the dead module, clearing the warnings, greening the build — changes behaviour out of proportion to its size. - **Direct effort with data.** Hotspot analysis (files with both high change-frequency and high complexity, popularized by Adam Tornhill) tells you *which* windows actually cost money, so repair effort isn't spent on cold code. ## Honest caveats - The original criminology research is disputed and its policing applications were harmful; the software use is an **analogy about signalling**, and it's a mark of maturity to say that rather than cite it as proven science. - Not all mess is a broken window. Code that is stable, isolated, working and rarely touched is cheap to leave alone — ugly is not the same as costly. The interest rate on debt is proportional to how often you must touch it. - Zero-tolerance taken too far becomes its own pathology: teams that block delivery on cosmetic rules train people to route around the gates entirely. ## How to answer Give the origin in one line, then spend your time on **the social mechanism** (people calibrate to observed standards → non-linear decay) and **observable signals** (what has the team stopped reacting to). Close with the countermeasure pairing: Boy Scout Rule for continuous repair plus zero-tolerance/ratcheting gates so the signal can't drift back.

  • Why is "zero warnings" easier to sustain than "under 200 warnings"?
    Because zero makes every new warning individually visible and attributable, so it gets fixed by whoever caused it. Any non-zero threshold makes each additional warning invisible in the noise, and the number only ratchets upward.
  • Your codebase already has 4,000 static-analysis issues. Where do you start?
    Freeze the bleeding first: baseline the existing issues and fail the build on any new one, so quality can only improve on touched code. Then use hotspot analysis — high churn combined with high complexity — to pick which legacy areas are actually costing money, rather than fixing issues in cold code.
  • Is a broken window always worth fixing?
    No. Cost is proportional to how often the code is read and changed. A gnarly but stable, isolated, rarely-touched module charges almost no interest; fixing it is a vanity project compared with cleaning a file the team edits weekly.

One piece of litter in a spotless stairwell gets picked up; in a stairwell already full of litter, nobody picks up anything and everyone adds. Same person, same effort — different signal.

saying these in an interview costs you the question

  • Presenting the criminology theory as settled science instead of an analogy about signalling
  • Claiming messy code is bad primarily for runtime performance rather than comprehension and change cost
  • Arguing every ugly file must be cleaned, ignoring that stable rarely-touched code charges little interest
  • Setting a non-zero warning threshold and expecting it to hold
  • Disabling or quarantining failing tests indefinitely instead of fixing, deleting, or ticketing them
  • Treating a permanently red CI pipeline as a tooling annoyance rather than the loudest broken window on the team

context