skip to content

How would you make opportunistic, incremental code cleanup a durable habit across many teams, rather than something that depends on a few conscientious individuals? What mechanisms and signals would you use?

level: principalimportance: nice to knowfreq 18%

answer

  1. Mechanism over exhortation
  2. Formatter + blame-ignore for the bulk commit
  3. Baseline + fail on new findings = automated campsite rule
  4. Hotspots: churn × complexity aims the effort
  5. Team-level DORA trends, never per-person cleanup counts

basics

~20 s

Don't rely on willpower: automate style, gate only newly changed code so quality can only improve, target effort using change-frequency data, give every area an owner, and make small cleanup commits normal and welcome in review.

solid answer

~50 s

Turn the norm into mechanism. First, remove everything humans shouldn't argue about: a repo-wide auto-formatter and a bulk format commit added to a blame-ignore list. Second, install a ratchet — baseline existing static-analysis findings and fail the build only on newly introduced ones, or gate on changed lines only ("clean as you code"). That encodes the Boy Scout Rule mechanically: the codebase can only improve where it's touched. Third, direct attention with data rather than taste — hotspot analysis (change frequency × complexity), defect density, and change lead time per component tell you which mess costs money. Fourth, make it affordable: explicit capacity for cleanup, ownership for every area, and review norms where a small refactor commit alongside a fix is expected rather than questioned. Fifth, watch second-order signals: lead time, change-failure rate, onboarding time. Beware gaming — never measure people on cleanup counts.

go deeper

for a junior

Say quality should be automated where possible — a formatter and CI checks — and that small cleanups should be normal and welcomed in review.

for a middle

Add baselines that only shrink, tests required with changed behaviour, and ownership so someone is accountable for each area.

for a senior

Present the ratchet properly (baseline + new-issue failure, changed-line coverage, fitness functions) and aim effort using hotspot analysis rather than taste; describe review norms that keep cleanup commits separate.

for a principal

Frame the whole thing as mechanism design: layers of automation, funding and ownership, outcome metrics at team level with an explicit Goodhart caveat, honest leadership behaviour on deferral, root-cause fixes, and the failure mode where over-strict gates get routed around.

## The problem with "culture" Telling people to leave code cleaner works while the advocates are present and decays the moment they leave or a deadline lands. At organizational scale you want **mechanisms**: things that keep working when nobody is paying attention, and that shift the default rather than relying on individual conscientiousness. The design goal is that *doing the clean thing is the path of least resistance*. ## Layer 1 — eliminate the arguments (automation) - **Auto-formatter, repo-wide, enforced in CI and on save.** Formatting stops being a review topic entirely, freeing the cleanup budget for meaning. Land the initial bulk reformat as a single mechanical commit and add it to a blame-ignore list so line-level history stays useful — this removes the standard objection to mass formatting. - **Automated refactoring tooling and codemods.** Org-wide mechanical migrations (an API rename, a deprecated call) should be a scripted change, not thousands of manual Boy Scout acts. - **Kill the whole class where possible.** If the same smell keeps appearing, add a lint rule, a template, or an API that makes the wrong thing hard, instead of repeatedly cleaning up instances. ## Layer 2 — the ratchet (quality can only go up on touched code) This is the single highest-leverage mechanism, and it *is* the Boy Scout Rule expressed as tooling: - **Baseline / suppression file** for existing findings, so legacy debt doesn't block delivery, plus a **build failure on any newly introduced finding**. The baseline may shrink but never grow. - **Changed-lines gating ("clean as you code")**: apply the strict standard only to lines added or modified in this change. New code is held to the current bar; old code is upgraded whenever it is touched — exactly the campsite rule, automated. - **Coverage on changed lines** rather than a global percentage. Global coverage targets are gameable and stall on legacy; diff coverage is achievable and directly incentivizes adding tests for code you touch. - **Architecture fitness functions** — automated tests asserting structural rules (module dependency direction, forbidden imports, layering, cyclic-dependency checks). They prevent the *structural* rot that no amount of local tidying can repair. ## Layer 3 — aim the effort (data, not taste) Cleanup effort should follow cost, not aesthetics: - **Hotspot analysis** — rank files by change frequency × complexity (popularized by Adam Tornhill). The intersection is where debt actually charges interest; cold ugly code can be ignored. - **Defect density and incident post-mortem clustering** — which components produce the outages. - **Change lead time / cycle time per component** — where does a typical change take three weeks instead of three days. - **Knowledge concentration** — files with a single author or a bus factor of one. Publishing a small, regularly refreshed hotspot list turns "the code is bad" into "these seven files are where the money goes", which is also how you fund the work with non-engineering stakeholders. ## Layer 4 — make it affordable and legitimate - **Explicit capacity.** Some teams reserve a percentage of each iteration for improvement; others attach cleanup to the ticket that touches the area. What matters is that it isn't invisible unpaid work that dies under deadline pressure. - **Ownership.** Every module needs an accountable team. Unowned code is where broken windows accumulate, because no one has standing to fix or to say no. - **Review norms.** Reviewers should *expect* a small refactor commit alongside a fix and treat a mixed fix-plus-refactor commit as a review defect. Codify: cleanup in separate commits, behaviour-preserving, cleanup diff not dwarfing the change. - **Definition of done** that includes tests for changed behaviour and no new warnings — enforced by the gate, not by memory. - **Make deferral cheap and visible.** A one-click way to file debt from a review, feeding a register that is actually looked at during planning, so "not now" doesn't mean "never recorded". ## Layer 5 — measure outcomes, not activity Measure the *effects* you actually want, and keep the measurements away from individual performance: - delivery outcomes: **lead time for change, deployment frequency, change-failure rate, time to restore** (the DORA set) — improvement work should show up here or it isn't paying off; - **onboarding time** to first meaningful change; - **trend** of the analysis baseline and of hotspot complexity — direction matters far more than absolute values; - **warning count held at zero**, build green rate, flaky-test count trending down. **Goodhart's law is the main hazard**: any metric used as a target for individuals will be gamed. Counting refactor commits produces refactor commits, not better code. Keep these as team-level trend indicators used for deciding where to invest, never as personal scorecards. ## Layer 6 — the human system - **Leadership must model deferral honestly.** If every quarter ends with "ship it, clean up later" and later never comes, no mechanism survives. - **Anchor on the code that survives.** During migrations, explicitly tell teams not to polish modules scheduled for retirement. - **Address root causes.** Chronic mess usually traces to no tests, no ownership, or no slack. Cleaning output without changing those regenerates the mess. - **Beware over-tight gates.** If the gate blocks urgent work with cosmetic complaints, people will find the bypass and stop trusting all of it. Gates should be strict on new code and forgiving of legacy, with a documented, audited override for emergencies. ## How to answer Lead with "mechanism over exhortation", then walk the layers: automate style, install a ratchet (baseline + new-code-only gates + diff coverage + fitness functions), aim effort with hotspot and delivery data, fund it and assign ownership, measure outcomes at team level with an explicit Goodhart caveat, and close with the failure modes — unowned code, leadership that always defers, over-strict gates that get routed around, and polishing code that's about to be deleted.

  • Why gate on changed lines rather than setting a global coverage or quality target?
    Global targets on a large legacy codebase are unreachable, so they're either ignored or met by gaming (tests that assert nothing, exclusions). Changed-line gates are always achievable by the person making the change, apply pressure exactly where work is happening, and monotonically improve the codebase along its actual usage path.
  • What goes wrong if you make cleanup a tracked individual metric?
    Goodhart's law: people optimize the counter. You get cosmetic refactor commits, split PRs, and churn that raises review load without improving anything — while genuinely valuable but hard-to-count work (deleting code, saying no to a bad abstraction) is discouraged. Keep it a team-level trend used to direct investment.
  • How do you get non-engineering stakeholders to fund this?
    Translate to their units. Show that a handful of hotspot files absorb most incidents and most of the lead time, and express the ask as reducing time-to-market and change-failure rate for the specific roadmap items that route through those files — not as an abstract quality initiative.

Cities don't stay clean by asking residents to be tidy; they put bins on every corner, run collection on a schedule, and assign each street to someone accountable. The norm is real, but it rides on infrastructure.

saying these in an interview costs you the question

  • Relying on exhortation, wikis or a values poster instead of mechanisms that work unattended
  • Demanding zero findings across a large legacy codebase, which blocks delivery and trains people to bypass the gate
  • Chasing a global coverage percentage instead of coverage on changed lines
  • Making cleanup an individual performance metric, which gets gamed into cosmetic churn
  • Choosing cleanup targets by aesthetics rather than churn, incidents and lead-time data
  • Leaving modules unowned and then wondering why they rot
  • Mass-reformatting without a blame-ignore mechanism, then blaming the tool for destroyed history
  • Funding cleanup of modules already scheduled for retirement

context