skip to content

How would you introduce complexity thresholds as a CI quality gate across a large legacy codebase without stalling delivery or provoking metric gaming?

level: principalimportance: nice to knowfreq 22%

answer

  1. new code only — Clean as You Code
  2. baseline + ratchet: may shrink, never grow
  3. calibrate from your own distribution, not a blog default
  4. exempt generated/parser/migration code, justified suppressions
  5. hotspots = complexity × change frequency; never a per-team KPI

basics

~20 s

Do not fail the build on existing code. Freeze current violations as a baseline and enforce the threshold only on new and changed code, so the codebase improves as it is touched instead of demanding a big-bang cleanup.

solid answer

~50 s

Three moves. **Scope**: apply the gate to new and modified code only — SonarSource calls this 'Clean as You Code'. Existing violations are recorded as a baseline (Sonar's new-code period, detekt/ESLint baseline files) so nobody is blocked by code they did not write. **Ratchet**: the baseline may shrink but never grow. Any file touched must come out at or below the threshold, so hot files — the ones being changed, therefore the ones that matter — improve first. **Calibrate and exempt**: pick a threshold from your own distribution (often 15, sometimes 10) rather than a blog default, and exclude generated code, parsers, migrations and test fixtures, which score high for legitimate reasons. Provide a documented, reviewable suppression rather than pushing people to disable the rule. **Guard against Goodhart**: never make the aggregate score a team KPI; keep it a review trigger, and watch whether change-failure rate and review latency actually improve.

go deeper

for a junior

Say that you would not fail the build on existing code and would apply the rule to new code first.

for a middle

Add baselines, the ratchet, a calibrated threshold, and exclusions for generated code.

for a senior

Cover the rollout ladder from warn to block, suppression policy with justification, and the Goodhart risk of dashboards.

for a principal

Own the trade-offs explicitly: hotspot analysis to cover the legacy gap, baseline rot, pairing with change-failure and review-latency outcomes, and preserving reviewer authority to override the number in both directions.

## Why a naive rollout fails Turning on a per-function complexity rule across a mature codebase typically produces thousands of violations. If the build fails, teams do one of three things, all bad: disable the rule, blanket-suppress it, or spend a quarter on a risky mass refactor of code nobody was going to touch. Meanwhile the real problem — new complexity being added daily — is untouched. ## Scope the gate to change, not to history The key insight is that complexity only costs you where code is *read and modified*. Files never touched impose no maintenance cost, so cleaning them is pure risk. Hence: - **Clean as You Code (SonarQube 'new code' period).** The gate evaluates only lines added or changed since a baseline (a date, a version, or the branch point). New code must meet the standard; legacy is reported but not blocking. - **Baseline files** for tools without a new-code concept — detekt's `baseline.xml`, ESLint/`eslint-baseline`, PHPStan/Psalm baselines, `# type: ignore` inventories. They freeze the known set of violations. - **Ratchet rule:** the baseline count may only decrease. CI fails if it grows. Some teams enforce 'boy-scout' terms: if you edit a function, it must exit at or under the threshold. The effect is that improvement follows churn, which is exactly the code with the highest read frequency. ## Calibrating the threshold - Compute the actual distribution of cognitive complexity across your repo first. Set the initial line near a percentile you can live with (say the 90th), then tighten over time rather than starting at an aspirational number nobody meets. - SonarQube's default is **15** per function for cognitive complexity (`S3776`); many teams choose 10–15 for new code. Cyclomatic thresholds around 10–15 are the traditional analogue. - Distinguish **warn** from **block**. A common ladder: warn everywhere → block on new code → tighten the threshold → optionally shrink the baseline on a schedule. ## Legitimate exemptions Some code is complex for reasons refactoring will not fix: - generated code (protobuf, ORM, parser output, API clients), - hand-written lexers/parsers and state machines, - exhaustive dispatch over a large closed set (big `switch`/`when` over an enum), which is *readable* even at a high cyclomatic score, - database migrations and long test fixtures, - performance-critical hot paths where an unrolled or specialised form is deliberate. Exclude generated directories by path. For hand-written exceptions require an **inline suppression with a justification comment**, visible in review — not a silent global disable. That keeps the exception auditable and rare. ## Avoiding metric gaming Goodhart's law: once the number is the target it stops measuring anything. Concretely: - **Never dashboard complexity per team or per person.** It produces extraction theatre — six helpers named `part1..part6` — and discourages honest refactoring. - Keep the gate as a **trigger for human review**, and give reviewers explicit authority to say 'the score is fine but this is unreadable' or 'this is over the line but genuinely clearest this way, approved with a suppression'. - Pair the metric with outcome signals: change-failure rate, mean time to restore, review turnaround, defect density in high-complexity files. If those do not move, the gate is decoration. ## Rollout mechanics 1. Measure and publish the current distribution; do not act yet. 2. Turn the rule on in **warn** mode repo-wide; generate the baseline. 3. Block on **new/changed code** at a threshold most current code already meets. 4. Add path exclusions for generated/vendor code; document the suppression procedure. 5. Add the ratchet (baseline may not grow) to CI. 6. Revisit quarterly: tighten the threshold or fund targeted refactors of the top hot-and-complex files, chosen by intersecting complexity with change frequency (the 'hotspot' analysis popularised by Adam Tornhill's behavioural code analysis). 7. Tie any large cleanup to work already planned in that area, so the refactor rides existing test coverage and review attention. ## Trade-offs to state out loud - New-code-only gating leaves genuinely dangerous legacy untouched; the hotspot analysis in step 6 is what covers that gap. - Baselines rot: entries persist after the code changes shape. Regenerate periodically or they mask new problems. - Per-function thresholds push complexity into more functions and more classes; the metric cannot see whether the resulting structure is coherent. Only review can.

  • New-code-only gating leaves the worst legacy files untouched. How do you address them?
    Intersect complexity with change frequency from version-control history to find hotspots — files that are both complex and frequently modified. Those carry nearly all the real cost. Fund targeted refactoring there, ideally attached to feature work already planned in that area so it inherits test coverage and review attention.
  • A team wants a global suppression because their parser legitimately exceeds the threshold. What do you allow?
    Path exclusions for genuinely generated code, and for hand-written exceptions a narrowly scoped inline suppression carrying a justification comment, visible in review. What you refuse is a silent repo-wide disable, because it removes the signal for everything else.
  • What would convince you the gate is working?
    Not the complexity number itself. Look for the downstream signals: fewer defects in previously hot files, shorter review turnaround, lower change-failure rate. If complexity falls while those stay flat, you are probably seeing extraction theatre.

Introducing a building code does not mean condemning every existing house. New construction must comply, and any house being renovated must be brought up to code for the part you touch — which is exactly where people are living and working.

saying these in an interview costs you the question

  • Enabling the rule repo-wide as a hard failure on day one
  • Using the aggregate complexity score as a team or individual performance metric
  • Allowing silent repo-wide suppressions instead of narrow, justified ones
  • Assuming a blog-default threshold fits your codebase without measuring its distribution
  • Believing new-code-only gating alone fixes legacy risk, with no hotspot analysis
  • Never regenerating baselines, so stale entries mask newly introduced violations

context