A teammate lowers a function's cognitive complexity from 24 to 8 by extracting six private helpers. How do you judge whether the code actually got easier to read?
answer
- metric is per-function; comprehension is global
- Goodhart: measure becomes target
- name must let you skip the body
- 6 params / shared fields = wrong seam
- score improves as a side effect, not a goal
basics
~20 sCheck whether each helper's name lets you skip its body. If you must open all six to understand the original function, the complexity just moved and reading now costs six jumps. Good extraction removes detail; bad extraction only relocates it.
solid answer
~50 sCognitive complexity is measured **per function**, so moving code into helpers always lowers the parent's score — the metric can be satisfied without improving anything. Judge the result by comprehension instead. Good signs: each helper has a domain-meaningful name that fully summarises its body, so the caller reads as a coherent narrative; helpers take few parameters, are pure or effect-explicit, and each is a concept that could be tested and named independently. Bad signs: names like `doStep2`/`handleRest`; helpers that take six parameters or a bag of state because the extraction cut across a data dependency; helpers used exactly once that only exist to satisfy the linter; shared mutable fields introduced so the pieces can communicate — that converts local variables into object state, which is strictly harder to reason about. The honest test: can a newcomer read the parent function alone and correctly explain what it does?
go deeper
Say that shorter functions are not automatically clearer, and check whether the helper names explain what they do.
Explain that the metric is per-function so extraction always lowers it, and give the naming and parameter-count tests.
Name Goodhart's law, discuss shared mutable state and yo-yo reading, and propose alternatives (guard clauses, dispatch tables, splitting the type).
Set policy: complexity as a review trigger on new code, never a KPI; invest in domain vocabulary and boundary validation so scores fall as a by-product of better design.
## Why the question arises Cognitive complexity, cyclomatic complexity and most linter thresholds are computed **per function**. Extracting code into another function therefore *always* reduces the parent's score, and often reduces the aggregate too, because the nesting penalty is charged from zero again inside each helper. That makes the metric trivially gameable: it measures a local property while comprehension is a global one. This is a textbook instance of **Goodhart's law** — when a measure becomes a target, it stops being a good measure. ## Extraction that genuinely helps: chunking Human working memory holds only a handful of live facts. Extraction helps when the helper's **name replaces its body** in the reader's head — one remembered concept instead of ten remembered statements. Conditions for that: 1. **The name is a true abstraction** in the domain's vocabulary (`isEligibleForRefund`, `applyLoyaltyDiscount`), not a positional label (`part2`, `helper`, `processData`). 2. **The interface is narrow.** Few parameters, one return value. A helper needing six arguments means you cut across a cohesive computation and the caller's context travels with the call. 3. **Effects are explicit.** A helper that quietly mutates caller state, or that reads/writes new mutable fields added just to pass data, replaces local variables with object state — a strictly wider scope and harder to reason about. 4. **It is independently meaningful.** You could name it in a design conversation, test it alone, and plausibly reuse it. Single-use is fine if the concept is real. 5. **Cohesion, not just size.** The split follows a seam in the logic (validate / compute / persist), not an arbitrary line count. ## Extraction that hurts - **Shotgun/yo-yo reading.** Understanding the whole now requires jumping through six definitions. Each jump is a context switch and leaves an unresolved item in working memory. The parent got shorter; total effort went up. - **Temporal coupling made invisible.** If the helpers must be called in a specific order, and that order is enforced only by the parent's line sequence plus shared fields, the extraction has hidden a real constraint. - **False reuse pressure.** A helper created for one caller invites a second caller with slightly different needs, which grows parameters and flags until it serves nobody well. - **Parameter/flag creep.** `applyRules(order, ctx, true, false, null)` is worse than the inline block it replaced. ## How to judge, concretely - **Read-aloud test:** read only the parent function. Can you state what it does and in what order, without opening a helper? If yes, the abstraction holds. - **Naming test:** could you have named each helper *before* writing it? Names that only made sense after the split usually describe position, not meaning. - **Parameter test:** count arguments and shared mutable fields. Growth means the seam was wrong. - **Diff-review test:** in review, does the change read like a design improvement independent of the linter? If the only justification is 'the rule was failing', that is the tell. - **Alternatives to consider instead:** guard clauses to flatten nesting (removes complexity rather than moving it); replacing a conditional chain with a lookup table or polymorphic dispatch; splitting the *type* rather than the function when the function is long because the class does too much. ## The mature position A high complexity score is a **reliable smell and an unreliable diagnosis**. Use it to trigger a look, then fix the underlying design — usually a function doing several jobs, a missing domain concept, or validation that belongs at a boundary. Judge the fix by human comprehension, in review; the score should improve as a *side effect*. Teams that invert this — treating the number as the goal — end up with codebases that pass every gate and still take a week to onboard into.
- Is a single-use private helper ever justified?Yes, whenever the name is a real concept that lets the caller be read without opening it. Reuse is not the point of extraction — naming is. The failure mode is single-use helpers named for position rather than meaning.
- What would you do instead if the function is long because it validates, computes and persists in one place?Split along that natural seam, ideally pushing validation to the boundary and persistence behind a port, so the middle becomes a pure computation. That reduces complexity by relocating responsibility, not just code.
- How do you keep a complexity rule useful rather than gameable?Apply it to new code, treat it as a review trigger rather than a scored target, require the reviewer to judge readability regardless of the number, and never turn the aggregate score into a team KPI.
Tidying a room by shoving everything into six unlabelled boxes lowers the visible clutter to zero. If the boxes are labelled 'winter clothes' and 'tax documents', you genuinely find things faster. If they say 'box 3', you now have the same mess plus six lids to open.
saying these in an interview costs you the question
- Treating 'the linter is green' as proof the refactor helped
- Believing extraction is always an improvement regardless of naming or parameter count
- Introducing mutable fields so extracted helpers can share state, then calling it cleaner
- Optimising the aggregate score as a team KPI
- Refusing all extraction because 'it causes jumping', and leaving 200-line functions intact