"Extract till you drop" produces very small functions. As a technical leader, how do you decide when further decomposition stops paying off, and what team-level policy do you set instead of a line-count rule?
answer
- Reading cost ↔ navigation cost trade
- Test: does the name let me skip the body?
- Stop signals: paraphrase name, param explosion, no domain concept
- Locality of behaviour is the honest counter-argument
- Metrics trigger conversation; gate only objective rules
basics
~20 sStop when a helper isn't a nameable concept — when its name just paraphrases the caller, it has one caller, or it needs many parameters to rebuild the caller's context. Set a policy on naming, cohesion, and abstraction level in review; use line-count linting only as a smell trigger, never as a gate.
solid answer
~50 sDecomposition trades **reading cost** (holding a long body in working memory) for **navigation cost** (jumping between many small functions). Past a point, the second exceeds the first. Practical stopping criteria: the only honest name is a paraphrase of the caller; the helper has exactly one caller and is meaningless outside its context; extraction requires threading many parameters or a mutable accumulator; the helper introduces an abstraction that has no counterpart in the domain. Counter-forces to weigh: **locality of behaviour** (code understood together should be visible together), debuggability (deep stacks, more step-into), and the risk of an ontology of single-use names that nobody can keep straight. As a policy: mandate the *properties* — one level of abstraction per function, intention-revealing names, no hidden effects, no flag args — and enforce them in review, since they're judgement calls. Use metrics (length, cyclomatic complexity, nesting depth, parameter count) as **warning thresholds that open a conversation**, not as build-breaking gates, because hard limits reliably get gamed into worse code.
go deeper
Say that functions can be split too far and that a helper should be a nameable idea, not an arbitrary slice of lines.
Give concrete stopping criteria — paraphrase names, single context-bound caller, parameter explosion — and note that length is a symptom rather than the rule.
Articulate the reading-cost versus navigation-cost trade explicitly, cover debuggability and name-budget costs, and argue why complexity and nesting beat raw length as signals.
Present a defensible team policy: enforce properties in review, gate only objective checks in CI, use metrics as triggers with a legacy baseline plus boy-scout adoption; state the locality-of-behaviour counter-position fairly and name the legitimate long-function exceptions.
## The core trade-off Decomposition converts one kind of cognitive cost into another: - **Reading cost** (long function): you must hold many statements, variables, and branches in working memory simultaneously. It grows super-linearly with length and nesting. - **Navigation cost** (many small functions): you must jump between definitions, rebuild context at each hop, and trust names you can't see the bodies of. It grows with the *depth and fan-out* of the call tree and with how untrustworthy the names are. Decomposition is a clear win while each extracted piece is a **real concept**: reading its name genuinely lets you skip its body. It becomes a loss when the name is untrustworthy or meaningless, because then you must open the body anyway — paying navigation cost *plus* the original reading cost. > The operative test: **does the name let me skip the body?** If yes, extraction paid. If I always open it, it didn't. ## Concrete stopping criteria 1. **Name paraphrases the caller.** If the only honest name for the block restates the parent (`processOrderStep2`, `doTheRest`), you've cut below the level where concepts exist. 2. **Single caller, context-bound.** One caller isn't disqualifying by itself (naming a block is a valid reason to extract), but a single caller *plus* a name that's meaningless in isolation *plus* dependence on the caller's local state is a strong stop signal. 3. **Parameter explosion.** If pulling the block out requires five parameters or a mutable accumulator passed in and out, the block wasn't independent — you found a seam that doesn't exist. 4. **No domain counterpart.** Good extractions usually name something a domain expert would recognize or that the team already says out loud. Invented plumbing names (`applyStep3Transform`) are a smell. 5. **The chain adds no information.** A function that only forwards to another with no added meaning is pure indirection tax. 6. **Temporal-only grouping.** Extracting "the lines that happen to run together" rather than "the lines that mean one thing" produces helpers that must be called in a fixed order — you traded a long function for temporal coupling. ## The counter-arguments a principal should be able to state fairly - **Locality of behaviour.** A well-known counter-position to aggressive extraction: the behaviour you need to understand should be visible in one place. Scattering a 40-line algorithm across 12 files can make the *system* harder to understand even as each function gets simpler. This is especially true when the pieces have no independent meaning. - **Debuggability and observability.** Deep chains produce long stack traces, more step-into during debugging, more frames in profiles, and noisier flame graphs. Usually an acceptable price; occasionally decisive in code you debug under pressure. - **Name budget.** Every extracted function adds a name to the codebase's vocabulary. Names are the scarcest resource in a large codebase; spending one on a non-concept devalues the whole namespace and creates near-synonym confusion (`validateOrder` vs `checkOrder` vs `verifyOrder`). - **Refactoring churn and review noise.** Aggressive re-decomposition of stable code produces large diffs with no behaviour change, consuming review attention that should go to risk. - **Performance** is almost never a real argument on JIT/optimizing runtimes; only raise it with profile data, and document any hand-inlining so it isn't "cleaned up" later. ## Why line-count gates fail (and what to do instead) A hard rule like "no function over 20 lines, build fails" is attractive because it's mechanical — and that's exactly why it fails: - **It gets gamed.** People split at line 20 regardless of meaning, producing `partA`/`partB`, or compress logic onto longer lines, or add parameters to move code out. All strictly worse than the original. - **It measures the proxy, not the property.** A 40-line orchestration of well-named steps is fine; an 8-line function with three hidden effects is not. Length correlates weakly with the actual defect. - **It has no escape valve for legitimate long forms:** exhaustive `switch`/`when` over a closed set, table/config literals, generated code, a state machine's transition table. **Better policy shape:** 1. **Enforce the properties in review** (they're judgement calls): one level of abstraction per function; intention-revealing name; no hidden side effects; no flag arguments; CQS respected or the violation named in the verb; arguments ≤ 3 or a parameter object. 2. **Use metrics as triggers, not gates.** Report cyclomatic complexity, nesting depth, parameter count, and length; flag outliers for a *conversation* in review. Cyclomatic complexity and nesting depth correlate with defects far better than raw length. 3. **Gate only the objective things** in CI: formatting, unused code, dependency direction, architecture rules, test coverage of new code. These have no judgement component and never get gamed into worse design. 4. **Use a baseline for legacy.** Freeze existing violations, block *new* ones. This is how you adopt a rule without a mass-refactor that nobody can review. 5. **Write the rationale down, with examples from your own codebase.** Rules without worked examples get applied literally and cargo-culted. 6. **Make the boy-scout rule the mechanism of change:** improve the function you're already touching, don't schedule sweeping re-decomposition. ## How to talk about it in an interview The expected principal-level answer is *not* "small functions good" or "small functions bad". It's: name the trade-off explicitly (reading vs navigation cost), give operational stopping criteria, acknowledge the locality-of-behaviour counter-position honestly, and describe a policy that enforces properties through review while using metrics only as conversation triggers — plus a migration strategy (baseline + boy-scout rule) that doesn't stop delivery.
- Your team ships a linter rule capping functions at 20 lines. Six months later, is the code better?Typically the metric improves and the design doesn't. You see mechanical splits at the line limit (partA/partB), logic compressed onto longer lines, and parameters added purely to move code out — each strictly worse than what it replaced. Cyclomatic complexity and nesting depth are better signals than length, and even those work best as review triggers rather than build-breaking gates.
- How do you introduce these standards to a large legacy codebase without a stop-the-world refactor?Baseline the existing violations so the build stays green, block new ones, and rely on the boy-scout rule — improve what you touch. Pair that with worked examples from your own repository in the guideline document, so reviewers apply the intent rather than the letter. Reserve dedicated refactoring effort for the small number of files that concentrate change and defects.
- When is a genuinely long function acceptable?When the length is inherent and decomposition adds no concepts: exhaustive dispatch over a closed set of cases, a state-machine transition table, configuration or fixture literals, generated code, or a tight numerical kernel where inlining was measured to matter. In each case the body is uniform in abstraction level and has no internal structure worth naming.
A book's table of contents helps only if chapter titles are meaningful. Split the same book into 500 one-paragraph chapters and the contents page becomes longer than the text — you now flip pages constantly and still don't know where anything is.
saying these in an interview costs you the question
- Answering only "smaller is always better" with no stopping criteria
- Proposing a hard line-count CI gate as the team policy without acknowledging gaming
- Citing function-call overhead as a general reason not to decompose, with no measurements
- Dismissing locality of behaviour instead of engaging with it as a real counter-position
- Proposing a codebase-wide re-decomposition sprint rather than baseline plus boy-scout adoption