You decomposed a payroll run top-down into nested subroutines — what signals that the decomposition has stopped paying for itself?
answer
- a split that only relocates code
- parameters forwarded but never read
- a name that paraphrases the body
- cut along control flow, not data
- a change that re-roots the tree
basics
~20 sDecomposition stops paying when a split only relocates code: routines called once whose names restate their bodies, parameters threaded through layers that never read them, and any new field requiring an edit at every level.
solid answer
~50 sA useful split names a step you could test, reuse or reason about alone. The signals that a split has stopped doing that are concrete: a routine whose name is a paraphrase of its body, a routine that cannot be understood without its single caller, parameters threaded through intermediate layers that never touch them, and a one-field change that must be edited into every level of the chain. Behind all four sits the structural weakness of a strictly top-down method — it decomposes by **control flow**, so the tree encodes the order of today's steps, and when requirements move the tree has to be re-rooted rather than extended. The practical rule is to keep splitting while each split buys a name, a test boundary or a reuse, and to stop when it buys only a smaller file.
code
pseudocode · 12 linesfunction runPayroll(date)
return payEveryone(date, loadEmployees())
function payEveryone(date, employees) // date: forwarded, never read
for each e in employees
call payOne(date, e)
function payOne(date, employee) // date: forwarded, never read
call writeSlip(date, computeNet(employee))
function writeSlip(date, amount) // the first routine that uses date
call emit(date, amount)go deeper
Know that breaking a long routine into named steps helps a reader, and that the name has to say something the body does not — otherwise the reader is just visiting two places for one fact.
Name the concrete signals: forwarded parameters, names that paraphrase bodies, a one-field change that edits every level. Tie them back to the axis the decomposition was cut along.
Explain why a control-flow cut re-roots under change and a data-affinity cut usually does not, and show how you would restructure a chain whose intermediate layers exist only to forward values.
Decide what the codebase's decomposition axis is and where the boundary between the two sits, so that routine reviews agree on what a legitimate split is instead of relitigating it per pull request.
## What top-down decomposition buys Stepwise top-down decomposition is the classic procedural design method: state the run as one step, break that step into a handful of sub-steps, and repeat until each leaf is small enough to write directly. It gives you three real things — a vocabulary of named steps, a reading order that matches the run, and small units you can verify one at a time. For a batch process with a genuinely sequential shape, it is still the fastest way to a correct program. It is a method with a cost curve, though, and the cost curve turns. ## The signals it has stopped paying - **The name restates the body.** A routine called `computeNetAfterDeductions` whose body is one subtraction has not abstracted anything; the reader now visits two places to learn one fact. - **Pass-through parameters.** A value is accepted by three layers that never read it, purely to reach the fourth. The intermediate signatures now advertise a dependency they do not have. - **A change touches every level.** Adding one field to the output requires editing the top routine, each intermediate and the leaf. That is the plumbing tax, and it scales with depth. - **The unit is not independently meaningful.** If a routine cannot be described without "it's the part of the caller that…", it is a bookmark in the caller, not a step. - **Single-call-site routines multiplying.** One such routine is often fine — it names a step. A layer of them, each called once, means the tree is recording the author's editing history rather than the problem's structure. | A split that pays | A split that only relocates code | |---|---| | Names a step someone would say aloud | Name paraphrases the body | | Can be tested with its own inputs | Needs the caller's context to make sense | | Takes only what it uses | Takes parameters it merely forwards | | Survives a requirements change | Has to be re-cut when the order changes | ## Why the tree has to be re-rooted The deeper issue is what the decomposition was *cut along*. A strict top-down pass decomposes by **control flow**: the tree mirrors the order in which things happen today. That is precisely the part of a program most likely to change. When the run gains a step in the middle, or a step must now happen per employee rather than per batch, the branch structure that encoded the old order no longer fits, and the repair is not a new leaf but a new root. Decomposing along the **data** instead — which routines touch the employee record, which touch the ledger, which touch the output file — yields groupings that change only when the data changes, which is less often. A practical procedural design uses both: control flow for the top two levels, where the run order really is the subject, and data affinity below that, where routines should cluster around what they read and write. ## A stopping rule you can apply in review 1. Say the routine's name out loud as a step in the process. If you cannot, it is not a step. 2. Ask what inputs it would need to be tested alone. If the honest answer is "whatever the caller happens to have", do not split it out. 3. Count the parameters it forwards without reading. More than one is a signal the chain is too deep, not that the names are wrong. 4. Ask what a plausible requirements change would do to the tree. If a common change re-roots it, you cut along the wrong axis. One caveat worth stating in an interview, because it is where candidates overcorrect: a routine with exactly one call site is not automatically a defect. If its name tells a reader what the block does and lets the caller read as a sequence of steps, it has paid for itself with legibility alone. The defect is a layer of such routines that each add a name and nothing else.
- Is a routine with exactly one call site always a sign of over-decomposition?No. If its name lets the caller read as a sequence of named steps, legibility alone pays for it. The defect is a whole layer of single-use routines that add a name and nothing else, or one whose name merely paraphrases its body. Judge the split by whether the name tells the reader something the body would not.
- What does decomposing along the data give you that a strict top-down cut does not?Groupings that change when the data changes rather than when the running order changes. Routines that touch the employee record cluster together, as do those that touch the ledger and those that touch the output. Since running order shifts more often than record structure, that cut survives more requirement changes and confines their edits.
- Two intermediate layers forward four parameters each. What does that tell you?That the chain is deeper than the problem. Either the leaf work belongs closer to where its inputs already are, or the forwarded values belong in one record passed as a unit so the layers carry one parameter instead of four. Long forwarding lists are a depth signal, not a naming signal.
saying these in an interview costs you the question
- Says more subroutines always means a better decomposition
- Claims every routine with one call site must be inlined back
- Thinks a long parameter list is always just a naming problem
- Says a top-down decomposition never needs revisiting once agreed
- Uses routine line count alone as the signal to split