One school derives a program's abstraction levels by refining a top-level statement downward, as in Wirth's stepwise refinement; another builds vocabulary upward until the top level reads as the problem itself, as with Forth word factoring, Lisp macros and Smalltalk protocols. Compare the two, and say when each produces the wrong levels.
answer
- Wirth 1971: refine the top statement downward
- middle levels = artefacts of the first guess
- Forth: factor mercilessly, top word = spec
- Lisp macros raise the language; debugger shows expansion
- combinators: level mismatch = type error, but a reader cliff
basics
~20 sTop-down refinement derives every intermediate level from one decomposition, so the levels freeze the first guess. Bottom-up vocabulary building (Forth words, Lisp macros, Smalltalk protocols) yields reusable levels but risks a private dialect and weak tooling. Mature designs alternate.
solid answer
~60 sBoth produce a hierarchy of levels; they disagree about where the levels come from. - **Wirth's stepwise refinement (Pascal, 1971)**: each step names sub-steps, and nested procedures give them a home, so levels follow one program's call tree. Fails when requirements shift: the middle levels were derived from the top-level guess and have no meaning of their own. - **Forth (Moore's "factor mercilessly")**: words are a line long and the top word reads as the specification. Price: an idiolect meaningful only inside that program, with a stack discipline that gives no types to check the level. - **Lisp/Racket macros**: raise the language toward the problem so the top level is a DSL. Price: debuggers and readers see the expansion, not your level — which is why Racket invests so heavily in syntax objects and source tracking. - **Smalltalk**: vocabulary lives on objects and is browsable in the image, not in syntax — legible live, less legible as text. - **Haskell/Scala combinators**: levels are typed values, so a level mismatch is a type error; price is the cliff for readers who do not know the algebra.
code
forth · 4 lines: CHECK-FUNDS ( amt acct -- amt acct ) 2DUP BALANCE < ABORT" insufficient" ;
: WITHDRAW ( amt acct -- ) CHECK-FUNDS DEBIT LOG-TXN ;
\ every word is a line, testable alone at the interpreter,
\ and meaningful only inside this program's vocabularygo deeper
Know both directions exist: break a big task into steps, or build small named operations you then combine.
Explain why refinement's middle levels rarely get reused, and give one concrete example of a language whose culture builds vocabulary upward.
Argue the alternation in practice — sketch down, build up, rewrite the top — and identify the scaffolding levels that should be deleted after convergence.
Weigh the tooling and onboarding tax of a DSL or combinator algebra against its expressive gain, and decide per team and domain stability which route to standardise on.
## Two ways to get a hierarchy A well-levelled program reads as policy at the top and mechanism at the bottom, with each layer expressed in the vocabulary of the one above it. There are two established routes to that shape, and they have different failure modes. **Top-down refinement.** Wirth's 1971 "Program Development by Stepwise Refinement" starts from a one-line statement of what the program does, then rewrites each unexplained phrase as a small routine that in turn contains unexplained phrases, until everything bottoms out in language primitives. Pascal supported it directly through nested procedures: each refinement gets a home inside its parent and no wider visibility. The method is teachable, produces a documented decision tree, and is still the right instinct when the problem is genuinely understood in advance — a known algorithm, a specified protocol. Its weakness is that every intermediate level is an artefact of the first decomposition. The middle routines exist because the top-level split needed them, not because they name anything meaningful; they take the parameters the split happened to require. Change the top-level requirement and the intermediate levels do not adapt — they are re-derived. That is why pure top-down design ages badly under changing requirements and yields little reuse: nothing in the middle was designed to be called by anyone but its one parent. **Bottom-up vocabulary building.** The other tradition grows the language upward until the problem can be stated in it. Forth is the extreme case: Chuck Moore's doctrine is to factor mercilessly into words that are often a single line, so the final top-level word reads as the specification and every word below it is independently testable at the interpreter. Lisp does it with macros — you are not writing a program in Lisp, you are extending Lisp toward the problem and then writing three lines in the extended language; Paul Graham's "programming bottom-up" is the essay-length version. Smalltalk puts the vocabulary on objects instead of syntax: Kent Beck's composed-method style plus method categories means the level structure is browsable in the live image. Haskell and Scala do it with typed combinators — parser combinators, `for`-comprehensions over monads — where a new level is an ordinary value with a type. Bottom-up levels are reusable by construction, because each word, macro or combinator was defined to stand alone. That is the reuse top-down rarely produces. ## When bottom-up produces the wrong levels The dialect problem. A Forth program's vocabulary is meaningful only inside that program, and the stack discipline offers no types to catch a level mismatch; a newcomer must learn a private language before reading anything. Lisp macros carry a specific tooling cost: stack traces, steppers and error messages naturally speak in terms of the expansion rather than your surface syntax, so the abstraction you built is exactly the abstraction the debugger removes — the reason Racket spends so much design effort on syntax objects, source locations and error-message contracts. Smalltalk's vocabulary is legible in the image and much less legible as a diff or a text file, which is a real cost for review-centric teams. Typed combinator libraries impose an abstraction cliff: the top level is beautifully short, and unreadable until you know the algebra it is written in. The second failure is speculative generality: vocabulary built before the problem is known, producing a rich language for a program nobody ended up writing. ## How practitioners actually work Alternate. Sketch top-down to discover which operations the problem needs; build those operations bottom-up as things that stand alone and are tested alone; then rewrite the top level in the vocabulary you now have and watch several intermediate levels from the sketch disappear because they were scaffolding. The tell that you have converged is that the top-level routine names steps a domain expert recognises, and every name below it survives being read out of context. ## The judgment call Choose refinement when the decomposition is dictated by a known algorithm or specification and reuse is not the goal; choose vocabulary building when the same nouns and verbs recur across many use cases and the top level is expected to change often. Weigh the tooling tax honestly: a macro-based DSL or a combinator algebra is an excellent level structure with a poor debugging story, so it pays where the domain is stable and the team is small, and costs where onboarding and incident response dominate.
- Why does pure top-down refinement tend to yield little reusable code?Because every intermediate routine is created to serve one parent's decomposition and takes whatever parameters that split required, so it names no independently meaningful concept. Reuse requires a unit designed to stand alone, which is what bottom-up construction produces by definition. Teams usually recover reuse afterwards, by noticing repeated middle levels across programs and promoting them into a real vocabulary.
- What is the tooling tax on macro-built or combinator-built abstraction levels?The level you invented is the level the tools tend to erase: stack traces and steppers naturally report the expansion or the desugared combinator chain, and type errors are phrased in the library's internal vocabulary rather than yours. Ecosystems that take DSLs seriously invest heavily in that gap — source-location tracking, custom error messages, contract systems — and it is a real cost when weighing a DSL against plain functions.
saying these in an interview costs you the question
- Presenting top-down and bottom-up as a correctness question rather than a tradeoff about where the levels come from
- Claiming stepwise refinement produces reusable components, when its middle levels exist only to serve one decomposition
- Ignoring the debugging and onboarding cost of a macro-built or combinator-built vocabulary
- Building the vocabulary before the problem is understood, producing a rich language for a program never written