David Parnas argued that modules should be decomposed by information hiding rather than by processing steps. What does that mean in practice, and how do you decide what each module hides?
answer
- Parnas 1972, KWIC index
- Module = secret, not step
- Interface must not name the secret
- Simulate a change: how many files?
- One secret, one owner
basics
~20 sInstead of splitting a system into one module per step of the process, split it so each module hides one design decision that is likely to change — a data format, an algorithm, a device, a policy. Everything that would change together lives behind one interface.
solid answer
~50 sParnas's 1972 paper contrasted two decompositions of the same program. The conventional one mirrored the flowchart: read input, build index, sort, format, print — each step a module. The alternative gave each module custody of a hidden secret: how lines are stored, how the index is built, how output is formatted. Both worked; only the second confined a change to one module. The practical method: enumerate the decisions you expect to change (storage representation, wire format, third-party vendor, business rule, hardware). Make each one the secret of exactly one module. The module's interface must be expressible without referencing that secret — if the interface mentions the file format, the format is not hidden. Two consequences follow. First, interfaces get defined by what callers need, not by what the implementation happens to have. Second, modules that share a secret must be merged, and a secret that appears in two modules is a design smell — that duplication is where change-time breakage comes from.
go deeper
Say modules should be split by what they hide, not by the steps of the algorithm, and give one concrete secret (the storage format).
Recount the KWIC contrast and the change-cost argument; state the interface test — if the interface names the format, it is not hidden.
Add method: derive secrets from evidence of volatility, one secret per owner, simulate changes to validate. Separate information hiding from layering/uses hierarchy and from encapsulation mechanisms.
Discuss the economics — hiding is a bet on anticipated change with real cost when wrong — and how to keep the criterion alive organizationally: change-coupling metrics, module ownership, fitness functions/ArchUnit-style enforcement, and mapping secrets onto team boundaries.
## The setting David Parnas, *"On the Criteria To Be Used in Decomposing Systems into Modules"* (1972). He implemented the same small program — a **KWIC index** (Key Word In Context: read lines, produce every circular shift of each line, sort them alphabetically, print) — two ways. **Decomposition 1 — by processing step (a flowchart made of modules):** 1. Input: read lines into a shared array 2. Circular shifter: build a shift index over that array 3. Alphabetizer: sort the index 4. Output: print 5. Master control Every module knew the shared data structures: the packed character array, the index-pair representation, the fact that shifts are precomputed and stored. **Decomposition 2 — by hidden secret:** 1. Line storage: hides *how* characters/words/lines are stored; offers `char(line, word, ch)`, `setChar(...)`, `words(line)` 2. Circular shift: hides *whether* shifts are stored or computed on demand; offers the same word-access interface over the shifted set 3. Alphabetizer: hides the sorting algorithm and whether ordering is precomputed or lazy 4. Output, Master control: as before ## The result: change cost Parnas walked through changes — pack characters differently, compute shifts on demand instead of storing them, sort incrementally instead of up front, keep lines on disk instead of in core. In decomposition 1 *each* change touched several modules, because the representation was common knowledge. In decomposition 2 each change was confined to one module because that module's interface never mentioned the thing that changed. The key line of the argument: **a module is not a subprogram; it is a responsibility assignment, and its interface should reveal as little as possible about how the responsibility is met.** ## The practical method 1. **List anticipated changes.** Formats, protocols, algorithms, storage engines, vendors, regulatory/business policies, UI, hardware, units, identifiers. Parnas's later term for the artifact is a *changeability list*; Uncle Bob's phrasing of the same instinct is "identify the axes of change." 2. **Assign each change to exactly one module as its secret.** Two modules sharing a secret means either merging them or introducing a third that owns it. 3. **Design the interface without the secret.** Write the operation signatures; if you cannot describe them without saying "the CSV row" or "the SQL row" or "the Redis key", you have not hidden the decision. This is the single sharpest test. 4. **Check by simulating a change.** Pick a change from step 1 and ask which files you would edit. More than one module means the secret leaked. ## Trade-offs and honest limits - **Guessing wrong is costly.** Hiding is anticipatory; if you hide the decisions that never change and expose the one that does, you have paid abstraction cost for nothing. This is why the technique is paired with evidence — actual change history, known roadmap, known volatility — rather than imagination. Parnas himself notes the criterion depends on foreseeing likely changes. - **Performance.** Decomposition 2 replaces field access with calls and forbids cross-module tricks. Parnas addresses this directly: the answer is not to break the interface but to allow implementation techniques (inlining, generation, co-compilation) that preserve it. - **Not the same as layering.** Information hiding says *what each module conceals*; layering says *who may call whom*. Parnas wrote separately about the **uses hierarchy** — a module's allowed dependencies — precisely because the two are independent design decisions. - **Not the same as encapsulation.** Encapsulation is the enforcement mechanism (`private`, module systems, package boundaries). Information hiding is the criterion for *what* to enforce. You can encapsulate the wrong things. ## Modern restatements you can cite - "Deep modules": John Ousterhout's *A Philosophy of Software Design* argues for a small interface over a large body of functionality — the same ratio Parnas was optimizing, stated as interface-to-implementation surface. - **Hexagonal / ports-and-adapters** and **Clean Architecture** are structural applications: the persistence and transport decisions are the secrets of adapter modules, and the domain's interface names none of them. - **Bounded contexts** in Domain-Driven Design apply the criterion at the model level: the internal model is the secret, the published language is the interface.
- How do you find the 'likely to change' decisions without guessing?Mine evidence: version-control churn (which files change together and often), incident and change-request history, the product roadmap, and known external volatility (vendor contracts, regulation, formats). Where evidence is thin, prefer a thin, cheap seam over an elaborate abstraction — you can deepen it once the change actually arrives.
- Two modules both need to know the on-disk record format. What does Parnas's criterion say?That the format is a shared secret and the decomposition is wrong. Give the format to one module that exposes record-level operations, and let the other two call it. If they truly need different views, that module can expose several interfaces — but only it knows the bytes.
- Is information hiding the same as making things private?No. `private` is a mechanism; information hiding is the criterion for what deserves to be private. A class can have all-private fields and still expose its representation through its method signatures, return types, and thrown exceptions.
Two ways to organize a restaurant kitchen: by the order of the meal (one station chops, one heats, one plates, all sharing the same pots and layout) or by what each station owns and conceals (the grill owns fire technique, the pastry station owns dough). Switch from gas to induction and the first kitchen retrains everyone; the second retrains one station.