David Parnas's 1972 paper "On the Criteria To Be Used in Decomposing Systems into Modules" argues for information hiding. What criterion does it propose for drawing module boundaries, and how does that differ from encapsulation as a language feature?
answer
- KWIC: flowchart split vs. hidden-secret split
- Hide decisions likely to change
- Encapsulation = mechanism; hiding = criterion
- Getters exposing internals ≠ hiding
- Descendants: ports, Common Closure, bounded contexts
basics
~20 sParnas says each module should hide one design decision that is likely to change — not represent one step of the processing flow. Encapsulation is the language mechanism (private fields, accessors) you might use; information hiding is the design decision about what to conceal and why.
solid answer
~50 sParnas compared two decompositions of the same KWIC index program: one split by processing steps (input → shift → alphabetize → output), one split so each module hides a design decision — the storage format of lines, the choice of sorting algorithm, the input format. Both worked; only the second survived change, because in the flowchart split every module knew the shared data representation, so changing it touched everything. His criterion: **hide the decisions most likely to change behind interfaces that reveal as little as possible about them**, and choose boundaries by anticipated change, not by temporal order of execution. Encapsulation is a language-level access-control mechanism — `private`, module exports, opaque handles. You can have full encapsulation and zero information hiding (private fields plus getters/setters that expose the exact internal shape), and information hiding without language support (disciplined C headers). Parnas's idea is the ancestor of abstract data types, ports, bounded contexts, and the Common Closure Principle.
go deeper
Say that a module should hide the decision most likely to change, and that encapsulation (private/public) is only the mechanism used to enforce it. One example — hiding a storage format behind a small interface — is enough.
Reference the KWIC comparison: flowchart split vs. secret-based split, and what happens when the shared representation changes. Give a concrete encapsulation-without-hiding example (getters that leak the internal collection).
Discuss the interface as the full set of assumptions callers may rely on (including order, errors, aliasing), the prediction risk in choosing secrets, and the modern descendants: ports/DIP, Common Closure, bounded contexts, service-owned data.
Treat 'what changes together' as the primary architectural input, backed by change history rather than intuition; discuss where you deliberately expose representation for performance or diagnostics, how you keep contracts narrow over time against Hyrum's Law, and how boundaries align with team ownership and release cadence.
## The setting In 1972 David Parnas published *On the Criteria To Be Used in Decomposing Systems into Modules*. The question of the day was not *whether* to modularize but *how to choose the pieces*. The default practice was decomposition by processing step, essentially drawing the flowchart and making each box a module. Parnas used a small worked example, the **KWIC (Key Word In Context) index**: read lines, produce all circular shifts of each line, alphabetize the shifts, print them. He gave two decompositions. **Modularization 1 — by processing step:** Input, Circular Shift, Alphabetizer, Output, Master Control. Each module reads a shared data structure produced by the previous one — say, lines stored as a character array plus an index array in core memory. **Modularization 2 — by hidden decision:** Line Storage (hides *how* lines and words are stored), Circular Shifter (hides *how* shifts are represented — computed on demand or materialized), Alphabetizer (hides the sort algorithm), Input (hides the input format), Output (hides the output format). Both produce the same output with similar performance. Parnas's point is what happens when a decision changes: - Change the line storage from "all in memory" to "paged from disk": in Modularization 1 every module breaks, because every module knows the representation. In Modularization 2 only Line Storage changes. - Decide not to materialize circular shifts but compute them lazily: again localized in Modularization 2, global in Modularization 1. - Change the sort algorithm, the input encoding, the output layout: each is one module in Modularization 2. ## The criterion, stated plainly > Begin with a list of **difficult design decisions** or **design decisions which are likely to change**. Each module is then designed to **hide** such a decision from the others. Two consequences follow: 1. **Modules are not steps in the processing.** The runtime call sequence and the module structure are different structures over the same program, and conflating them is the error. (Parnas later formalized this as separate *uses*, *is-a-component-of*, and *module* hierarchies.) 2. **The interface must reveal as little as possible** about the hidden decision. An interface that returns the internal array is not hiding it. The interface should be the *smallest set of assumptions* the callers need — Parnas's later work calls this the "secret" of the module and the interface its "assumptions". ## Information hiding vs. encapsulation vs. abstraction These three are routinely conflated in interviews; distinguishing them is the whole point of this question. - **Information hiding** — a *design principle*: deliberately choosing which design decision each module conceals, driven by likelihood of change. It answers *what* and *why*. - **Encapsulation** — a *language mechanism* for enforcing access: `private`, package-private, module exports (JPMS, ES modules), opaque pointers, closures. It answers *how it is enforced*. It is enforcement without judgment. - **Abstraction** — presenting a simplified model that omits detail (an abstract data type, a port, a domain concept). It answers *what the client sees*. The combinations matter: - **Encapsulation without information hiding** (the common failure): every field private, every field with a public getter and setter of the same name and type. The internal representation is fully exposed through the interface; changing a `List` to a `Map` still ripples out. Language rules satisfied, principle violated. Likewise a "private" field exposed via a getter that returns the mutable internal collection. - **Information hiding without language encapsulation**: classic C with an opaque `struct Foo*` in the header and the definition in the .c file, or a team convention plus review. The principle can be honored with only discipline. - **Hiding the wrong secret**: you can hide a decision that never changes (wasted indirection) while exposing the one that changes weekly. The criterion is *anticipated change*, so it requires domain judgment, not mechanics. ## Interface = a contract about assumptions, not a list of methods Parnas emphasized that the interface is everything the caller may rely upon — including things not in the signature: valid call orders, error behavior, performance characteristics, thread-safety, and whether returned objects alias internal state. Anything callers can observe and depend on becomes part of the contract in practice (Hyrum's Law is the modern restatement: with enough users, every observable behavior of your system will be depended upon by somebody). So "hiding" means shrinking the *observable* surface, not just marking fields private. ## Where it shows up today - **Abstract data types and objects** — a class whose representation is a secret is exactly Parnas's module. - **Ports and adapters / DIP** — the port hides which mechanism is used. - **Common Closure Principle** (component-level): gather into a component the classes that change for the same reasons at the same times — the same criterion at package scale. - **Bounded contexts** in domain-driven design — hide a model's internal language behind a translated contract. - **Microservice boundaries** — the standard advice "a service owns its data store and no one else reads its tables" is information hiding; shared-database integration is Modularization 1 at deployment scale, and it fails the same way. - **API design and semantic versioning** — the smaller the revealed surface, the fewer breaking changes. ## Limits and trade-offs - **Prediction risk.** The criterion depends on guessing which decisions will change. Wrong guesses produce ceremony around stable things and exposure of the volatile thing. Mitigations: use history (what has changed before), start with the well-known volatility axes (I/O formats, storage, third parties, UI, pricing/regulatory rules), and prefer refactoring toward hiding once a change actually recurs rather than speculating broadly. - **Performance.** Hiding a representation may block a cross-cutting optimization that needs global knowledge of the data layout. Sometimes the honest answer is a deliberately exposed representation with a documented, narrow contract. - **Debuggability.** Deep hiding can make failures harder to trace; good diagnostics/observability must be designed as part of the interface rather than by prying it open. - **Over-hiding.** Wrapping every third-party type in your own can cost more than it saves if that dependency is genuinely stable.
- Give a concrete example of full encapsulation with zero information hiding.A class with private fields and a public getter/setter for each field, matching the field types exactly — for example `getItems(): ArrayList<Item>` returning the internal list. The representation is fully published through the interface (and even mutable through the returned reference), so any representation change is a breaking change for every caller. The language rule is satisfied; the design principle is not.
- How do you decide *which* decision a module should hide when you can't predict the future?Use evidence rather than speculation: look at what has actually changed in version history, apply the known volatility axes (storage, wire formats, third-party vendors, UI, regulatory/pricing rules), and hide those. For everything else, keep the design simple and refactor toward a hidden boundary the second time a change proves recurrent.
- How does this principle apply to service boundaries?A service's secret should be its data model and internal logic; other services get a contract, never the tables. Integrating through a shared database is decomposition-by-flowchart at deployment scale: every service knows the representation, so a schema change breaks all of them, and you get distributed coupling with none of the independence benefits.
- Parnas separated the 'uses' hierarchy from the module hierarchy. Why does that matter?Because the runtime call structure and the change-locality structure are different views of the same program. Optimizing the module structure for the call sequence (the flowchart) is exactly the failure Parnas demonstrated; the module structure should be optimized for containing change, while the uses hierarchy is kept acyclic for buildability and testability.
Two ways to organize a restaurant kitchen. Split by step — chopping station, cooking station, plating station, all sharing one giant open counter of ingredients — and any change to how ingredients are stored disrupts every station. Split by secret — the pantry owns storage and hands out prepared items through a hatch, the grill owns cooking technique — and you can reorganize the pantry overnight without a single other station noticing.
saying these in an interview costs you the question
- Saying information hiding *is* encapsulation, or that private fields + getters/setters constitute information hiding.
- Decomposing by processing step/flowchart order and calling it modular design.
- Defining the criterion as 'hide implementation' without the key qualifier — hide the decisions *likely to change*.
- Believing an interface consists only of its method signatures, ignoring call order, error behavior, aliasing, and timing that callers will inevitably depend on.
- Hiding everything uniformly, including stable decisions, and calling the resulting ceremony good design.
- Claiming a microservice split gives modularity while all services share one database schema.