At architecture scale, when does replacing type conditionals with a polymorphic subtype hierarchy make a system HARDER to evolve, and what alternatives would you reach for?
answer
- expression problem = who pays for the new axis
- multiple axes → composition, not class-per-combination
- wrong axis = shotgun surgery, worst outcome
- enumeration escapes into DB/API/UI/analytics
- new operation across team-owned impls = migration
basics
~20 sIt hurts when the variation isn't really by type, when several variation axes multiply into too many classes, when the variant list is duplicated in databases, APIs and configs, or when adding new operations means changing code owned by other teams. Then prefer data-driven rules, composition, or closed types with exhaustive checks.
solid answer
~50 sPolymorphic hierarchies optimize for one thing: **adding variants without editing existing code**. They become liabilities when that isn't the dominant change. Failure modes: (1) *combinatorial explosion* — variation on two or more independent axes (region × channel × tier) yields a class per combination unless you switch to composition/Strategy; (2) *the wrong axis* — you subtyped by product but change arrives by jurisdiction, so every change is shotgun surgery; (3) *the enumeration escapes the code* — DB check constraints, API enums, UI dropdowns and analytics all list variants, so "one new class" is a multi-system release; (4) *operation growth across ownership boundaries* — a new interface method must be implemented in modules other teams own or that ship as plugins, which is a versioning and coordination problem, not a coding one; (5) *reviewability loss* — logic thinned across dozens of files makes cross-variant consistency (pricing, compliance) unauditable. Alternatives: data-driven rules/tables, composition of capabilities, sealed types with exhaustive matching, and default methods or versioned interfaces at ownership seams.
code
text · 12 lines// Combinatorial explosion: class per cell (region x channel x tier)
class EuOnlineGold ... ; class EuOnlineSilver ... ; class UsRetailGold ... // n*m*k
// Composition per axis: n + m + k pieces, each axis evolves alone
class OrderPricing(
taxRule: TaxRule, // EuVat | UsSalesTax | ...
shipping: ShippingPolicy, // Online | Retail | ...
discount: DiscountPolicy // Gold | Silver | ...
) { price(order) = discount.apply(taxRule.apply(base)) + shipping.cost(order) }
// Or fully data-driven, when business tunes the numbers without a deploy:
// rules table: (region, channel, tier) -> {vatRate, shipFlat, discountPct}go deeper
Recognise only that too many classes can be worse than one clear conditional, and that variants sometimes differ only in data.
Name combinatorial explosion and propose composition/Strategy per axis instead of a class per combination.
Add the wrong-axis failure, the expression problem for growing operations, and when a table-driven or sealed-type design fits better.
Frame it as an organizational and evolutionary bet: mine change history for the real axis, decide open vs closed variant sets, plan interface evolution across ownership boundaries, keep the enumeration single-sourced, and expect a deliberate hybrid.
## 1. What the hierarchy is actually optimizing GRASP Polymorphism buys **cheap addition of new types** and pays with **expensive addition of new operations** — the *expression problem*, formally: no conventional language mechanism lets you extend both the set of types and the set of operations without modifying existing code. At small scale you pick a side and move on. At architecture scale the choice determines who has to coordinate with whom to ship a change, which is an organizational commitment, not a coding style. So the principal-level question is never "is this polymorphic?" It is: **which axis actually changes, how often, and who owns the code that must change?** ## 2. Failure mode: combinatorial variation Behavior varies by region *and* channel *and* customer tier. Subtyping gives `EuOnlineGold`, `EuOnlineSilver`, `UsRetailGold`… — a class per cell, and each new region multiplies the grid. Symptoms: deep inheritance, copy-pasted overrides, a fat base class of shared helpers, and the impossibility of answering "what is the tax rule for EU?" without reading twelve files. **Fix: composition over inheritance.** Model each axis as its own small polymorphic role — `TaxRule`, `ShippingPolicy`, `DiscountPolicy` — and assemble an order-processing object from one implementation of each (Strategy per axis). n+m+k classes instead of n×m×k, and each axis evolves independently. This is what "favor composition over inheritance" concretely means at scale. ## 3. Failure mode: the wrong variation axis A hierarchy encodes a bet about where change will happen. Subtype by `ProductType` and every product-specific change is a one-file edit — but if change actually arrives as "new VAT regime in France", it touches every product class. The hierarchy is now *against* you: it not only fails to localize the change, it multiplies it. **Detection:** look at the last 20 changes and count how many touched one class versus many. If the modal change is cross-cutting, the seam is on the wrong axis. **Fix:** re-cut along the observed axis, or externalize the volatile dimension into data/rules. This is expensive, which is why PV insists on *predicted* (evidence-backed) variation rather than imagined variation, and why deferring the abstraction until the second or third real variant is usually correct. ## 4. Failure mode: the enumeration escapes the code Even with a perfect hierarchy, the variant list often also exists as: a DB `CHECK` constraint or lookup table, an OpenAPI/protobuf enum, a UI dropdown, feature-flag keys, analytics dimensions, a data-warehouse mapping, and partner documentation. Now "add one class" is a coordinated release across several systems, and OCP holds in the code while failing in the system. **Fix:** make the variant registry the single source of truth and *generate* the peripheral artifacts (schema values, API enums, UI options) from it; or keep the enumeration in data so that adding a variant is a configuration change with no deployment at all. Deciding this consciously is architecture; discovering it during a release is not. ## 5. Failure mode: operation growth across ownership boundaries Adding `refundWindow()` to a widely implemented interface is trivial in a monorepo with one team. It is a migration when implementations live in other teams' services, in third-party plugins, or in released library versions. Options and their costs: - **Default implementations** on the interface — additive and safe, but defaults silently produce wrong behavior for variants that should have overridden. - **A new, narrower interface** (`RefundPolicyAware`) that callers feature-detect — preserves compatibility, reintroduces a capability check. - **Interface versioning** (`PaymentMethodV2`) — explicit, but doubles surfaces for a migration window. - **Move the operation out** into a Visitor or an external dispatch table so variant classes never change — flips the expression-problem side for that operation specifically. A hybrid is normal and correct: the hierarchy for core operations, a Visitor or table for operations that must not touch third-party code. ## 6. Failure mode: unauditable logic Compliance, pricing, and safety logic must be reviewable *as a whole*: "show me every fee we charge" or "prove no path skips sanctions screening". Spread across 40 subclasses, that review is a manual grep. A conditional table, a rules DSL, or a decision table keeps the whole policy visible and diff-able — which is why insurance, tax and trading systems are so often table-driven rather than hierarchy-driven. **Mitigation if you keep the hierarchy:** a single golden-master/table-driven test that enumerates every variant's outputs side by side, so the review artifact exists even though the code is distributed. ## 7. Failure mode: dynamic-dispatch cost at the wrong layer Megamorphic call sites (many receiver types at one point) defeat inlining and branch prediction, and object-per-variant allocation adds GC pressure in hot paths. This rarely decides an architecture, but in per-packet, per-tick, or per-row code it does: data-oriented designs deliberately batch by type and use a single branch per batch rather than a virtual call per element. Decide with measurements only. ## 8. The alternative toolkit | Alternative | Best when | Cost | |---|---|---| | **Data/table-driven rules** | Variants differ by parameters; business wants to change them without a deploy | Needs validation, versioning, and its own testing story; logic becomes untyped | | **Composition / Strategy per axis** | Multiple independent variation axes | More wiring; assembly must be explicit and validated | | **Sealed types + exhaustive matching** | Closed variant set, operations keep growing | No third-party extension | | **Visitor / external dispatch** | Must add operations without touching variant classes | Verbose; new variants expensive | | **Plugin registry (open set)** | Third parties genuinely extend the system | Runtime rather than compile-time safety; needs contract tests and versioning | | **Keep the conditional** | One branch site, or the enumeration is genuinely local | Reassess at the second duplicate | ## 9. The judgement to articulate At this level the expected answer sounds like: *"I'd ask which axis has actually changed in the last year, whether the variant set is open or closed, who owns the implementations, and whether the enumeration exists outside the code. Open set plus stable operations → hierarchy or registry. Closed set plus growing operations → sealed types and exhaustive matching. Multiple axes → composition. Business-tuned parameters → data. And I'd expect a mature system to use several of these, deliberately, with the reason recorded."*
- How would you decide, with evidence, whether an existing hierarchy is cut on the right axis?Mine the version-control history for the module: for the last N changes, count how many files each touched and group them by intent. If most changes are localized to one subclass, the axis is right. If the modal change fans out across all subclasses (a regulation, a pricing rule, a new required field), the volatile axis is different from the one you subtyped on, and that dimension belongs in composition or data.
- When would you deliberately choose a data-driven rules table over a polymorphic hierarchy for business logic?When non-engineers must change the rules, when the rules must be auditable as a single artifact, when variants differ only in parameters, or when change frequency exceeds your deployment cadence. The costs you accept: validation and versioning of the data, a testing strategy for configurations, loss of compile-time typing, and the risk that the table grows an accidental programming language — at which point a constrained DSL with its own tests is the honest next step.
- A widely implemented interface needs a new operation, but implementations live in other teams' modules. What are your options?Add a default implementation (safe and additive, but silently wrong for variants that should override); introduce a narrow secondary interface that callers feature-detect (compatible, reintroduces a capability check); version the interface with a migration window; or move the operation entirely into a Visitor or external dispatch table so no variant class changes. Choose by whether a silent default is acceptable and how long the migration window can stay open.
Org design: a hierarchy is a reporting structure. Reorganizing by product makes product changes easy and regional changes agonizing. The cost of the wrong structure is not the org chart — it is that every real change now requires cross-team coordination. Class hierarchies pick the same kind of bet, and are just as expensive to undo.
saying these in an interview costs you the question
- Assuming more polymorphism is always better architecture, without asking which axis changes or who owns the implementations.
- Modelling multi-axis variation as one class per combination instead of composing one strategy per axis.
- Ignoring that the variant list is duplicated in schemas, API enums, UI options and analytics, so 'just add a class' is really a multi-system release.
- Adding methods to a widely implemented interface without a migration plan for implementations owned by others.
- Distributing compliance or pricing logic across dozens of subclasses with no single auditable view.
- Introducing hierarchies for speculative variation — Protected Variations targets *predicted*, evidence-backed change.
- Treating dynamic dispatch cost as always negligible, including in per-element hot loops.