What is the data/object anti-symmetry, and how does it guide whether you should model something as polymorphic objects or as plain data plus procedures?
answer
- OO: new types cheap, new operations expensive
- Procedural: new operations cheap, new types expensive
- Expression Problem / rows vs columns of a 2x2 grid
- Visitor flips the trade-off, does not remove it
- Closed type set + exhaustive match = pick data
basics
~20 sObjects make it easy to add new types (write one class; existing code is untouched) but hard to add new operations (every class must change). Plain data plus procedures is the exact opposite. So pick based on which change you expect more often.
solid answer
~60 sThe anti-symmetry: **object-oriented code makes adding new *types* easy and new *functions* hard; procedural code over data structures makes adding new *functions* easy and new *types* hard.** Add a Triangle to a polymorphic Shape hierarchy and no caller changes; add a `perimeter()` operation and every shape class must be edited. With a `Shape` data structure and a switch-based `area(shape)` function, adding `perimeter(shape)` is one new function, but adding Triangle forces every switch to change. Neither style is superior - they trade the same coupling in opposite directions, which is why the same tension appears as the Expression Problem in language design and why the Visitor pattern exists (it buys easy new operations over a fixed type set, at the cost of making new types expensive). Practically: stable set of types, growing set of operations - use data plus functions (compilers, reports, serializers). Growing set of variants, stable operations - use polymorphism (payment methods, plug-in strategies). Mixed and volatile on both axes - keep the data dumb at boundaries and localize the dispatch in one place you can change.
code
pseudocode · 7 lines// Object layout: new type is cheap, new operation touches every class
class Circle { area() {...} perimeter() {...} }
class Square { area() {...} perimeter() {...} }
// Data layout: new operation is cheap, new type touches every function
function area(shape) { match shape { Circle -> ..., Square -> ... } }
function perimeter(shape) { match shape { Circle -> ..., Square -> ... } }go deeper
State the two halves: objects make new types easy, data plus functions make new operations easy - and give one concrete example of each.
Add the mirrored costs, and explain the decision rule of matching the style to the axis of expected change.
Name the Expression Problem, position Visitor and sealed types with exhaustive matching, and describe layering data structures at boundaries with objects in the core.
Reason about it as an organizational coupling decision: which axis different teams own, whether variants come from outside the codebase, versioning and serialization constraints at seams, and the cost of migrating a chosen layout later.
## Statement of the anti-symmetry From *Clean Code*, paraphrased: > Procedural code (code using data structures) makes it easy to add new functions without changing the existing data structures. Object-oriented code makes it easy to add new classes without changing existing functions. The complement is also true: procedural code makes it hard to add new data structures because all the functions must change; OO code makes it hard to add new functions because all the classes must change. The two styles are not good and bad. They are **mirror images**, each cheap on one axis and expensive on the other. ## The 2x2 view Imagine a grid: rows are types (Circle, Square, Triangle), columns are operations (area, perimeter, draw). Any program must fill every cell. - **Object-oriented layout:** group cells by *row*. Each class owns all its operations. Adding a row = one new file, zero edits elsewhere. Adding a column = edit every row. - **Procedural layout:** group cells by *column*. Each function switches over all types. Adding a column = one new function, zero edits elsewhere. Adding a row = edit every function. This is precisely the **Expression Problem** (named by Philip Wadler): no mainstream single mechanism makes both axes cheap without loss of type safety or without extra machinery. ## How to decide Ask which axis actually moves in your system. **Favour objects / polymorphism when:** - New *variants* arrive regularly: payment providers, notification channels, file formats, pricing rules, device drivers. - The set of operations is small and stable. - Variants must be plugged in from outside (plug-ins, modules, third-party extensions) - polymorphism gives you Open/Closed behaviour at the boundary. **Favour data structures + functions when:** - The *type set* is closed and rarely changes: AST node kinds, protocol message kinds, geometry primitives, event kinds. - New *operations* keep arriving: a new report, a new validation pass, a new export format, a new optimization stage. - You need exhaustive-match safety: with sealed/tagged unions the compiler tells you every function that missed a case, which is often *better* than silently inheriting a default method. ## Escape hatches and their price - **Visitor pattern:** re-groups by column inside an OO language. New operation = one new visitor, no type edits. New type = touch every visitor. Same anti-symmetry, flipped. - **Sealed hierarchies / algebraic data types with pattern matching:** closed type set + exhaustiveness checking. Adding a type produces compile errors at every match site - which is a feature when the type set is meant to be closed. - **Double dispatch, multimethods, type classes / traits with coherence rules:** each buys some of both, at the cost of indirection or language-specific machinery. ## Common architectural resolution Most real systems split by layer rather than choosing globally: - **Edges** (HTTP payloads, DB rows, queue messages, config) - dumb data structures. They must be serializable, versionable, diffable; behaviour on them is a liability. - **Core domain** - objects with hidden state and behaviour, so invariants live with the data they protect. - **Cross-cutting operations over a closed type set** (rendering, exporting, auditing) - one dispatch site (visitor / match), not behaviour smeared across every domain class. The failure mode is doing neither consistently: hybrids everywhere, `instanceof`/type-tag switches scattered through the codebase *and* deep hierarchies, so both axes are expensive.
- Where does the Visitor pattern sit in this trade-off?It re-groups behaviour by operation inside an OO language: a new operation is one new visitor class with no edits to the types, but every new type forces a change to every visitor. It flips the anti-symmetry rather than escaping it.
- Why can a sealed hierarchy with exhaustive pattern matching be preferable to polymorphism for a closed type set?Because the compiler flags every match site that has not handled the new case. With polymorphism a forgotten override silently inherits a base implementation, turning a compile-time error into a runtime bug.
A restaurant menu can be organised by dish (each dish lists its steps) or by station (each station lists what it does for every dish). Adding a dish is trivial in the first layout and painful in the second; adding a new prep station is the reverse.
saying these in an interview costs you the question
- Claiming object-orientation is strictly superior and procedural style is legacy
- Saying polymorphism removes all conditionals - it moves the dispatch, and a closed type set may be better served by an explicit exhaustive match
- Ignoring how the system will actually change and choosing by habit or framework default
- Treating Visitor as a way to make both axes cheap
- Putting rich behaviour on wire/DB data structures 'so it is object-oriented'