skip to content

Boundaries, Objects & Data Structures

Objects hide data and expose behavior, data structures do the opposite, and mixing the two produces hybrids that are bad at both. This node also covers keeping third-party libraries behind your own wrapper so their API changes stop at one file.

part ofSoftware design & architectureoverview, primer and where to startread it →
on this pageshow

questions

6

In Robert Martin's Clean Code terminology, what is the difference between an "object" and a "data structure"?

level: juniorimportance: must knowfreq 62%

answer

  1. Object = hide data, expose behaviour
  2. Data structure = expose data, no behaviour
  3. Getters+setters everywhere = data structure in disguise
  4. Hybrid = worst of both
  5. Objects add types easily, data adds operations easily

basics

~20 s

An object hides its data and exposes behaviour - you tell it to do something. A data structure exposes its data and has almost no behaviour - other code reads its fields and does the work.

solid answer

~50 s

The line is drawn by what is public. An **object** keeps its representation private and publishes *behaviour*: callers say shape.area() without knowing whether a circle stores a radius or a bounding box. A **data structure** publishes its *data* (public fields, or getters/setters that are just field access with ceremony) and has little meaningful behaviour: a Point{x,y}, a JSON payload, a DTO, a database row. Neither is wrong; they solve opposite problems. Objects let you add new *types* without touching existing callers, because behaviour travels with the type. Data structures let you add new *operations* in one place without touching the types. The mistake is the hybrid: a class with public getters/setters for everything *and* significant business rules. It gets the drawbacks of both - callers reach into its data, yet it hides enough that you cannot treat it as plain data. Pick one side per type, deliberately.

code

pseudocode · 9 lines
pseudocode
// Data structure: exposes data, no behaviour
struct Rectangle { width, height }
function area(r) { return r.width * r.height }   // logic lives outside

// Object: hides data, exposes behaviour
class Rectangle {
  private width, height
  function area() { return width * height }      // logic lives inside
}

go deeper

for a junior

State the one-line contrast (hide data + expose behaviour vs expose data + no behaviour) and give one example of each, e.g. an Account object vs a Point data structure.

for a middle

Add that trivial getters/setters do not create encapsulation, and name the hybrid as the thing to avoid.

for a senior

Frame the choice by expected axis of change - new types favour objects, new operations favour data structures - and describe where in a system each belongs (edges vs core domain).

for a principal

Connect it to the Expression Problem and to architectural boundaries: dumb data at the seams for versioning and serialization, behaviour-rich objects inside; and set team conventions so hybrids do not accumulate.

## The distinction The vocabulary comes from the *Objects and Data Structures* chapter of Robert C. Martin's *Clean Code*, but the idea is older and language-independent. **Object** = hides its representation, exposes behaviour. - Fields are private/internal; nobody outside can see how state is stored. - Public methods are domain verbs: `withdraw(amount)`, `area()`, `render()`, `isEligible()`. - Callers *tell* it what to do rather than pulling out its parts and deciding for it. **Data structure** = exposes its representation, has (almost) no behaviour. - Fields are public, or accessed via trivial getters/setters that add nothing. - Behaviour lives in *other* code: functions, services, procedures that read the fields. - Examples: geometric `Point{x, y}`, a request/response DTO, a row mapped from a database, a parsed JSON document, a struct in a C-style language. ## Getters and setters do not make it an object The single most common misconception: "my fields are private and I wrote getX/setX, therefore it is encapsulated." It is not. If every field has a public getter and setter, the representation is fully visible and fully mutable - you have written a data structure with extra typing. Real encapsulation means an outside caller *cannot* learn how the data is stored. `account.getBalance()` may still be legitimate if balance is part of the abstraction the class promises; `account.getInternalLedgerEntries()` almost certainly is not. ## Why the distinction matters Because it determines *where new code goes* when requirements change: - With objects, adding a new **type** (a new shape, a new payment method) means writing one new class. Existing callers do not change - they already call the abstract method. - With data structures plus procedures, adding a new **operation** (a new report, a new export format) means writing one new function. Existing types do not change. Those are exactly opposite strengths, and the trade-off is symmetric (this is the *data/object anti-symmetry*, and it is the same force behind the Expression Problem in language design and the Visitor pattern). ## Choosing - Domain concepts with invariants to protect (money, an order, a policy) - make them objects: hide the fields, publish verbs. - Transport / edge shapes (API payloads, config, query results, coordinates, events) - make them data structures: plain, public, dumb, easy to serialize. - Boundary rule of thumb: data structures are fine *at* the boundary of the system; convert them into real objects once inside. ## The hybrid anti-pattern A class with public accessors for all its state *and* important business methods is a hybrid. It invites callers to bypass the behaviour (`if (order.getStatus() == PAID && order.getItems().isEmpty())`) while also being too opinionated to serialize or diff cleanly. Hybrids are the usual habitat of train-wreck chains and Law-of-Demeter violations, because callers get used to reaching in.

  • If a class has private fields but a public getter and setter for every one of them, is it an object or a data structure?
    Effectively a data structure. The representation is fully exposed and mutable, so callers can (and will) build logic on its shape. The accessors add ceremony, not encapsulation.
  • Is a DTO carrying data across a service boundary a bad design because it has no behaviour?
    No. A DTO is a deliberate data structure: dumb, serializable, versionable. The mistake would be letting business rules grow on it, turning it into a hybrid.

An object is a vending machine: you press a button and get a drink; the mechanism is sealed. A data structure is a supermarket shelf: everything is visible and you assemble the meal yourself.

saying these in an interview costs you the question

  • Claiming private fields plus getters/setters automatically means encapsulation
  • Saying data structures are always bad / everything must be an object
  • Confusing the term 'data structure' here with algorithmic data structures like hash maps or trees
  • Treating DTOs, config records, and API payloads as design smells
  • Believing objects are always the safer default regardless of how the code will change

context

open as a page

What does the Law of Demeter state, and why is a chain like order.getCustomer().getAddress().getCity().toUpperCase() considered a problem?

level: middleimportance: must knowfreq 70%

basics

~20 s

The Law of Demeter says a method should only talk to its immediate neighbours: itself, its own fields, its parameters, and objects it created. Long chains like that one couple your code to the internal structure of three other classes, so any of them changing breaks you.

open as a page

How should you manage the boundary between your code and a third-party library or external API, and what are "learning tests" in that context?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Do not let a third-party type spread through your codebase. Wrap it behind an interface you own, expressed in your domain's terms, and keep the vendor types inside that wrapper. Learning tests are small tests you write against the library to verify how it actually behaves - and they warn you when a version upgrade changes that behaviour.

open as a page

What is a "hybrid" class in the objects-versus-data-structures sense, and how do Tell-Don't-Ask and feature envy relate to it?

level: middleimportance: should knowfreq 40%

basics

~20 s

A hybrid is a class that both exposes all its data through getters/setters and carries important business behaviour. It gets the downsides of both styles. Feature envy is code that uses another object's data more than its own; Tell-Don't-Ask fixes it by moving the behaviour to the data.

open as a page

What is the data/object anti-symmetry, and how does it guide whether you should model something as polymorphic objects or as plain data plus procedures?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Objects make it easy to add new types (write one class; existing code is untouched) but hard to add new operations (every class must change). Plain data plus procedures is the exact opposite. So pick based on which change you expect more often.

open as a page

When does hiding an external dependency behind your own abstraction stop paying for itself, and how do you decide where to place - or not place - a boundary?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

A boundary pays off when the thing behind it is likely to change, is hard to test, or would otherwise spread everywhere. It stops paying when the abstraction just mirrors the dependency, when the dependency is a stable standard, or when its behaviour leaks through anyway.

open as a page