In Robert Martin's Clean Code terminology, what is the difference between an "object" and a "data structure"?
answer
- Object = hide data, expose behaviour
- Data structure = expose data, no behaviour
- Getters+setters everywhere = data structure in disguise
- Hybrid = worst of both
- Objects add types easily, data adds operations easily
basics
~20 sAn object hides its data and exposes behaviour - you tell it to do something. A data structure exposes its data and has almost no behaviour - other code reads its fields and does the work.
solid answer
~50 sThe line is drawn by what is public. An **object** keeps its representation private and publishes *behaviour*: callers say shape.area() without knowing whether a circle stores a radius or a bounding box. A **data structure** publishes its *data* (public fields, or getters/setters that are just field access with ceremony) and has little meaningful behaviour: a Point{x,y}, a JSON payload, a DTO, a database row. Neither is wrong; they solve opposite problems. Objects let you add new *types* without touching existing callers, because behaviour travels with the type. Data structures let you add new *operations* in one place without touching the types. The mistake is the hybrid: a class with public getters/setters for everything *and* significant business rules. It gets the drawbacks of both - callers reach into its data, yet it hides enough that you cannot treat it as plain data. Pick one side per type, deliberately.
code
pseudocode · 9 lines// Data structure: exposes data, no behaviour
struct Rectangle { width, height }
function area(r) { return r.width * r.height } // logic lives outside
// Object: hides data, exposes behaviour
class Rectangle {
private width, height
function area() { return width * height } // logic lives inside
}go deeper
State the one-line contrast (hide data + expose behaviour vs expose data + no behaviour) and give one example of each, e.g. an Account object vs a Point data structure.
Add that trivial getters/setters do not create encapsulation, and name the hybrid as the thing to avoid.
Frame the choice by expected axis of change - new types favour objects, new operations favour data structures - and describe where in a system each belongs (edges vs core domain).
Connect it to the Expression Problem and to architectural boundaries: dumb data at the seams for versioning and serialization, behaviour-rich objects inside; and set team conventions so hybrids do not accumulate.
## The distinction The vocabulary comes from the *Objects and Data Structures* chapter of Robert C. Martin's *Clean Code*, but the idea is older and language-independent. **Object** = hides its representation, exposes behaviour. - Fields are private/internal; nobody outside can see how state is stored. - Public methods are domain verbs: `withdraw(amount)`, `area()`, `render()`, `isEligible()`. - Callers *tell* it what to do rather than pulling out its parts and deciding for it. **Data structure** = exposes its representation, has (almost) no behaviour. - Fields are public, or accessed via trivial getters/setters that add nothing. - Behaviour lives in *other* code: functions, services, procedures that read the fields. - Examples: geometric `Point{x, y}`, a request/response DTO, a row mapped from a database, a parsed JSON document, a struct in a C-style language. ## Getters and setters do not make it an object The single most common misconception: "my fields are private and I wrote getX/setX, therefore it is encapsulated." It is not. If every field has a public getter and setter, the representation is fully visible and fully mutable - you have written a data structure with extra typing. Real encapsulation means an outside caller *cannot* learn how the data is stored. `account.getBalance()` may still be legitimate if balance is part of the abstraction the class promises; `account.getInternalLedgerEntries()` almost certainly is not. ## Why the distinction matters Because it determines *where new code goes* when requirements change: - With objects, adding a new **type** (a new shape, a new payment method) means writing one new class. Existing callers do not change - they already call the abstract method. - With data structures plus procedures, adding a new **operation** (a new report, a new export format) means writing one new function. Existing types do not change. Those are exactly opposite strengths, and the trade-off is symmetric (this is the *data/object anti-symmetry*, and it is the same force behind the Expression Problem in language design and the Visitor pattern). ## Choosing - Domain concepts with invariants to protect (money, an order, a policy) - make them objects: hide the fields, publish verbs. - Transport / edge shapes (API payloads, config, query results, coordinates, events) - make them data structures: plain, public, dumb, easy to serialize. - Boundary rule of thumb: data structures are fine *at* the boundary of the system; convert them into real objects once inside. ## The hybrid anti-pattern A class with public accessors for all its state *and* important business methods is a hybrid. It invites callers to bypass the behaviour (`if (order.getStatus() == PAID && order.getItems().isEmpty())`) while also being too opinionated to serialize or diff cleanly. Hybrids are the usual habitat of train-wreck chains and Law-of-Demeter violations, because callers get used to reaching in.
- If a class has private fields but a public getter and setter for every one of them, is it an object or a data structure?Effectively a data structure. The representation is fully exposed and mutable, so callers can (and will) build logic on its shape. The accessors add ceremony, not encapsulation.
- Is a DTO carrying data across a service boundary a bad design because it has no behaviour?No. A DTO is a deliberate data structure: dumb, serializable, versionable. The mistake would be letting business rules grow on it, turning it into a hybrid.
An object is a vending machine: you press a button and get a drink; the mechanism is sealed. A data structure is a supermarket shelf: everything is visible and you assemble the meal yourself.
saying these in an interview costs you the question
- Claiming private fields plus getters/setters automatically means encapsulation
- Saying data structures are always bad / everything must be an object
- Confusing the term 'data structure' here with algorithmic data structures like hash maps or trees
- Treating DTOs, config records, and API payloads as design smells
- Believing objects are always the safer default regardless of how the code will change