skip to content

In document data modelling, what makes a document an aggregate rather than just a stored record?

level: juniorimportance: should knowfreq 50%

answer

  1. Think about the unit, not the field
  2. What gets loaded and saved together
  3. One boundary, three properties
  4. Consistency scope, not just relatedness
  5. The document is the aggregate

basics

~20 s

An aggregate is the clump of data you load, write and keep consistent as one unit. A document is an aggregate when its contents are always used together and nothing outside it must change atomically with it.

solid answer

~40 s

A record is just a row of fields; an **aggregate** is a boundary. It says: this data is fetched together, written together, and is consistent together. In a document store the document *is* that boundary — one read returns the whole thing, one write replaces or updates it atomically. So an order document that carries its line items, its shipping address snapshot and its totals is an aggregate: the application never wants a line item without its order, and never needs to change a line item and something outside the order in the same indivisible step. The practical test is not "do these things belong together conceptually" but "are they read together and do they have to change together". If the answer is no, they are probably two aggregates linked by an identifier.

go deeper

for a junior

Be ready to define an aggregate as the unit that is read, written and kept consistent together, and to say that in a document store the document is that unit. One worked example, such as an order carrying its lines, is enough.

for a middle

Explain the three properties behind the boundary and why relatedness alone is not the criterion. Be able to point at a field in a sample document and justify why it is nested rather than linked.

for a senior

Show that you treat the boundary as a physical decision with measurable costs — extra round trips when it is too small, write contention and wasted I/O when it is too large — and that you derive it from observed access, not from a domain diagram.

for a principal

Own the consequence at system scale: aggregate boundaries decide which operations can ever be atomic, which teams contend on the same documents, and how far a schema change ripples. Be ready to argue why a boundary is worth defending even when a new feature asks to cross it.

## The word "aggregate" An **aggregate** is a group of data treated as a single unit for retrieval, modification and consistency. The term predates document databases, but document stores are the family that took it literally: the storage engine's unit — the document — *is* the aggregate. That is why this style of database is often called *aggregate-oriented*, in contrast to the relational model, where a single conceptual thing is normally spread across several tables and reassembled by joins at query time. A stored record, by itself, carries no such claim. A row in a table is just a tuple of values; whether it forms a meaningful unit with other rows is expressed elsewhere, through keys and joins. A document, by contrast, physically contains its parts. When you store an order with its line items nested inside it, you have asserted three things at once: these are read together, they are written together, and they are consistent together. ## The three properties that make it an aggregate **Read together.** One lookup by identifier returns everything the operation needs. There is no second round trip to assemble the shape the screen or the API response wants. This is the property that makes document reads fast and predictable: the cost of loading an order is one key lookup, regardless of how many lines it has. **Written together.** A document store gives you atomicity over a single document essentially for free — the update either lands entirely or not at all. Anything inside the boundary can be changed in one indivisible step. Anything outside it cannot, without extra machinery. **Consistent together.** Rules that must never be observed broken — a total that must equal the sum of its lines, a status that must not be *shipped* while a line is unallocated — can only be enforced cheaply *inside* one aggregate, because that is the widest scope a single write covers. ## An example ```json { "_id": "ord-4821", "customerId": "cus-77", "status": "placed", "lines": [ { "sku": "KB-01", "qty": 2, "unitPrice": 4900 }, { "sku": "MS-14", "qty": 1, "unitPrice": 2900 } ], "total": 12700 } ``` The lines live inside because no use case wants a line without its order, and because `total` must agree with them at all times. The customer does **not** live inside: customers are read on their own, updated on their own, and one customer has an open-ended number of orders. `customerId` is a link across an aggregate boundary — a plain value the application resolves with a second lookup when it actually needs the customer's name. ## What this changes about how you design Because the aggregate is the storage unit, the boundary is not a documentation artefact — it is a physical decision with runtime consequences. Draw it too small and every operation becomes several reads and a multi-step write. Draw it too large and you load megabytes to change one field, and unrelated writers collide on the same document. The question that decides the boundary is behavioural, not taxonomic: *what does the application load in one go, and what must change in one go?* Two things that sound related on a whiteboard — a product and its reviews, a user and their activity log — often belong to different aggregates, because they are read on different screens, written by different flows, and grow at wildly different rates. Conversely, two things that look like separate entities — an order and its shipping address as captured at purchase time — often belong in one document, because the order's copy is a historical fact that must never change when the customer edits their address book. ## Where the boundary stops Outside the aggregate you have identifiers, not nesting. Crossing that line at runtime means a second read the application issues itself, and a multi-step write with no automatic all-or-nothing guarantee. Those costs are the price of the boundary and should be visible when you draw it, not discovered later. ## How to answer this in an interview Say what the aggregate is (a unit of retrieval, modification and consistency), say that the document is that unit in a document store, and then give the behavioural test — read together, written together, consistent together — with one concrete example of something you would nest and something you would link. Avoid answering purely in terms of "related data": relatedness is not the criterion, joint access and joint consistency are.

  • Are aggregate boundaries and entity boundaries the same thing?
    No. An entity is a conceptual thing with an identity; an aggregate is a storage and consistency unit that may contain several entities, or may be one entity split off from its neighbours. An order line is an entity in a discussion but usually not an aggregate: it never leaves its order. The boundary is decided by joint access and joint consistency, not by the noun list.
  • Can a single document hold two aggregates?
    It can physically, but that is a design smell. If two parts of a document are read on separate screens, written by separate flows, and have no rule spanning them, keeping them in one document means every writer contends on the same record and every reader pays for data it discards. Split them and link by identifier.

An aggregate is like a paper folder in a filing cabinet: you pull out the whole folder, mark it up, and put it back as one thing. A cross-reference note pointing at another folder is cheap to write, but acting on it means a second trip to the cabinet.

saying these in an interview costs you the question

  • Says an aggregate is just any group of related fields
  • Nests data because it is conceptually related, not jointly used
  • Thinks the boundary is documentation, with no runtime cost
  • Treats every entity in the domain as its own document by default
  • Cannot say what consistency guarantee the boundary buys

context