How would you choose aggregate boundaries for a new document-database system when the access patterns are still uncertain?
answer
- Some parts of a domain change slower than others
- Anchor on what the business does, not the screen
- Modest aggregates are easier to widen than to split
- One mistake is asymmetric: unbounded nesting
- Instrument early so patterns become evidence
basics
~20 sAnchor the boundaries on what is stable — the invariants and the units the business writes as a whole — keep aggregates modestly sized, avoid unbounded nesting, and treat everything derived as replaceable so shapes can change as real access patterns appear.
solid answer
~50 sQuery-first modelling assumes you know the queries; early on you often do not. So anchor on the parts that change slowest. Invariants are stable: what must always be true tends to outlive any screen. The **write unit** — what the business creates or updates as one act — is also stable, and it is usually a good aggregate. Beyond those anchors, stay conservative: modest documents, no unbounded nesting, no copied field you cannot justify, and identifiers on every cross-boundary link so a shape can be assembled later. Then make the model observable: instrument access counts and shapes from day one so the real patterns become data rather than argument. Treat derived shapes as replaceable and the primary aggregates as the thing you defend. Reserve the expensive reshaping decisions until the workload tells you which way to lean.
go deeper
Be ready to say that a first model should be conservative — modest documents, identifiers between them, and no unbounded nesting — because shapes are easier to widen later than to split.
Explain which parts of a domain are stable enough to anchor a boundary — invariants, the unit the business writes as one act, lifecycle, cardinality — and which parts should be deferred until a workload exists.
Show that you would instrument the model from the first release: accesses per request, document size distribution, contention hot spots. Describe how those measurements turn a reshaping argument into an evidence-based decision.
Own the strategy: which decisions are load-bearing and taken now, which are deliberately deferred, what evidence would reopen them, and how the codebase stays reshapeable so moving a boundary later is a project rather than a rewrite.
## The tension The standard advice is to model from access patterns. Early in a product's life that advice has a gap: the patterns are guesses. Guessing produces one of two failures — an over-optimised shape built for a workload that never arrives, or a shape so generic that it is slow at everything. The way through is to notice that not everything about a domain is equally uncertain. Some things change every sprint; others outlive several rewrites of the interface. Anchor on the stable parts and stay flexible about the rest. ## What is actually stable **Invariants.** "A booking never exceeds capacity." "An order's total matches its lines." These come from how the business works, not from how it is displayed, and they rarely change. They fix parts of the boundary and they fix them durably, so they are the best possible anchor. **The unit of work.** What does the business create, submit or approve as one act? An order is placed as a whole. An invoice is issued as a whole. A message is sent as a whole. These write units are stable because they reflect real-world procedure, and they are usually good aggregates independently of how anyone queries them later. **Lifecycle and ownership.** Data that is created, changed and deleted together, by the same actor, at the same time, tends to belong together. Data with divergent lifecycles — an account created once and a login event stream appended forever — tends not to, no matter how related it looks. **Cardinality shape.** Whether a relationship is one-to-few or one-to-unbounded is a property of the domain, not the workload. A collection that can grow without limit must not be nested, and knowing that early prevents the most expensive class of mistake. ## What to defer Everything derived. Precomputed figures, read-oriented shapes, copies of fields for filtering — these are consequences of a workload, so they should be built when the workload shows up, and built so they can be discarded when it changes. Keep them clearly separated from the aggregates that hold truth, so that deleting one is a local operation rather than a data migration. ## Principles for a conservative first model 1. **Keep aggregates modest.** A small aggregate can be widened by copying data in; a bloated one is painful to split because references to its inner parts have already spread through the code. 2. **Never nest an unbounded collection.** This is the one asymmetric mistake: it degrades reads and writes together, and it gets worse with success. If the count has no natural ceiling, it is a separate aggregate. 3. **Use identifiers everywhere you cross a line**, even where you suspect you will later copy fields across. Adding a copy later is easy; recovering a lost identifier is not. 4. **Prefer explicit fields to clever generic containers.** Shapes designed to hold anything are hard to query and harder to reason about, and they defer the modelling problem instead of solving it. 5. **Do not distribute prematurely.** Boundaries chosen to satisfy a scale requirement that has not arrived make every early feature slower to build, and the eventual real requirement is rarely the one that was guessed. ## Make the model observable The fastest way to stop guessing is to measure. From the first release, capture accesses per request, which entry points are actually used, the size distribution of documents, and how result sets grow. Within weeks the real access patterns are data rather than opinion, and the reshaping conversation becomes concrete: *this endpoint issues eleven accesses per request and runs forty times a second* is an argument; *users will probably want this on one page* is not. Size distribution deserves special attention, because the tail is where the damage is. A collection whose ninety-ninth-percentile document is a hundred times the median is telling you that some nesting decision has no ceiling. ## Plan for reshaping Accept that boundaries will move and make moving them survivable. Route data access through a layer thin enough that the stored shape is not spread across the whole codebase. Keep the primary aggregates the source of truth so a derived shape can be rebuilt rather than migrated. And when a boundary does need to move, treat it as a real change with a transition period rather than a switch — the mechanics of running a shape change against live data belong to their own discipline, but the design principle is simply that the model must be reshapeable at all. ## What to say no to A product owner who cannot describe a workload will still describe requirements. "It has to be fast" and "it should scale" are not access patterns and cannot justify a shape. The useful counter-question is behavioural: *what will people do most often, and what happens if this number is a minute old?* Two honest answers to those often do more for the model than a week of diagramming. ## How to answer this in an interview Say explicitly that you would not pretend to know the queries. Name the stable anchors — invariants, write units, lifecycle, cardinality — say what you defer, give the conservative principles with the unbounded-nesting rule called out as the asymmetric one, and finish with instrumentation as the mechanism that converts guesses into evidence. This is a judgment question with no single right answer; what is being assessed is whether you can commit to a defensible starting point while keeping the expensive decisions open.
- What is the single most expensive boundary mistake to make early?Nesting a collection with no natural ceiling. It degrades reads and writes simultaneously, worsens as the product succeeds, and is painful to undo because code has already been written against the nested shape. Almost every other early mistake — a document slightly too small, a missing copied field — is cheap to correct later.
- How do you push back when nobody can describe the access patterns?Replace the abstract question with behavioural ones: what will users do most often, what must be on screen within a second, and what happens if this figure is a minute stale. Those are answerable without a design document, and the answers constrain the model far more usefully than 'it must be fast and scale'.
- How do you know when it is time to revisit the boundaries?When the instrumentation says so: an endpoint whose access count scales with result size, a document-size tail far above the median, or a write path contending on one hot document. Those are measurable signals. Re-model on evidence, not on the feeling that the shape has aged.
Framing a building before the tenants are chosen: you fix the load-bearing structure, which is expensive to change, and leave the partitions light, because those will move once people move in.
saying these in an interview costs you the question
- Claims access patterns can always be known fully up front
- Nests a collection with no ceiling to save an early read
- Builds derived shapes before any workload justifies them
- Optimises the first model for a scale that has not arrived
- Ships with no instrumentation of access counts or document sizes