Once data is split into separate aggregates, what does an operation spanning two of them actually cost?
answer
- The database will not follow the link for you
- Count the round trips per list element
- Two writes are not one write
- Nothing checks that the target exists
- Batch, snapshot, reconcile
basics
~20 sCrossing an aggregate boundary costs an extra access the application must issue and order, no automatic all-or-nothing across the two, and no engine-enforced referential integrity. The usual failures are N+1 read patterns and half-finished multi-document writes.
solid answer
~50 sA cross-aggregate reference is just an identifier, so the database will not resolve it for you at no cost. On the read side you pay a second round trip, and if you resolve references inside a loop you pay one per item — the N+1 pattern that turns a 20-row list into 21 accesses. On the write side there is no free atomicity across the two documents: an update can land on the first and fail on the second, leaving a state no invariant covers. And nothing stops a reference pointing at a document that was never created or has since been deleted, because the engine enforces no foreign key. The mitigations are batching reads by collecting identifiers and fetching them in one multi-key lookup, keeping a small snapshot of the fields you always display alongside the identifier, ordering multi-document writes so a partial failure is recoverable, and reconciling dangling references with a background sweep.
code
javascript · 5 lines// N+1: accesses grow with the size of the list
const posts = await findPostsByTag("databases");
for (const p of posts) {
p.author = await findAuthorById(p.authorId);
}go deeper
Be ready to say that a reference between documents is just a stored identifier, that the application must issue a second read to resolve it, and that nothing guarantees the target exists.
Explain the N+1 pattern concretely and the batching fix, and be able to state that a write spanning two documents can land on one and fail on the other.
Show that you design for these costs up front: batched resolution, idempotent and ordered multi-step writes, a deletion policy, and a reconciliation sweep. Describe how each cost looks in production metrics.
Own the rule of thumb and its consequence: hot paths should cross few boundaries, and a path that crosses several is evidence the boundary is misplaced. Decide where the organisation is allowed to spend coordination and how that is reviewed.
## What a cross-aggregate reference is When the boundary says two things are separate aggregates, the link between them is a stored identifier — `customerId`, `authorId`, a list of `tagIds`. The database treats it as an ordinary value. It does not follow it, does not validate it, and does not include it in any consistency guarantee. Everything that happens when you cross that line is work the application arranges. That is a deliberate trade, not a defect. The boundary bought independent write paths, bounded document size and reduced contention. This section is about pricing the other side. ## Cost one: the extra read, and N+1 Resolving a reference is a second access. One extra round trip on a page that already does one is usually irrelevant. The problem is the shape where resolution happens per element: ```javascript // N+1: one read for the list, then one per element const orders = await getRecentOrders(customerId); for (const o of orders) { o.customer = await getCustomer(o.customerId); } ``` Twenty orders become twenty-one accesses, each with its own network latency, and the cost grows with the page size rather than staying flat. The fix is almost always the same: collect the identifiers first, fetch them in a single multi-key lookup, and stitch in memory. ```javascript // Batched: two accesses regardless of list length const orders = await getRecentOrders(customerId); const ids = [...new Set(orders.map(o => o.customerId))]; const customers = await getCustomersByIds(ids); ``` The deeper fix, when the same fields are needed on every read, is to stop crossing the boundary at read time at all: keep a small copy of the two or three fields the caller always displays next to the identifier, and cross the boundary only when the full document is genuinely needed. That buys read locality at the price of keeping the copy current — a trade to make consciously, and only for fields that change rarely or where staleness is harmless. ## Cost two: no free atomicity across the line A write that must touch both aggregates has no single-operation guarantee. It becomes a sequence, and sequences fail halfway. If step one succeeds and step two fails — process death, network partition, a validation error — the system is left in a state no invariant describes. Three practical responses: - **Order the steps so a partial result is safe.** Write the record that makes the operation *knowable* first, so a recovery pass can detect and finish it. A half-finished state that is detectable and completable is very different from one that is silently wrong. - **Make the steps idempotent** so a retry cannot double-apply. Conditional updates keyed on the current state, or a stored operation identifier, both work. - **Reconcile.** A periodic job that finds incomplete sequences and finishes or reverses them is unglamorous and is what actually keeps such systems honest. Platform-level multi-document transactions, where available, collapse this to one step — at a cost in latency and contention, and with limits of their own. They are a tool for the operations that genuinely need them, not a substitute for a boundary drawn where the rules are. ## Cost three: no referential integrity Nothing prevents storing `customerId: "cus-999"` when no such customer exists, and nothing stops the customer being deleted while orders still point at it. Dangling references appear from bugs, from out-of-order writes, from imports, and from deletions. So the read path must be defensive: a missing target is a normal case to handle, not an exception to crash on. Deletion needs a policy decided up front — refuse while references exist, mark inactive rather than remove, or cascade explicitly — and whichever you pick, a sweep that reports dangling references is worth having, because the engine will never tell you. ## Cost four: querying across the line Filtering by a field that lives on the other side of a boundary is awkward. "Orders placed by customers in Leeds" cannot be answered from the order documents alone; you either resolve the customer set first and filter orders by those identifiers, or you copy the field you filter on into the order. Copying a filterable field across the boundary is one of the most common and most defensible reasons to duplicate a value — and it is a decision to record, because a future reader will see a stray `customerCity` and wonder why. ## Reading the symptoms In production, boundary-crossing costs show up as characteristic shapes: request latency that scales with result-set size (unbatched resolution), a trickle of errors on documents that reference something absent (dangling references), and support tickets about records that exist on one side of an operation but not the other (partial multi-document writes). All three are boundary costs, and all three are cheaper to design for than to debug. ## How to answer this in an interview Name the four costs — extra access, no cross-document atomicity, no referential integrity, awkward cross-boundary filtering — and pair each with its mitigation. Then say the honest part: the right number of boundary crossings on a hot path is small, and if a hot path crosses several, the boundary is probably in the wrong place. Interviewers are listening for someone who prices the trade rather than someone who either fears references or scatters them freely.
- How do you detect an N+1 read pattern before it reaches production?Instrument the data-access layer to count accesses per request and alert when the count scales with result-set size. In tests, assert an upper bound on accesses for a representative list endpoint. The signature is unmistakable once measured: latency that grows linearly with page size while each individual access stays fast.
- When is copying a field across an aggregate boundary the right answer rather than resolving the reference?When the field is needed on nearly every read of the referring document, changes rarely, and a brief stale value is harmless — a display name or a city used for filtering. Record the copy and where it is refreshed. If the field changes often or must be exact, resolve the reference instead and pay the read.
- How do you handle deleting a document other aggregates reference?Pick a policy explicitly: block deletion while references exist, mark the document inactive so readers still resolve it, or cascade with an explicit job. Whichever you choose, make readers tolerate a missing target, because concurrent operations and past bugs will produce dangling references the engine never prevented.
saying these in an interview costs you the question
- Assumes the database enforces referential integrity on stored identifiers
- Resolves references one at a time inside a loop
- Treats a two-document write as if it were atomic
- Crashes on a reference whose target no longer exists
- Reaches for cross-document transactions before fixing the boundary