When should you embed related data in a document and when should you store a reference?
answer
- Start from the access pattern, not the diagram
- Ask whether the child is read alone
- Bounded and read together points one way
- Shared or unbounded points the other way
basics
~10 sEmbed when the related data is bounded in size, read with its parent, and changes with it. Reference when it is unbounded, shared by many parents, edited independently, or queried on its own.
solid answer
~40 sThe choice comes from the access pattern, not from the entity diagram. **Embed** when the child data is small and bounded, is almost always read together with the parent, and belongs to that parent alone — addresses on a user, line items on an order. One read returns everything, and one write updates it all atomically. **Reference** when the child can grow without limit (comments, events, telemetry), when the same entity is shared by many parents and edits must be seen everywhere, or when the child is queried and paginated on its own. Referencing costs an extra lookup per read and gives up cross-document atomicity, but it keeps the parent small and the child independently addressable. The honest answer in an interview is a checklist: cardinality, is-it-read-together, does-it-grow, is-it-shared, how-often-does-it-change.
go deeper
Be ready to state the two options and give one concrete example of each — addresses embedded on a user, comments referenced from a post — and to say that the deciding factor is how the data is read.
Explain the mechanics behind the tradeoff: one document is fetched and written as a unit, so embedding buys locality and atomic updates while referencing buys unbounded growth and a single shared copy.
Show judgment on a real domain: name the bound on the many side, say what happens when that bound is wrong, and describe the hybrid you would use when the hot read needs only a couple of fields from the referenced record.
Own the standard. Decide how the organisation resolves conflicting access patterns over the same entity, when a second read-optimised copy is worth its divergence risk, and how new services are prevented from each inventing their own shape.
## What the decision actually is A document database stores whole, self-describing records containing nested objects and arrays, and retrieves a record by identifier in a single operation. When one entity relates to another — an order and its line items, a user and their addresses, a post and its comments — there are two physical choices. **Embedding** nests the related data inside the parent record, so the relationship is expressed by containment. **Referencing** stores the related data as separate records and keeps only their identifiers in the parent, so the relationship is a value you resolve later — with a second query, an application-side join, or a server-side join stage where the product offers one. Neither is a default. You pick from how the data is read, how it is written, and how it grows. ## What embedding buys **Read locality.** The parent and its children arrive in one operation, already assembled, stored contiguously and cached as one unit. There is no second round trip and no join, which is exactly why document stores exist. **Atomic writes.** A write to a single document is all-or-nothing. Fields spread across the parent and its embedded children can be updated together, with no reader ever observing a half-applied change and no coordination protocol. **Simpler code.** There is nothing to resolve, nothing to be missing, and no ordering problem between two writes. ## What embedding costs The whole document is the unit of storage and, in most engines, of write. Adding one element to an embedded array can mean rewriting the record. Every read of the parent carries the children unless you project them away. If the same entity is embedded in many parents, an edit to it has to be applied in every copy. And an embedded child cannot easily have a lifecycle of its own — it exists only for as long as, and only inside, its parent. ## What referencing buys and costs Referencing gives the child an independent lifecycle, unlimited growth, and one authoritative copy that all parents see. It lets the child be queried, sorted and paginated on its own terms, and keeps the parent document small enough to stay in cache. The costs are symmetrical: an extra round trip or join on every read that needs both, no atomicity across the two writes, more application code, and no engine-enforced integrity — document stores generally do not stop you deleting a record that other documents still point at, so the reader must tolerate a missing target. ## Cardinality as the first cut A useful first pass is the size of the "many" side: - **One-to-few** — a handful, with a natural ceiling: addresses, phone numbers, product variants. Embed. - **One-to-many, bounded in practice** — tens to a few hundred, read with the parent: order line items, a document's revisions-summary. Usually embed, and watch the tail. - **One-to-unbounded** — comments on a viral post, sensor readings, audit events. Always reference. There is no ceiling, and a design with no ceiling eventually meets the store's maximum document size. - **Many-to-many** — if either side is unbounded, reference; often with link records of their own. ## The questions to ask out loud 1. Is the child ever read without the parent? 2. Is the parent often read when the child is not needed? 3. Does the number of children have a real upper bound? 4. Is the child shared by more than one parent, and must every parent see edits immediately? 5. How often does the child change relative to the parent? 6. Do parent and child have to change together to preserve a rule? Answers of "no, no, yes, no, rarely, yes" describe an embed. The opposite pattern describes a reference. ## The middle ground exists The choice is not binary. You can keep a reference *and* a small copy of the one or two fields the hot read needs, so the common case is one read and the full record is still resolvable. You can keep a bounded slice — the most recent few children — inside the parent while the complete set lives in its own collection. These hybrids are usually the right answer for a read-heavy page that also needs an unbounded history. ## A worked example An order: line items embedded, because they are bounded, always read with the order, and their price is a point-in-time fact rather than a copy of the current price. The customer referenced, because the customer is shared across many orders, edits their own profile, and has a lifecycle of their own — with the customer's name copied onto the order if the order-history screen needs it without a second read. Shipping events referenced, because a slow international parcel can accumulate an unpredictable number of them.
- What happens when a referenced document is deleted while parents still point at it?You get a dangling reference. Document stores generally do not enforce that the target exists, so nothing stops the delete and nothing repairs the pointers. The application has to decide: tolerate a missing target on read and render a placeholder, cascade the delete itself, or mark the target inactive instead of removing it. Whichever you pick, it has to be a deliberate rule, because the engine will not remind you.
- Does embedding remove the need to index the nested fields?No. A query that filters on an embedded field still needs an index on that path, or it scans every parent document. Embedding changes where the data lives and how it is fetched once found; it does not change how the engine locates matching documents. Indexes over array fields also carry one entry per element, so a large embedded array makes the index larger too.
- How do you decide when a child is read on its own about as often as it is read with its parent?Measure both paths rather than guessing. If the standalone reads need sorting, filtering or pagination across children of many parents, that argues strongly for a separate collection, because those operations are awkward and expensive over arrays scattered inside parents. If the standalone read is a rare admin screen, embed and let that path do the extra work.
saying these in an interview costs you the question
- Always embed, because joins are slow in document databases
- Always reference, mirroring a normalized relational schema
- Assumes documents can grow to any size
- Expects the database to keep referenced documents consistent automatically
- Picks the model from the entity diagram instead of the queries