In a document database, when should documents of different shapes share one collection?
answer
- ask how the documents are read
- one feed, one sort, one index
- a field that names the variant explicitly
- shared fields must match in name and type
- split when growth or retention diverges
basics
~20 sKeep different shapes together when they are read by the same queries, share a common core of fields, and need the same indexes. Split them when access patterns, growth rates or retention differ. Always carry an explicit type field to tell variants apart.
solid answer
~50 sThe deciding question is not "are these the same thing?" but **"are they read together?"** If a single feed, list or timeline has to return card payments, bank transfers and vouchers ordered by date, keeping them in one collection means one query and one index; splitting them means several queries merged in application code. That is the case for polymorphism. Three rules make it work: carry an explicit **discriminator** field naming the variant; keep the fields that all variants share identical in name *and* type, because those are what the shared queries and indexes touch; and push variant-specific fields into a nested sub-object so the top level stays predictable. Split into separate collections when the variants are queried separately anyway, when one variant is far larger or grows faster, when retention or access-control rules differ, or when the union shape has drifted so far that only the discriminator is genuinely shared.
code
json · 7 lines{ "_id": 1, "type": "card", "amount": 1200, "currency": "EUR",
"createdAt": "2026-03-01T10:00:00Z",
"details": { "last4": "4242", "brand": "visa" } }
{ "_id": 2, "type": "voucher", "amount": 500, "currency": "EUR",
"createdAt": "2026-03-01T10:04:00Z",
"details": { "code": "XMAS26" } }go deeper
Know that one collection can legitimately hold documents with different fields, and that a type field is how code tells them apart.
Explain the access-pattern test — are they returned by the same query and sort — and the three rules: explicit discriminator, stable common core, variant data nested in a sub-object.
Demonstrate the split criteria you would actually apply in production: divergent growth rates, different retention or access rules, and shared indexes dominated by one variant. Mention the summary-plus-payload hybrid.
Own the long-run consequence: a polymorphic collection is a shared contract between every team that writes a variant, so adding a variant must be additive to the core and the core must be governed.
## The real decision driver: access pattern, not taxonomy A classic modelling instinct is to group by what things *are*. In a document model the better instinct is to group by how they are *read*. Documents that always appear together in the same result — a payment feed containing three kinds of payment, an activity stream containing several kinds of event, a catalogue containing several kinds of product — belong in one collection, because that lets one query with one sort and one index serve the whole feature. Conversely, two document kinds that are never queried together gain nothing from sharing storage. Putting them together just means every query must filter by kind, every index carries entries for documents it will never serve, and any shape check must accept the union of both. ## The three rules that make polymorphism work **1. An explicit discriminator.** Every document carries a field — `type`, `kind`, `eventType` — whose value names the variant. Never infer the variant from the presence of some other field ("if it has `iban` it must be a transfer"), because that inference breaks the moment a fourth variant also has that field. The discriminator should be a short, closed set of values, and it is one of the best candidates for a database-side shape check. **2. A stable common core.** Whatever the shared queries and sorts touch must be present on every variant, under the same name, with the same type: an identifier, a timestamp, an owner, a status, an amount. This is the part of the document that behaves like a table. If one variant calls it `createdAt` and another `created_at`, the feed query silently loses half its results. **3. Variant data in a sub-object.** Put the fields that only one variant has inside a nested object — `details`, `payload`, `attributes` — rather than spreading a dozen optional top-level fields across the collection. The top level stays small and predictable, readers can see at a glance what is universal, and adding a variant does not widen the shared surface. ## What it costs Polymorphic collections are not free: - **Every query must filter by the discriminator** unless it genuinely wants all variants. Forgetting that filter is a common bug and usually returns too much rather than erroring. - **Indexes span all variants.** An index on a field that only one variant carries still lives in the same collection; whether that costs anything depends on the engine's options, but it is at least a shared resource that one variant's growth affects. - **Shape enforcement becomes a union.** A single check must accept every variant, so it can only require the common core plus a valid discriminator. Enforcing the variant-specific parts either becomes conditional, which not every validator can express, or stays in application code. - **Readers must branch.** Any code touching the variant-specific parts needs a switch on the discriminator, and forgetting a case is a silent bug rather than a compile error in most stacks. ## When to split instead Split into separate collections when any of these hold: - The variants are queried separately in practice — the shared feed you imagined never materialised. - One variant is vastly larger or grows much faster, so it dominates every shared index and any collection-level operation. - Retention, archival or access-control rules differ. If one kind must be purged after a short window and another kept for years, storing them together makes both policies harder. - The union has drifted until only the discriminator and an id are truly shared, at which point "one collection" is providing nothing but a name. A hybrid is common and legitimate: keep a small, uniform, queryable summary document per item in a shared collection, and store the bulky variant-specific payload separately, referenced by id. The feed reads the shared collection; the detail view fetches the payload. ## Worked example A payments collection carrying `type` with the values `card`, `transfer` or `voucher`, plus `amount`, `currency`, `status` and `createdAt` on every document, plus a `details` sub-object holding the card's masked digits or the transfer's account reference. The transaction list sorts by `createdAt` across all three types using one index. The card-specific reconciliation job filters on the discriminator and reads `details`. Adding a fourth payment type adds one discriminator value and one shape of `details`; it does not touch the feed query, the index, or the shared core. ## What interviewers listen for The strong answer starts from the query, names the discriminator without being prompted, insists that shared fields keep identical names and types, and can state at least two conditions under which it would split the collection instead. The weak answer either splits by taxonomy reflexively, reproducing a table-per-type layout that then needs merging in application code, or dumps everything into one collection with no discriminator and infers the variant from which fields happen to be present.
- Why not infer the variant from which fields a document happens to contain?Because presence-based inference is not stable. As soon as a new variant carries the same field, or an old variant stops populating it, the inference misclassifies documents and nothing reports an error. An explicit discriminator with a closed set of values is unambiguous, cheap to index, and is the one rule a write-time check can enforce for every variant.
- What is the hybrid between one shared collection and one collection per variant?Keep a small uniform summary document per item in the shared collection — id, type, timestamp, status, amount — so feeds and listings need one query and one index, and store the bulky variant-specific payload in its own collection referenced by id. The list view stays fast and uniform; the detail view pays one extra read.
- What breaks first when the shared core is not type-stable across variants?Sorting and range queries. A timestamp stored as a string on one variant and a date type on another will not order correctly together, and range filters silently exclude documents. The result is a feed that looks plausible but is missing records, which is far harder to notice than an outright error.
saying these in an interview costs you the question
- Splits by taxonomy without asking how the data is queried
- Infers the variant from which optional fields are present
- Lets shared fields differ in name or type between variants
- Spreads dozens of variant-specific fields across the top level
- Forgets that every query must filter by the discriminator