What is MongoDB's outlier pattern, and how does it stop a few huge documents from shaping the whole schema?
answer
- a few documents are wildly bigger
- don't design for the celebrity account
- a flag marks the exceptional document
- remainder spills into overflow documents
- only flagged documents pay a second read
basics
~20 sThe outlier pattern keeps the common document shape optimal and handles the rare extreme case separately: a flag field marks documents whose data overflows into extra documents, and only flagged documents pay for the second lookup.
solid answer
~50 sSome collections have a long tail. Nearly every user has a few hundred followers; a handful have millions. If you design the follower list for the millions, every ordinary user carries the machinery — an extra collection, an extra query, a wider document — for a case that affects one in a hundred thousand. The outlier pattern refuses that trade: keep the array embedded for the normal case, and when a document exceeds a threshold set a marker such as `hasOverflow: true` and spill the remainder into overflow documents keyed back to the parent. Application code reads the parent, checks the flag, and issues the extra query only when it is set. The gains are that the 99.99% path stays a single point read and documents stay far from the 16 MB BSON limit; the costs are a branch in the application, a threshold you have to choose and revisit, and duplicated read logic that must be tested on both paths.
code
javascript · 8 linesconst user = db.users.findOne({ _id: id })
let followers = user.followerIds
// only outliers pay for the second query
if (user.hasOverflow) {
const pages = db.followerOverflow.find({ userId: id }).toArray()
followers = followers.concat(...pages.map(p => p.followerIds))
}go deeper
Know the idea in one line: keep the normal document shape simple and handle the rare enormous case with a flag plus extra documents, instead of penalizing every document.
Be able to describe both paths — how the write detects the threshold and sets the flag, and how the read branches on it — and why the flag must live in the document you already fetched.
Show that you would derive the threshold from a measured size distribution, monitor flagged documents and maximum document size, and test the rare branch that almost never executes in production.
Judge when the tail is rare enough for a conditional design at all, weigh two code paths against a uniformly split model, and set the policy for who owns the threshold as the data distribution shifts.
## Long tails break uniform schemas Real datasets are rarely uniform. Followers per account, comments per post, line items per order, transactions per customer — all follow steep distributions where the median is tiny and the maximum is enormous. Schema design under those conditions has a trap: you notice the extreme case, design for it, and impose its cost on everything. Concretely, suppose a `users` collection embeds `followerIds`. Almost every account has under a thousand. A few celebrity accounts have millions, which would blow past the 16 MB BSON document limit and, long before that, make every read of those documents ruinous. The tempting conclusion is "arrays don't work, move followers to their own collection". That is a defensible answer — but it makes the common read two queries forever, to accommodate accounts you could count on one hand. ## The pattern Keep the optimal shape for the common case; detect and divert the exception. ```json { "_id": "u-7", "handle": "kim", "followerCount": 640, "followerIds": ["u-11", "u-12"] } ``` and for the rare account: ```json { "_id": "u-9001", "handle": "popstar", "followerCount": 4200000, "hasOverflow": true, "followerIds": ["…first N…"] } ``` with the remainder in overflow documents such as `{ userId: "u-9001", page: 7, followerIds: [ … ] }` in a companion collection. The write path checks the size or count before appending; once the threshold is crossed it sets `hasOverflow` and starts writing to overflow documents instead. The read path fetches the user and, only if `hasOverflow` is set, issues the second query. ## Why the flag is the whole idea Without the marker, the application cannot know whether more data exists without asking, and asking is the cost you were trying to avoid. The flag makes the exception *self-declaring* in the document you already fetched. It also gives operations something to query: `db.users.find({ hasOverflow: true })` enumerates every outlier in the system, which is how you monitor whether the tail is growing and whether your threshold is still right. ## Choosing and revisiting the threshold Pick it from measurement, not intuition: look at the actual distribution of array length or document size and choose a cut that leaves the overwhelming majority in the simple path while keeping the largest embedded document a comfortable margin below 16 MB. Two things move over time — the distribution itself as the product grows, and the width of each element as you add fields to it. A threshold set in elements can silently become a threshold in megabytes. Alerting on the count of flagged documents and on maximum document size is the operational counterpart of the pattern. ## When it is the wrong answer If the tail is not rare — if a meaningful percentage of documents overflow — you no longer have outliers, you have a bimodal collection, and two code paths for a coin flip is worse than one uniform design. If the overflowed data must be queried and sorted globally ("most recent followers across all accounts"), splitting it across an embedded array and an overflow collection makes every such query awkward; a plain child collection is cleaner. And if the branch would be duplicated across many services, the complexity may cost more than the extra read it saves. ## Relationship to neighbouring patterns Outlier is often confused with **subset**, but they answer different questions. Subset is applied uniformly to every document: everyone keeps N embedded and the rest elsewhere. Outlier is applied *conditionally*: the vast majority keep everything embedded and only the exceptions split. They can also combine — subset everywhere, with the flag telling you whether the remainder is one page or ten thousand. Bucketing sits nearby too: when the outlier's overflow is itself huge, the overflow documents are usually bucketed pages rather than one document per element. ## How to answer it Name the distribution problem first, then the pattern: keep the common shape, mark the exception with a flag, spill the remainder, branch on the flag at read time. Then show judgment — how you'd pick the threshold from real data, how you'd monitor flagged documents, and the honest condition under which you'd abandon the pattern for a plain child collection.
- How do you choose the overflow threshold?From the measured distribution of array length and document size, not intuition. Pick a cut that leaves the overwhelming majority on the embedded path while keeping the largest embedded document a wide margin below the 16 MB BSON limit. Then monitor it: alert on the count of flagged documents and on maximum document size, because both element width and the distribution drift as the product grows.
- How does the outlier pattern differ from the subset pattern?Subset applies uniformly — every parent keeps a bounded slice embedded and the full set lives elsewhere, always. Outlier applies conditionally — nearly every document keeps everything embedded, and only the rare oversized one sets a flag and spills into overflow documents. Subset optimizes the common read for all; outlier protects the common read from the exceptions.
- When would you drop the pattern and just use a child collection?When the tail is not actually rare, so a large share of documents take the overflow branch and you are maintaining two paths for a coin flip; or when the data must be queried and sorted globally across all parents, which is awkward when half of it is embedded and half is not. A plain child collection with a good compound index is simpler in both cases.
saying these in an interview costs you the question
- Designs every document around the largest one
- Splits the data with no flag, so every read needs a probe
- Applies it when a large share of documents overflow
- Assumes MongoDB moves oversized documents automatically
- Sets the threshold once and never re-measures it