What does allowDiskUse: true do in a MongoDB aggregate() call?
answer
- a per-operation aggregate() option
- about memory, never about results
- tied to the 100 MB per-stage budget
- lets $sort and $group use temp files
basics
~20 sallowDiskUse: true lets blocking aggregation stages such as $sort and $group write temporary files on disk when their working set exceeds the 100 MB per-stage memory budget, instead of the pipeline failing with a memory-limit error.
solid answer
~40 sEvery aggregation stage has a **100 MB RAM budget** for the state it accumulates. Streaming stages like `$match`, `$project` and `$limit` never come close, but blocking stages such as `$sort` and `$group` must hold state for the whole input, and on a large collection they can blow past it. `allowDiskUse: true` authorises those stages to use an external, on-disk algorithm — spilling temporary files under the server's data directory — rather than aborting the operation. From MongoDB 6.0 the server parameter `allowDiskUseByDefault` is true, so pipelines spill by default and you pass `allowDiskUse: false` to get the old hard failure back; before 6.0 you had to opt in per operation. Spilling keeps the query alive but is much slower than staying in memory, so treat it as a safety valve, not a tuning technique.
code
javascript · 8 linesdb.events.aggregate(
[
{ $match: { createdAt: { $gte: ISODate("2026-01-01") } } },
{ $group: { _id: "$sessionId", hits: { $sum: 1 } } },
{ $sort: { hits: -1 } }
],
{ allowDiskUse: true }
)go deeper
Be ready to say what the flag permits in one sentence: blocking stages may use temporary files instead of failing at the memory limit. Knowing that $sort and $group are the stages involved is enough at this level.
Explain the mechanics: a 100 MB per-stage budget, blocking versus streaming stages, and the switch to an external sort or group. Mention that MongoDB 6.0 made spilling the default via allowDiskUseByDefault.
Frame a spilling pipeline as a diagnosis, not a setting. Show that you would find out which stage holds the state and why, and that disk spilling competes for I/O with live traffic on the same node.
Own the policy: whether analytics pipelines are allowed to spill on production nodes at all, where such work should run, and what limits or alerting stop one ad-hoc aggregation from degrading the cluster for everyone.
## What the option is `allowDiskUse` is an option on the `aggregate` command (and on `db.collection.aggregate()` in `mongosh`, and on the corresponding driver method). It takes a boolean and applies to the whole pipeline execution, not to one stage: ```javascript db.orders.aggregate(pipeline, { allowDiskUse: true }) ``` It changes nothing about the *results* of the pipeline. It only changes what the server does when a stage runs out of the memory it is allowed to use. ## The 100 MB per-stage limit MongoDB caps the memory an individual aggregation stage may use for accumulated state at 100 MB. This is a per-stage budget, not a per-pipeline one, and it applies to the state the stage holds — not to the total volume of data that flows through it. A pipeline can happily stream a terabyte through `$match` and `$project`, because those stages look at one document at a time and forget it immediately. Blocking stages are different. `$sort` cannot emit its first output document until it has seen its last input document, so in the general case it holds the whole input. `$group` builds one entry per distinct grouping key, each carrying its accumulator state (`$sum`, `$push`, `$addToSet` …), so its footprint grows with the number of distinct keys — and `$push`/`$addToSet` can make each entry large on their own. ## What spilling actually does When a blocking stage exceeds its budget and disk use is permitted, it switches to an external algorithm: partial results are written to temporary files in the server's data directory and merged back later. The operation completes, but it has traded RAM for disk I/O. On a busy production node that I/O competes with everything else the server is doing, so a pipeline that spills is usually both slow itself and disruptive to neighbours. When disk use is *not* permitted, the stage fails the whole operation with a memory-limit error naming the offending stage and telling you to pass `allowDiskUse:true`. ## The default changed in 6.0 Historically the default was "fail loudly": you had to pass `allowDiskUse: true` yourself, and forgetting to do so was one of the most common aggregation errors. From MongoDB 6.0 the server parameter `allowDiskUseByDefault` defaults to true, so a pipeline that exceeds a stage's budget spills instead of erroring, and you can pass `allowDiskUse: false` on a specific operation to opt back into the hard failure. This is worth knowing in an interview because it explains why the same pipeline behaves differently on an older cluster: the failure mode moved from a visible error to a silent slowdown. ## Which stages care Only stages that accumulate state need it — `$sort` and `$group` are the two you should be able to name without hesitating. Streaming stages (`$match`, `$project`, `$addFields`, `$limit`, `$skip`, `$unwind`) hold at most a document (or a document's array) at a time, so the option is irrelevant to them. ## What allowDiskUse is not It does not raise the 16 MB BSON limit that applies to every document the pipeline emits — a `$group` that builds one enormous array with `$push` will still fail on document size no matter how much disk it is allowed. It does not grant a stage more RAM; the budget stays at 100 MB and the overflow goes to files. It does not make anything faster; if anything, a pipeline that starts spilling has just got slower. And it is not a substitute for the real fixes — filtering earlier so less data reaches the blocking stage, carrying fewer fields, or reshaping the work so the blocking stage sees a smaller input. ## How to talk about it The answer an interviewer wants is two-part: what the flag permits (external sort/group when the 100 MB stage budget is exceeded), and the judgment that needing it is a signal. "We turned on allowDiskUse and the job stopped failing" is a fine first-aid story; "we turned it on, then went and found out why a stage was holding that much state" is the senior version.
- Does turning on allowDiskUse make a large aggregation faster?No — it usually makes it slower. The stage abandons a purely in-memory algorithm for an external one that writes and re-reads temporary files. What it buys you is completion instead of a memory-limit error. If a pipeline is spilling regularly, the fix is to reduce what reaches the blocking stage (earlier index-backed `$match`, fewer carried fields, a smaller grouping key), not to leave it spilling.
- Which pipeline stages actually spill, and which never need to?Stages that accumulate state across the whole input spill — `$sort` and `$group` are the canonical pair. Streaming stages hold at most one document at a time: `$match`, `$project`, `$addFields`, `$limit`, `$skip` and `$unwind` pass data through and keep no growing buffer, so the per-stage memory budget is never a factor for them regardless of collection size.
- A $group with $push still fails after you enable allowDiskUse. Why?Almost certainly the 16 MB BSON document limit rather than the stage memory budget. `$push` builds an array inside a single output document, and every document a pipeline emits must fit in 16 MB. Disk use has no bearing on document size. The fix is to stop materialising the array — aggregate to counts or sums, or emit one document per element instead of one per group.
saying these in an interview costs you the question
- Claims allowDiskUse makes a large aggregation run faster
- Thinks it lets a result document exceed the 16 MB limit
- Says it raises the per-stage memory budget above 100 MB
- Cannot name a single stage that would actually spill
- Treats it as the permanent fix rather than a symptom