skip to content

What does $unwind do to a document whose array field is empty or missing?

level: middleimportance: must knowfreq 72%

answer

  1. one output document per array element
  2. zero elements means zero documents out
  3. the short form silently loses documents
  4. an option restores the dropped parent
  5. preserveNullAndEmptyArrays is the outer-join switch

basics

~20 s

$unwind emits one document per array element, copying the other fields. A document whose array is empty, missing or null produces no output at all — it is dropped — unless you use the object form with preserveNullAndEmptyArrays: true, which passes it through once.

solid answer

~40 s

`$unwind` turns each array element into its own document, repeating every other field, so a document with three tags becomes three documents differing only in `tags`. The trap is the empty case: if the path is an empty array, missing, or `null`, the document is **dropped entirely** by the short form `{ $unwind: "$tags" }`. That silently turns an unwind into an inner-join-like filter and is a common cause of "my count went down". The object form fixes it: `{ $unwind: { path: "$tags", preserveNullAndEmptyArrays: true } }` lets such a document through once. The object form also offers `includeArrayIndex: "idx"`, which records each element's original position. Since MongoDB 3.2 a path holding a non-array, non-null value is treated as a one-element array, so the document passes through unchanged rather than erroring.

code

javascript · 8 lines
javascript
// input:
// { _id: 1, sku: "A", tags: ["x", "y"] }
// { _id: 2, sku: "B", tags: [] }
db.items.aggregate([{ $unwind: "$tags" }])
// output:
// { _id: 1, sku: "A", tags: "x" }
// { _id: 1, sku: "A", tags: "y" }
// _id 2 produced nothing

go deeper

for a junior

Know that $unwind produces one document per array element and that you usually pair it with a following $group. Recognise the object form with a path key as the alternative spelling.

for a middle

Explain the empty/missing/null case and name preserveNullAndEmptyArrays as the fix. Be able to state that a non-array scalar passes through as a single-element array in current versions.

for a senior

Show that you think about fan-out: filter and project before unwinding, know when $filter avoids the expansion entirely, and be able to explain why a report's totals dropped after an unwind was added.

for a principal

Own the modelling consequence: if every read has to unwind a large embedded array, that array may belong in its own collection. Frame it as a document-shape decision, not a pipeline tweak.

## What the stage does `$unwind` is the pipeline's array-flattening stage. For each input document it looks at the field named by `path` and, if that field holds an array, emits one output document per element. Every other field is copied verbatim; the array field itself is replaced by the single element. A document with `tags: ["x", "y"]` becomes two documents, one with `tags: "x"` and one with `tags: "y"`. This is the standard preparation step for grouping *by* array elements: you cannot ask "how many documents carry each tag" with `$group` alone, because the grouping key would be the whole array. Unwind first, then group on the now-scalar field. ## The empty, missing and null cases The short form has an important asymmetry. If the path resolves to an **empty array**, is **missing** from the document, or is **null**, there are zero elements to emit — and `$unwind` emits nothing, so the document disappears from the pipeline. Nothing warns you. A pipeline that reads `{ $match: {...} }, { $unwind: "$items" }, { $group: {...} }` quietly excludes every order with no items, and the totals downstream are wrong in a way that looks plausible. The object form is the fix: ```javascript { $unwind: { path: "$items", preserveNullAndEmptyArrays: true } } ``` With that option the document passes through once instead of being dropped, with no unwound element in that field. Reach for it whenever you are unwinding an optional array and the parent document must survive — for example a "customers and their orders" report where customers with no orders still need a row. ## Non-array values Since MongoDB 3.2, a path that holds a scalar — a string, a number, a sub-document — is treated as a one-element array: the document is emitted once, unchanged. This means `$unwind` is safe over a field whose type is inconsistent across documents, a common situation in a schema-flexible collection where some documents stored `tag: "x"` and others `tags: ["x", "y"]`. Earlier behaviour raised an error, which is worth knowing only because old answers still circulate. ## includeArrayIndex The object form takes `includeArrayIndex: "<newField>"`, which adds a field holding the zero-based position of the element within its original array. That is how you preserve ordering information that unwinding otherwise throws away — useful when you must later re-sort elements into their document order, or when you want "the first line item" identified positionally rather than by value. Combined with `preserveNullAndEmptyArrays`, documents that had nothing to unwind get `null` in that index field. ## Fan-out is the cost `$unwind` multiplies documents. A collection of 1 million orders averaging 8 line items produces 8 million documents in the stream, each carrying a full copy of the order's non-array fields. Everything downstream — sorts, groups, further stages — works on that inflated stream. Two habits keep it under control: - **Filter before you unwind.** Put the selective `$match` on the parent document first, so you never expand documents you are about to discard. - **Project away what you do not need before unwinding**, so the copies are small. If the order has a large `notes` field and your report only needs `items.sku`, dropping `notes` first cuts the bytes copied per element. You can also filter *after* the unwind on the element itself — `{ $unwind: "$items" }, { $match: { "items.sku": "A1" } }` — but note the semantics are different from matching before: the pre-unwind form `{ "items.sku": "A1" }` keeps whole documents that contain such an element, while the post-unwind form keeps only the matching elements themselves. Interviewers like this distinction because it is exactly the mistake that produces wrong per-element aggregates. ## Undoing an unwind A pipeline that unwinds, filters or transforms elements, and then needs the parent shape back rebuilds it with `$group`: group on `"$_id"`, `$push` the transformed elements into an array, and `$first` any parent fields you need to keep. The result is the original document with a filtered array. It is a well-worn idiom, and it is also a reminder that unwinding to filter array elements is not free — for simple cases `$filter` inside `$project`/`$set` does the same work without expanding the stream at all. ## Quick mental model Think of `$unwind` as a stream-level cross product between the document and its array. Zero elements means zero rows out — the same reason an inner join drops the parent row when the child table has nothing. `preserveNullAndEmptyArrays` is the outer-join switch.

  • What does includeArrayIndex give you, and when do you need it?
    `includeArrayIndex: "idx"` adds a field holding each element's zero-based position in its original array. Use it when unwinding destroys ordering you still need — restoring document order after a sort, or identifying "the first line item" positionally. Documents that had nothing to unwind, when combined with preserveNullAndEmptyArrays, get null in that field.
  • How is $match on an array field before $unwind different from the same $match after it?
    Before the unwind, `{ "items.sku": "A1" }` keeps whole documents that contain at least one matching element — every element stays. After the unwind, the same predicate keeps only the matching elements themselves, since each element is now its own document. Aggregates over the two differ, and confusing them is a classic source of inflated totals.
  • When can you avoid $unwind entirely for array work?
    When you only need to filter or transform elements in place. `$filter` inside `$set`/`$project` narrows an array without expanding the stream, `$map` transforms each element, `$size` counts, and `$reduce` folds. Unwinding is required mainly when you must group *across* documents by element value or sort the stream by element.

saying these in an interview costs you the question

  • Assumes a document with an empty array passes through unchanged
  • Thinks $unwind errors on a non-array field in current versions
  • Unwinds before filtering the parent documents
  • Believes unwinding preserves element order without includeArrayIndex
  • Says $unwind modifies the stored documents

context