What makes a MongoDB index multikey, and what does it store for an array field?
answer
- You do not declare this kind of index
- One document, many index entries
- Element-level keys, not the whole array
- Arrays of subdocuments behave the same way
- Index size follows total elements, not documents
basics
~20 sAn index becomes multikey automatically the first time a document holds an array in an indexed field. MongoDB then stores one index key per array element, so a document with five tags contributes five entries pointing at that one document.
solid answer
~50 sYou never declare a multikey index — `createIndex({ tags: 1 })` is an ordinary spec, and MongoDB flags the index multikey as soon as it indexes a document whose `tags` field is an array. For each such document it writes **one key per array element**, all pointing back at the same document, so `{ tags: ["red", "blue"] }` produces two entries. That is what makes `find({ tags: "red" })` an index lookup rather than a scan: the query value is compared against individual elements, not against the array as a whole. It works through arrays of subdocuments too — indexing `items.sku` produces one key per element of `items`. The costs follow from the same mechanic: index size grows with total array length rather than document count, and once the multikey flag is set on an index it stays set even if all the arrays are later removed, until the index is dropped and rebuilt.
code
javascript · 4 linesdb.products.insertOne({ _id: 1, name: "shirt", tags: ["red", "cotton", "sale"] })
db.products.createIndex({ tags: 1 })
// index now holds three keys for _id:1 -> "cotton", "red", "sale"
db.products.find({ tags: "sale" })go deeper
Recall that indexing an array field just works and that a query for a single element uses the index. You do not have to declare anything special to get it.
Explain the one-key-per-element mechanic, show it on an array of subdocuments such as items.sku, and state how index size grows with total element count.
Bring the operational consequences: unbounded arrays inflating index size and write cost, the sticky multikey flag surviving data cleanup, and looser index bounds on array predicates.
Own the modelling rule that follows: cap or externalize growing arrays at design time, since the index for an unbounded array is a cost that never stops climbing across the system's life.
## The definition A multikey index is an ordinary index whose indexed field turns out to contain arrays. There is no `multikey: true` option — the flag is set by the server, in the collection catalog, the first time a document with an array value in that field is indexed. From then on the planner treats the index as multikey and applies multikey rules to it. ## What actually gets stored For a scalar field, one document contributes one index key. For an array field, one document contributes **one key per element**: ``` { _id: 1, tags: ["red", "blue", "red"] } ``` with `createIndex({ tags: 1 })` produces index entries for `"blue"` and `"red"` (duplicate elements do not produce duplicate keys for the same document), each pointing at `_id: 1`. Because the elements are indexed individually, `find({ tags: "red" })` is a normal index seek: MongoDB looks up the key `"red"` and finds the document. This is exactly the mental model people arriving from SQL do not have — there, an array-ish column would be an opaque value and you would need a join table. ## Arrays of subdocuments The same rule applies one level down. Given ``` { _id: 7, items: [ { sku: "A1", qty: 2 }, { sku: "B4", qty: 1 } ] } ``` an index on `{ "items.sku": 1 }` stores keys `"A1"` and `"B4"`. This is the single most common multikey case in real schemas: the order-lines-inside-an-order shape, indexed by the field inside the line. You can also index the whole subdocument (`{ items: 1 }`), but then the key is the entire embedded document and matching requires exact field order and full equality — almost never what you want. ## Compound indexes that are multikey A compound index may be multikey. `createIndex({ customerId: 1, "items.sku": 1 })` is legal and useful: keys are the cross product of the (single) `customerId` value with each element of `items`, so a document with three items writes three keys. The hard rule is that at most **one** of the indexed fields may be an array in any given document — a document with arrays in two indexed fields is rejected as a parallel-array error. ## Costs and consequences - **Index size tracks total array length, not document count.** A collection of a million documents each carrying fifty tags produces fifty million index keys. Every insert writes them all; every update to the array rewrites the affected keys. This is the reason unbounded arrays are a modelling smell, not just a document-size concern. - **The multikey flag is sticky.** Once set it is not cleared when the arrays disappear, because the server does not rescan the collection to prove no arrays remain. Only dropping and rebuilding the index clears it. That matters because the planner is more conservative with a multikey index. - **Index bounds can be looser.** For a query with two bounds on the same array field, the index cannot always tighten to a single contiguous range, because different elements may satisfy different halves of the predicate. The planner may use one bound from the index and re-check the rest against the documents. - **Covering is limited.** A query whose projection needs the array field itself cannot be answered from the index alone, since the index holds individual elements rather than the array. ## How to talk about it The strong answer states the mechanic (one key per element, set automatically), shows the array-of-subdocuments case because that is what production schemas look like, and then names one consequence — index-size growth with array length, or the parallel-array restriction — rather than reciting the definition alone.
- How large does the index get for a million documents that each carry fifty tags?Roughly fifty million index keys, because a multikey index stores one key per array element rather than one per document. Index size, build time, and per-write maintenance all scale with total element count. That is a strong argument against unbounded arrays: the document may stay small while the index for it grows without limit.
- If every array is later replaced by a scalar, does the index stop being multikey?No. The multikey flag is stored on the index in the collection catalog and is not cleared when arrays disappear, because the server will not rescan the collection to prove none remain. Dropping and rebuilding the index is what clears it. Until then the planner keeps applying the more conservative multikey rules.
- Can an index on items.sku, where items is an array of subdocuments, be multikey?Yes, and this is the common production case. Each element of items contributes one key holding that element's sku value, all pointing at the same document. A query on items.sku is then an index seek. The array being one level down changes nothing about the one-key-per-element rule.
saying these in an interview costs you the question
- Says you must pass a multikey option to createIndex
- Thinks the whole array is stored as one index key
- Assumes array fields cannot be indexed at all
- Claims the multikey flag clears once arrays are removed
- Ignores that index size grows with total array elements