What can a MongoDB text index do, and what restriction applies per collection?
answer
- The key spec holds a type name, not a direction
- Words are stored, not raw values
- There is a hard count limit per collection
- Relevance is opt-in, not automatic
- Searching inside a word finds nothing
basics
~20 sA text index tokenizes string fields into stemmed, case- and diacritic-insensitive terms so $text queries can match words in them. A collection may have at most one text index, though that single index can cover many fields.
solid answer
~40 sYou create it with a `"text"` key type — `createIndex({ title: "text", body: "text" })` — and query it with `$text: { $search: "..." }`. The index tokenizes the indexed string fields into terms, applies language-specific stemming and stop-word removal, and matches case- and diacritic-insensitively, so a search for `running` matches `Run`. The hard structural limit is **one text index per collection**, so covering more fields means dropping and recreating it with all of them, optionally weighting fields with the `weights` option; `{ "$**": "text" }` indexes every string field. Relevance is not returned by default — you project and sort on `{ score: { $meta: "textScore" } }` for that. The big expectation-setter is that it is word-based: no substring or prefix matching, so searching for part of a word finds nothing.
code
javascript · 9 linesdb.articles.createIndex(
{ title: "text", body: "text" },
{ weights: { title: 10, body: 1 }, default_language: "english" }
)
db.articles.find(
{ $text: { $search: "\"index types\" -deprecated" } },
{ score: { $meta: "textScore" } }
).sort({ score: { $meta: "textScore" } })go deeper
Recall the two facts: a text index is declared with the "text" key type and queried with $text, and a collection can only have one of them.
Explain tokenization, stemming and case/diacritic insensitivity, and show how relevance is retrieved with $meta textScore rather than appearing automatically.
Set expectations in production terms: no substring matching, one text index shared by every query shape, and the drop-and-recreate cost of adding a searchable field.
Own the boundary decision — when keyword filtering inside the database is sufficient and when the search workload should move to a purpose-built search system instead.
## What it is A text index is a special index type declared by putting the string `"text"` where a direction would go: ``` db.articles.createIndex({ title: "text", body: "text" }) ``` Instead of storing field values, it stores **terms**: each indexed string field is split into words, stop words for the configured language are removed, and each remaining word is reduced to a stem. The resulting terms are what the index holds, so the index is a mapping from word-stems to documents. ## How you query it The query operator is `$text` with a `$search` string: ``` db.articles.find({ $text: { $search: "mongo indexing" } }) ``` Space-separated terms are ORed by default; a phrase in double quotes must appear as a phrase; a term prefixed with `-` is excluded. Matching is case-insensitive and diacritic-insensitive, and it is stem-based, so `running`, `runs` and `ran`-style variants collapse to a common root depending on the language rules. ## The one-per-collection rule A collection may have **at most one** text index. It may span many fields, but you cannot have two separate text indexes on two different field sets. In practice that means adding a searchable field is a drop-and-recreate operation on the single text index, and it makes the text index a shared, collection-wide resource rather than something each query pattern gets its own copy of. Two related options matter when the one index has to serve everything: - **`weights`** — per-field multipliers that make matches in some fields count more towards relevance than others, for example title over body. - **`default_language`** — chooses the stemming and stop-word rules; a per-document language override field is also supported. - **`{ "$**": "text" }`** — a wildcard text index over every string field in every document. Convenient, and correspondingly expensive to build and maintain. ## Relevance scoring Text matches carry a relevance score, but it is not part of the result unless you ask for it. You project it and sort by it with the `$meta` expression: ``` db.articles.find( { $text: { $search: "indexing" } }, { score: { $meta: "textScore" } } ).sort({ score: { $meta: "textScore" } }) ``` Sorting by an ordinary field instead is legal and simply ignores relevance. ## What it deliberately does not do This is where candidates most often over-claim: - **No substring or infix matching.** The index holds whole word-stems, so searching for a fragment inside a word finds nothing. It is not a `LIKE '%x%'` replacement. - **No prefix/autocomplete matching** out of the box. - **One `$text` expression per query.** You cannot combine two independent text searches in a single filter. - **In an aggregation pipeline, `$text` must appear in a `$match` at the start of the pipeline**, not after arbitrary stages. - **Compound restrictions.** A compound index that includes a text key can carry ordinary equality-prefix keys alongside it, but it cannot combine a text key with other special index types such as geospatial keys. ## When it is the right tool A text index is the built-in, self-managed option for word matching within a collection: good enough for filtering on keywords, tagging, and simple search boxes over modest corpora. Once the requirements grow into analyzers, faceting, fuzzy matching or ranked multi-field relevance tuning, teams generally move that workload to a dedicated search product rather than pushing the text index further. ## Answering well Name the declaration syntax and the `$text` operator, state the one-index-per-collection rule plainly, then volunteer the limitation that matters in practice — word-based, no substring matching — and mention `$meta: "textScore"` for relevance. That combination shows you have actually used it rather than read the feature list.
- How do you make matches in the title field count more than matches in the body?Pass the weights option when creating the text index, giving each indexed field a multiplier — for example title 10 and body 1. Weights affect only the relevance score, not whether a document matches. You then have to project and sort on { score: { $meta: "textScore" } } for the weighting to have any visible effect.
- A user searches for a fragment in the middle of a product name and gets nothing. Why?A text index stores whole word-stems, so it matches words, not substrings. There is no infix matching, and no prefix matching either. Fragment and autocomplete-style search needs a different mechanism; the built-in text index is not a substitute for a wildcard LIKE pattern.
- How do you get the relevance score of a $text match into the results?Project it with { score: { $meta: "textScore" } } and, if you want ranked output, sort on the same expression. Without that projection the score is computed for matching but never returned, and the result order is unspecified rather than relevance-ranked.
saying these in an interview costs you the question
- Thinks a collection can have several text indexes
- Expects substring or infix matching from $text
- Assumes results come back relevance-ranked by default
- Says text search is case-sensitive
- Believes each searchable field needs its own text index