skip to content

In an Atlas Search index, when do you replace dynamic mapping with static field mappings?

level: middleimportance: should knowfreq 45%

answer

  1. One choice indexes everything, one indexes what you name
  2. Defaults are fine until you need a specific analyzer
  3. Autocomplete and facets are not default behaviours
  4. Every indexed field costs memory and write work

basics

~20 s

Dynamic mapping indexes every supported field with default settings, which is ideal for prototyping. Switch to static mappings once you need a chosen analyzer, autocomplete, facet or filter field types, or want to stop indexing fields nobody searches.

solid answer

~40 s

With `mappings.dynamic: true`, mongot indexes every field of every supported BSON type using default settings — strings get the `lucene.standard` analyzer, and new fields appear in the index automatically. That is a good default while you explore, and it is why a first `$search` "just works". You move to `mappings.dynamic: false` plus an explicit `fields` block when you need control: a different analyzer such as `lucene.keyword` for exact whole-value matching or `lucene.english` for stemming, an `autocomplete` type for type-ahead, `stringFacet`/`numberFacet`/`dateFacet` for facet buckets, or a `token` type for exact-match and sort. Static mappings also keep the index small: a collection full of large descriptions or blobs nobody searches costs disk, memory and indexing throughput for nothing. You can also map one field several ways with `multi`.

code

json · 14 lines
json
{
  "mappings": {
    "dynamic": false,
    "fields": {
      "title": [
        { "type": "string", "analyzer": "lucene.english" },
        { "type": "autocomplete", "tokenization": "edgeGram", "minGrams": 2, "maxGrams": 15 }
      ],
      "sku": { "type": "string", "analyzer": "lucene.keyword" },
      "genre": { "type": "stringFacet" },
      "inStock": { "type": "boolean" }
    }
  }
}

go deeper

for a junior

Know that the search index is a JSON definition on the collection, and that dynamic mapping is what makes a first $search work without configuring anything.

for a middle

Be able to name concrete triggers for static mappings: a specific analyzer, an autocomplete or facet field type, and cutting index size by not indexing unsearched fields.

for a senior

Show you treat index definitions as deployable artifacts — rebuilds on edit, memory footprint, and the silent failure mode where a query works only if the field was mapped for it.

for a principal

Own the policy: who reviews index definitions, whether search runs on dedicated Search Nodes, and how index size and rebuild windows are budgeted as the corpus grows.

## What the index definition is An Atlas Search index is a JSON definition attached to a collection and materialised by `mongot` into Lucene structures. Its core is the `mappings` object, which has two shapes. **Dynamic** — `{ "mappings": { "dynamic": true } }` — tells mongot to index every field it encounters, in every document, using the default treatment for that BSON type. Strings are analyzed with `lucene.standard`; numbers, dates, booleans and objectIds get their natural index types; nested objects and arrays are traversed. New fields that appear later are picked up without editing the definition. **Static** — `{ "mappings": { "dynamic": false, "fields": { ... } } }` — indexes only the paths you name, each with the type and options you choose. The two can be mixed: a static `fields` block can contain a subdocument with `dynamic: true` inside it, so you can pin the fields you care about and still let a nested object be indexed wholesale. ## Why dynamic is the right starting point Dynamic mapping removes the chicken-and-egg problem of schema-flexible data: you do not have to enumerate a shape you have not settled yet, and a `$search` with `text` over a wildcard path works immediately. For a prototype, a small collection, or a genuinely heterogeneous document set, that is often the correct permanent answer too. ## The four reasons to go static **Analyzer control.** The default `lucene.standard` tokenizes on word boundaries and lowercases, so a query for `york` matches a title containing "New York City". Sometimes that is wrong. `lucene.keyword` indexes the whole field value as a single token, giving exact, case-sensitive whole-string matching — right for SKUs, tags and status codes. `lucene.english` adds stemming and stop-word removal so `running` matches `run`, at the cost of precision. The `analyzer` key sets the index-time analyzer and `searchAnalyzer` the query-time one; they normally have to agree, and mismatching them deliberately (for example an edge-gram index analyzer with a plain search analyzer) is an advanced technique, not a default. Custom analyzers let you assemble a char filter, tokenizer and token filter chain when none of the built-ins fit. **Feature types.** Several capabilities exist only as explicit field types. Type-ahead needs `type: "autocomplete"`, whose `tokenization` (`edgeGram`, `nGram`, `rightEdgeGram`) plus `minGrams`/`maxGrams` decide how prefixes are indexed — and it costs real index size, so you enable it per field. Facets need `stringFacet`, `numberFacet` or `dateFacet`. Exact matching with `equals`, and sorting on a string, need the `token` type. A vector field for `$vectorSearch` lives in a separate `vectorSearch`-type index definition entirely. **Index size and write cost.** Every indexed field is Lucene work on every change. A collection whose documents carry a long HTML body, a base64 payload or an audit trail nobody queries pays for all of it under dynamic mapping. Naming the six fields you actually search can shrink the index dramatically, which matters because search indexes want to be resident in memory. **Predictability.** Under dynamic mapping, a new field introduced by a producer silently changes the index and can change ranking. A static definition makes the searchable surface a reviewed artifact. ## Multi-analyzer fields You rarely want a field indexed one way only. The `multi` option indexes the same path under alternative analyzers, and a query names the variant through the path object. A product name can be indexed once with `lucene.standard` for recall and once with `lucene.keyword` for exact matches, and a `compound` query can boost the exact variant above the analyzed one. ## Changing the definition Editing an index definition triggers a rebuild by mongot; the old index keeps serving until the new one is ready, but a large collection takes time and consumes resources, which is an argument for dedicated Search Nodes. Treat index definitions like migrations: version them with the application, because a query using `autocomplete` against a field that was never mapped that way returns nothing rather than an obvious error. ## What to say in an interview Dynamic is the fast path and a legitimate production choice for small or unpredictable data. Static earns its keep the moment you need a specific analyzer, an autocomplete or facet type, or you want to stop indexing fields nobody searches.

  • What does lucene.keyword do differently from lucene.standard?
    `lucene.standard` splits text into lowercased tokens, so a query for `york` matches "New York City". `lucene.keyword` emits the entire field value as one token, so only the full string, case included, matches. Use it for SKUs, status codes and tags where partial word matches are wrong.
  • What is the difference between analyzer and searchAnalyzer in a field mapping?
    `analyzer` transforms the field value when the document is indexed; `searchAnalyzer` transforms the query text at query time. They usually match, because tokens only compare if both sides were processed the same way. Deliberately differing them supports patterns like edge-gram indexing with plain-token queries.
  • What happens to a running index when you edit its definition?
    mongot rebuilds it. The existing index keeps answering queries until the new one is ready, so there is no outage, but the rebuild consumes CPU and memory on whichever nodes run mongot and can take a long time on a large collection — a reason to isolate search onto dedicated Search Nodes.

saying these in an interview costs you the question

  • Thinks dynamic mapping enables autocomplete and facets automatically
  • Believes indexing every field is free because search is separate
  • Assumes lucene.standard gives exact whole-string matching
  • Queries an autocomplete operator against a plain string mapping
  • Sets analyzer and searchAnalyzer differently by accident

context