skip to content

Why should an edge_ngram autocomplete field in Elasticsearch use a different analyzer at search time?

level: middleimportance: must knowfreq 60%

answer

  1. The index already holds every prefix
  2. One typed word should become one term
  3. Very short terms match nearly every document
  4. Both sides expanding is the trap
  5. Nothing longer than max_gram was ever stored

basics

~20 s

Because n-gramming the query too turns each typed word into many short prefixes, which match almost everything and destroy precision and scoring. Prefixes belong in the index; the query should contribute the whole typed token as one term.

solid answer

~50 s

The `edge_ngram` design is intentionally asymmetric. At index time the analyzer expands each token into every prefix, so `"star"` stores `s`, `st`, `sta`, `star` and any typed prefix can be found with an exact term lookup. If the same analyzer also runs at search time, the query `"star"` is expanded to `s`, `st`, `sta`, `star` as well: the short terms match unrelated documents such as `"sunset"` and `"stop"`, so a `match` query with the default OR operator returns junk, and scoring is skewed because the very common one-character terms carry weight. Set `search_analyzer` to a plain analyzer — typically `standard` or the lowercasing chain used before the n-gram filter — so the query contributes exactly one term. The same reasoning explains why an index-time `edge_ngram` cannot match input longer than `max_gram`, since no such prefix was ever stored.

code

json · 32 lines
json
{
  "settings": {
    "analysis": {
      "analyzer": {
        "autocomplete_index": {
          "tokenizer": "standard",
          "filter": ["lowercase", "autocomplete_filter"]
        },
        "autocomplete_search": {
          "tokenizer": "standard",
          "filter": ["lowercase"]
        }
      },
      "filter": {
        "autocomplete_filter": {
          "type": "edge_ngram",
          "min_gram": 2,
          "max_gram": 20
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "title": {
        "type": "text",
        "analyzer": "autocomplete_index",
        "search_analyzer": "autocomplete_search"
      }
    }
  }
}

go deeper

for a junior

Recall that autocomplete stores prefixes in the index and that the search box's text should be looked up as a single whole token, not expanded again.

for a middle

Explain concretely why re-expanding the query wrecks precision and BM25 scoring, and write the mapping with the two analyzers correctly.

for a senior

Bring the operational consequences: index growth from n-gram expansion, the max_gram cutoff and its truncate fix, and when to choose a query-time prefix approach instead.

for a principal

Own the trade-off across the product: where prefix matching should live, what it costs in storage and write throughput, and when a suggester or a separate field is the better architecture.

## The shape of the data An `edge_ngram` tokenizer or token filter takes a token and emits every prefix of it between `min_gram` and `max_gram` characters. With `min_gram: 1` and `max_gram: 10`, indexing `"star wars"` (after lowercasing and splitting) stores the terms `s, st, sta, star` and `w, wa, war, wars`. The point is to convert a prefix *search* problem into an exact *term lookup* problem: the user types `sta`, you look up the term `sta`, and you get a posting list — no wildcard scan, no per-query prefix expansion. ## What goes wrong when both sides n-gram If the field declares only `analyzer` and no `search_analyzer`, the same expansion runs on the query. Typing `star` produces four query terms instead of one: - **Precision collapses.** A `match` query ORs its terms, so the term `s` alone pulls in every document whose text contains any word starting with `s`. The user typed four characters and got the whole corpus back. - **Scoring is distorted.** BM25 weights terms by inverse document frequency, and one-character prefixes are the most common terms in the index. They contribute noise, and documents can outrank better ones simply by containing many short-prefix hits. - **Phrase and position semantics get strange.** The n-grams of a single word all sit at the same position, so phrase-style clauses behave in ways nobody intends. - **The query gets expensive.** Four to ten term lookups per typed word, each on a very long posting list, for what should be one cheap lookup. Raising `minimum_should_match` to 100% is a common but wrong-headed patch: it makes the extra terms harmless only because every prefix of the typed word is also a prefix of the stored word, so you are paying for redundant clauses that can never discriminate. Fixing the analyzer is strictly better. ## The correct configuration Declare the n-gram chain as the field's `analyzer` and a plain chain as its `search_analyzer`. The search analyzer must apply the *same normalization* — lowercasing, ASCII folding, whatever the index chain did before n-gramming — and simply stop short of the n-gram step. Then the typed `sta` becomes the single term `sta`, which the index already contains for every document whose text has a word starting with `sta`. ## The max_gram consequence Index-time n-gramming stores prefixes only up to `max_gram`. If `max_gram` is 10 and the user types 14 characters, the query term is 14 characters long and no stored term equals it, so a document that clearly starts with that string returns no hit. Two standard mitigations exist: truncate the query side (a `truncate` token filter in the search analyzer, sized to `max_gram`), or set `max_gram` high enough for realistic input at the cost of index size. Notice that this failure mode only exists *because* the sides are asymmetric — it is the price of the design, and interviewers like to see that you know the price rather than only the trick. ## Index cost and the alternatives Every document's terms are multiplied roughly by the average token length, so the field's postings grow substantially; keep the n-gram field as a dedicated sub-field or separate field rather than n-gramming everything, and do not enable it on large bodies of text. Alternatives with different trade-offs exist: `match_phrase_prefix` and `match_bool_prefix` do prefix expansion at query time with no index blow-up but more per-query work; the `search_as_you_type` field type packs a similar structure behind a single mapping declaration; a completion suggester serves a separate, sorted, in-memory structure optimized for suggestions rather than full search. The n-gram approach wins when you need fast prefix matching combined with ordinary relevance ranking and filtering over the same field. ## Diagnosing it in an interview The tell that a field is misconfigured is trivially checkable: run `_analyze` with the field, then with the search analyzer, on the same string. If the query side yields several short tokens rather than one, the field is n-gramming twice, and the symptom the team reported — "autocomplete returns irrelevant results as soon as you type one letter" — is fully explained.

  • What happens when a user types more characters than max_gram?
    Nothing matches. The index only ever stored prefixes up to `max_gram`, so a longer query term has no equal term to find. Fix it either by truncating the query side to `max_gram` with a `truncate` token filter in the search analyzer, or by raising `max_gram` and paying for the extra terms on disk.
  • Why is raising minimum_should_match not a real fix for double n-gramming?
    Because every prefix of the typed word is also a prefix of any word it matches, so requiring all clauses does not exclude anything the single full-length term would not already exclude. You keep paying for redundant lookups on the longest posting lists in the index, and the scoring noise from very common short terms remains.
  • What does the edge_ngram approach cost compared with query-time prefix matching?
    Index size and write cost: each token becomes many terms, so postings grow roughly with average token length, and reindexing is required to change gram sizes. In exchange, queries are single exact term lookups with normal BM25 ranking and filtering, instead of per-query prefix expansion that grows with the number of matching terms.

saying these in an interview costs you the question

  • Uses the same n-gram analyzer on both sides and blames relevance tuning
  • Thinks the search analyzer changes which terms are stored
  • Fixes junk results with minimum_should_match instead of the analyzer
  • Unaware that input longer than max_gram cannot match
  • Applies edge n-grams to large text fields without considering index growth

context