What tokens do the ngram and edge_ngram tokenizers emit, and how does that drive index size?
answer
- One anchors, the other slides
- Prefix versus substring matching
- Term count per word differs by word length
- Defaults are one and two
- An index setting caps the range
basics
~20 sedge_ngram emits only prefixes anchored at the start, so Quick with grams 2 to 4 gives Qu, Qui, Quic. ngram emits substrings at every position, producing far more terms and a much larger index for infix matching.
solid answer
~40 sBoth tokenizers cut text into character grams; the difference is anchoring. `edge_ngram` emits grams that all start at the beginning of the token, so `Quick` with `min_gram: 2` and `max_gram: 4` yields `Qu`, `Qui`, `Quic` — roughly `max_gram - min_gram + 1` terms per word, which is what makes it the standard prefix-autocomplete tool. `ngram` slides a window over every position, so the same word yields grams starting at `Q`, `u`, `i`, `c`, and `k` — order-of-length-times-gram-range terms, the price of matching in the middle of a word. Both default to `min_gram: 1` and `max_gram: 2`, which is almost never what you want, and `index.max_ngram_diff` (default 1) rejects a wider range until you raise it. `token_chars` controls what counts as part of a token; left empty, the tokenizer grams whitespace and punctuation too.
code
json · 22 linesPUT /products
{
"settings": {
"index": { "max_ngram_diff": 4 },
"analysis": {
"tokenizer": {
"prefix_tok": {
"type": "edge_ngram",
"min_gram": 2,
"max_gram": 10,
"token_chars": ["letter", "digit"]
},
"infix_tok": {
"type": "ngram",
"min_gram": 3,
"max_gram": 5,
"token_chars": ["letter", "digit"]
}
}
}
}
}go deeper
Be able to state the outputs: edge grams are prefixes of the word, plain ngrams are substrings from every position, and both are configured with min_gram and max_gram.
Explain the term-count arithmetic that follows from anchoring, why the defaults of 1 and 2 are almost never right, and what token_chars changes about word boundaries.
Show judgment about index size and query behaviour: choose grams on a dedicated subfield, pick a min_gram that does not match a third of the corpus, and know the alternatives before reaching for n-grams at all.
Weigh n-gram indexing against other approaches for the same feature and against its ongoing cost in storage, merge pressure, and cluster capacity, rather than treating autocomplete as a free feature.
## The two shapes Character n-grams let a search engine match on fragments of words rather than whole terms. Elasticsearch ships two tokenizers for this and the choice between them is mostly a choice about index size. `edge_ngram` produces grams anchored at the start of the token. With `min_gram: 2` and `max_gram: 4`, the input `Quick` produces `Qu`, `Qui`, `Quic` — and nothing else. The number of terms per word is bounded by the gram range, independent of how long the word is (beyond `max_gram`). `ngram` produces grams starting at every position. With `min_gram: 1` and `max_gram: 2`, the documented behaviour on `Quick Fox` is a stream that includes `Q`, `Qu`, `u`, `ui`, `i`, `ic`, `c`, `ck`, `k`, and — because the default `token_chars` is empty — grams containing the space as well. Term count grows with word length multiplied by the gram range. ## What each one buys you Edge grams answer prefix queries as ordinary term lookups. Index `Qu`, `Qui`, `Quic` at index time and a user typing `qui` matches an exact term — no wildcard, no prefix scan, constant-ish cost per query. That is why edge grams underpin as-you-type search. Full n-grams answer infix and substring queries: finding `nick` inside `Ravenclaw-Nickerson`, or matching a part number in the middle of a longer string. Nothing about prefix matching needs them, and using them where edge grams would do is one of the classic causes of an index several times larger than the source data. ## The parameters that matter `min_gram` and `max_gram` bound the gram lengths and both default to `1` and `2`. Those defaults are rarely useful — one-character grams match nearly everything and produce enormous postings lists — so you set them explicitly nearly every time. `index.max_ngram_diff` is an index-level setting, default `1`, that caps `max_gram - min_gram` for the ngram tokenizer and filter. Creating an index with `min_gram: 1, max_gram: 10` fails until you raise it, and the setting exists exactly because that configuration is a footgun. The analogous cap for the shingle filter is `index.max_shingle_diff`. `token_chars` lists the character classes that may appear inside a token: `letter`, `digit`, `whitespace`, `punctuation`, `symbol`, and `custom` (paired with `custom_token_chars`). The tokenizer splits on everything not listed. The default is an empty list, meaning nothing is treated as a separator and the entire input — spaces and all — is grammed as one run. Setting `["letter", "digit"]` is the usual fix so grams stay inside words. ## Tokenizer versus token filter Both n-gram flavours also exist as token filters. The tokenizer form splits the raw input directly; the filter form runs after some other tokenizer has produced words, and grams each word. In practice the filter form is more common and more controllable: `standard` tokenizer, then `lowercase`, then an `edge_ngram` filter, gives grams of clean, lowercased words without having to encode word-splitting rules into `token_chars`. Note the filter is bounded by `min_gram`/`max_gram` as well, and the same `index.max_ngram_diff` cap applies. ## Sizing the cost honestly For a corpus of average word length L and a gram range of R: - edge grams add on the order of `R` terms per word, capped by word length. - full grams add on the order of `L × R` terms per word. So the difference is not marginal; it is the difference between an index that grows by a modest multiple and one that can grow several-fold. On top of raw size, the term dictionary grows, merges get more expensive, and very short grams create postings lists that match most documents, which hurts both scoring quality and query latency. ## Practical guidance Start with the question the feature actually needs. Prefix completion: edge grams, `min_gram` around 2 or 3 so you do not index a term matching a third of the corpus, `max_gram` around the length beyond which the user has typed enough to be unambiguous. Substring search inside identifiers: full n-grams on a dedicated subfield, not on the main text field, with a tight gram range. Hierarchical paths: neither — `path_hierarchy` is the right tokenizer for `/a/b/c`, emitting `/a`, `/a/b`, `/a/b/c`. And measure. `_analyze` shows you the exact token stream for a sample input, and comparing index sizes for two candidate chains on a realistic sample is a ten-minute experiment that settles most arguments.
- What is the difference between the edge_ngram tokenizer and the edge_ngram token filter?The tokenizer grams the raw input directly and decides word boundaries itself through `token_chars`. The filter runs after another tokenizer, so it grams tokens that have already been split and possibly lowercased or folded. The filter form composes better: you get proper word segmentation from `standard` and can normalize before gramming.
- Which tokenizer would you choose for indexing filesystem-style paths such as /usr/local/bin?`path_hierarchy`. It emits successive prefixes — `/usr`, `/usr/local`, `/usr/local/bin` — so a term query for a parent directory matches every descendant document. Its `delimiter` defaults to `/` and can be changed, and `reverse` handles suffix-style hierarchies such as domain names. N-grams would be both wrong and enormously wasteful here.
- Why does creating an index with min_gram 1 and max_gram 10 fail by default?`index.max_ngram_diff` defaults to `1`, capping the difference between `max_gram` and `min_gram` for the ngram tokenizer and filter. The guard exists because a wide gram range multiplies term count dramatically. You can raise the setting deliberately, but the failure is a prompt to check whether that range is really needed.
saying these in an interview costs you the question
- Says edge_ngram matches substrings anywhere in a word
- Uses full ngrams for plain prefix autocomplete
- Leaves min_gram and max_gram at their defaults
- Ignores token_chars so whitespace gets grammed
- Assumes n-gram indexes cost about the same as normal ones