When would you add a shingle token filter to an Elasticsearch analysis chain, and what does it cost?
answer
- Word n-grams, not character n-grams
- Adjacency becomes an ordinary term
- Far more pairs than single words
- Put it on its own subfield
- Removed tokens leave a placeholder
basics
~20 sShingles turn adjacent tokens into word n-grams, so word adjacency becomes an ordinary term you can match and boost. The cost is a much larger term dictionary, so they belong on a dedicated subfield rather than the main text field.
solid answer
~40 sThe `shingle` filter concatenates adjacent tokens into fixed-size word n-grams: `the quick fox` becomes `the quick` and `quick fox`. That converts adjacency into a plain term, which is useful when you want to reward documents where query words appear next to each other without paying full phrase-query cost, and for suggestion or language-model-style features that need bigram statistics. Configuration is `min_shingle_size` and `max_shingle_size` (both default 2), `output_unigrams` (default true), plus `token_separator` and `filler_token` — the latter marking gaps where a previous filter removed a token. `index.max_shingle_diff` (default 3) caps the size range. The cost is term explosion: bigrams are far more numerous than unigrams, so the term dictionary and index grow substantially. The usual pattern is a dedicated shingled subfield with `output_unigrams: false`, boosted in a `multi_match`, leaving the main field untouched.
code
json · 16 lines"analysis": {
"filter": {
"bigrams": {
"type": "shingle",
"min_shingle_size": 2,
"max_shingle_size": 2,
"output_unigrams": false
}
},
"analyzer": {
"title_shingled": {
"tokenizer": "standard",
"filter": ["lowercase", "bigrams"]
}
}
}go deeper
Know that the shingle filter joins neighbouring words into combined tokens such as "quick brown", and that this is about word pairs rather than characters.
Explain the parameters and their defaults — shingle sizes, output_unigrams, the filler token — and why the resulting terms are far more numerous than the originals.
Show the deployment judgment: a dedicated subfield, unigrams off, a measured boost, awareness of stopword interaction, and a before-and-after read on both index size and relevance.
Frame it as a capacity decision as much as a relevance one, and compare it against other levers for the same goal before committing every index in the fleet to a materially larger term dictionary.
## What shingles are A shingle is a word n-gram — the token-level analogue of a character n-gram. The `shingle` token filter takes the stream `the`, `quick`, `brown`, `fox` and, with a shingle size of 2, emits `the quick`, `quick brown`, `brown fox`. Because these are ordinary tokens, they land in the inverted index as ordinary terms and can be matched, counted, and scored like any other term. ## Why you would want them **Adjacency as a term.** A phrase query proves adjacency by reading positions from the postings, which costs more than a term lookup. Indexing bigrams turns "these two words were adjacent" into a single term match. On a field where phrase-like matching drives relevance, matching a shingled subfield is a cheap way to reward word order. **Relevance boosting.** The common pattern is a `multi_match` across the plain field and a shingled subfield with a higher boost on the shingled one. Documents containing the exact word pair score above documents that merely contain both words somewhere. This gives a phrase-ish ranking signal without turning the whole query into a phrase query. **Bigram statistics.** Features that need to know how often two words co-occur — suggesters, query analysis, ranking experiments — read those counts straight from the shingled field's term statistics. ## The parameters `min_shingle_size` and `max_shingle_size` both default to 2, giving bigrams only. Raising `max_shingle_size` to 3 adds trigrams, multiplying terms again. `index.max_shingle_diff`, an index-level setting with a default of 3, caps the difference between the two sizes, for the same reason the n-gram cap exists. `output_unigrams` defaults to `true`, meaning the original single tokens are emitted alongside the shingles. On a dedicated subfield that is usually wrong — you already have the unigrams in the main field, and keeping them doubles the subfield's work and muddies its scoring. Set it to `false` there. The related `output_unigrams_if_no_shingles` emits unigrams only when no shingle could be produced, which matters for single-word inputs that would otherwise analyze to nothing. `token_separator` (default a single space) joins the words in a shingle, and `filler_token` (default `_`) is inserted where a previous filter removed a token. That second one is the surprise: run `stop` before `shingle` and `the cat sat` yields shingles containing `_`, because the removed stopword still occupies a position. Whether that is desirable depends on what you are matching — it does preserve the fact that *something* was there, but it means the shingles no longer read as natural word pairs. ## The cost Term count is the whole story. A corpus has far more distinct word pairs than distinct words, and the distribution is much flatter, so a shingled field enlarges the term dictionary disproportionately rather than just adding postings to existing terms. Consequences: bigger index on disk, more memory for the terms index, slower merges, and longer shard recovery. Trigrams make all of that dramatically worse. There is also a relevance cost if you are careless. Scoring a field that mixes unigrams and shingles conflates two different signals, and boosting a shingled field too heavily makes exact word-order matches dominate results that users expected to be more forgiving. ## How to deploy it sensibly Put shingles on their own subfield with their own analyzer, keep `min_shingle_size` and `max_shingle_size` at 2 unless you have measured a reason to go further, set `output_unigrams: false`, and decide deliberately whether stopwords are removed before shingling. Then measure: index size before and after, and relevance before and after on a real query set. Shingles are one of the few relevance levers where the storage cost is large enough to be a capacity decision, not just a mapping choice. ## When not to use them If you simply need exact phrase matching, a phrase query on a normally-analyzed field already does that correctly and cheaply enough for most workloads. If you need substring matching inside words, that is character n-grams, a different filter entirely. Shingles earn their place when word adjacency needs to be a *scoring* signal blended with other signals, or when you need bigram counts as data.
- Why set output_unigrams to false on a dedicated shingled subfield?The unigrams already exist on the parent field, so emitting them again doubles the subfield's terms and mixes two signals in one score. With it off, the subfield scores purely on word adjacency, which is the signal you are boosting. Watch single-word inputs, where output_unigrams_if_no_shingles may be needed so they still produce a token.
- What appears in your shingles if the stop filter runs before the shingle filter?A filler token — an underscore by default — occupies the position of each removed word, so "the cat sat" can yield shingles such as `_ cat` and `cat sat`. It preserves the fact that a token was there, but the shingles stop reading as natural word pairs. Decide the ordering deliberately and inspect the output.
- When is a phrase query the better tool than a shingled field?When you need exact phrase semantics rather than a ranking signal. A phrase query reads positions and answers the adjacency question precisely, with no extra storage. Shingles are worth their index cost when adjacency must be blended into scoring alongside other fields and boosts, or when you need bigram term statistics as data.
It is like pre-computing every pair of neighbouring words in a book's index so you can look up a two-word phrase as fast as a single word — accurate, but the index gets much thicker than the book.
saying these in an interview costs you the question
- Confuses shingles with character n-grams
- Adds shingles to the main text field
- Leaves output_unigrams on for a shingled subfield
- Ignores the term-dictionary growth from bigrams
- Uses shingles where a phrase query would do