skip to content

Why does synonym_graph handle multi-word synonyms correctly where the older synonym filter does not?

level: seniorimportance: should knowfreq 40%

answer

  1. The problem is positions, not words
  2. One token becoming two needs a branch
  3. Flat streams cannot express spans
  4. Indexing requires collapsing the structure
  5. Rules are analyzed by preceding filters

basics

~10 s

synonym_graph emits a graph token stream, so a multi-word replacement occupies the right span of positions. The plain synonym filter stacks the replacement at one position, corrupting positions and breaking phrase and proximity matching.

solid answer

~50 s

A rule like `ny, new york` replaces one token with two, or two with one. The plain `synonym` filter stacks alternatives at a single position, so `new york` injected for `ny` ends up with wrong position increments — phrase queries and proximity matching then behave unpredictably. `synonym_graph` instead produces a graph token stream in which a multi-token alternative spans the correct number of positions, and Elasticsearch's query parser can traverse that graph to build the right phrase alternatives. Because a graph stream cannot be written directly into the index, `synonym_graph` is designed for search-time analysis; putting it in an index-time chain requires a following `flatten_graph`, which collapses the graph and loses some of the multi-word fidelity. Rules come from an inline `synonyms` list or `synonyms_path`, and are themselves analyzed by the filters preceding the synonym filter in the chain.

code

json · 14 lines
json
"analysis": {
  "filter": {
    "city_syn": {
      "type": "synonym_graph",
      "synonyms": ["ny, new york", "sf, san francisco"]
    }
  },
  "analyzer": {
    "city_search": {
      "tokenizer": "standard",
      "filter": ["lowercase", "city_syn"]
    }
  }
}

go deeper

for a junior

Know that a synonym filter rewrites tokens according to rules, and that multi-word synonyms are the case that needs the graph-aware variant.

for a middle

Explain positions: two tokens at one position mean alternatives, and a replacement spanning several positions cannot be expressed in a flat stream without corrupting phrase matching.

for a senior

Diagnose the production symptom — phrase queries matching inconsistently after synonyms were added — and reason about flatten_graph, rule parsing through the preceding chain, and the cost of large rule sets.

for a principal

Own the synonym programme itself: who maintains the rules, how changes are validated and rolled out, and whether curated synonyms remain the right lever versus other relevance investments.

## What a synonym filter actually does A synonym filter rewrites the token stream according to rules. In the default Solr-style format, `ny, new york` declares those forms equivalent, and `laptop => notebook, portable` maps explicitly in one direction. Rules can be inline under `synonyms`, or loaded from a file with `synonyms_path` relative to the node's config directory; recent 8.x versions also expose a synonyms API and a `synonyms_set` parameter so rule sets can be managed and reloaded without touching files on every node. The hard part is not the rules. It is positions. ## Positions and why they matter Every token carries a position. Phrase queries, `match_phrase` with a slop, and proximity scoring all work by comparing positions: `new` at position 3 and `york` at position 4 make an adjacent phrase. Position increments also encode alternatives — two tokens at the same position mean "either of these matches here". Single-word synonyms fit that model perfectly: replace `car` with `car` and `automobile` at the same position and everything downstream works. Multi-word synonyms do not. If the rule turns one token into two, the stream needs a branch that is *two positions long* running in parallel with a branch that is one position long. A flat token stream with position increments cannot express that faithfully. ## The old filter's failure mode The original `synonym` filter emits the replacement tokens with position increments that squeeze the multi-token alternative into the space of the original. The result is a stream where `york` may sit at the same position as the following real word, or where the injected phrase overlaps its neighbours. Symptoms in production: a phrase query for the synonym form matches documents it should not, or fails to match documents it should; slop values behave inconsistently; highlighting picks odd spans. The rules look right, the analysis output looks roughly right, and the search results are subtly wrong — which is why this is a senior-level question. ## The graph model `synonym_graph` emits a *graph* token stream: tokens carry both a position and a position length, so an alternative that spans two input positions is recorded as such. The query parser reads that graph and enumerates the valid paths through it, generating the correct set of phrase alternatives — effectively "either the single token `ny` here, or the two-token sequence `new york` covering the same span". Phrase and proximity semantics are then preserved for both forms. Other graph-producing filters exist for the same reason, notably `word_delimiter_graph`, which is the graph-aware replacement for the older `word_delimiter` filter. ## flatten_graph and the index-time constraint The index format stores positions, not graphs. So a graph stream cannot be indexed as-is. `synonym_graph` is therefore designed for the analyzer used at query time. If you do want the expansion baked into the index, the chain must end with `flatten_graph`, which collapses the graph into a flat stream — and in collapsing it, reintroduces exactly the multi-word imprecision the graph was designed to avoid. That trade-off, not the syntax, is the substance of the question. ## Rule parsing depends on the chain Synonym rules are text, and Elasticsearch analyzes them using the tokenizer and the filters that appear *before* the synonym filter. Two consequences follow. First, if a stemmer runs before the synonym filter, your rules are stemmed too, and a rule written as `running shoes` becomes something else — which may be what you want or may silently stop it matching. Second, some preceding filters are not usable for rule parsing, and Elasticsearch will either ignore them or reject the configuration depending on version; a chain that fails to create an index with a synonym filter in it is very often this. ## Operational notes `expand` (default `true`) controls how equivalence lists behave: with `expand: false`, `foo, bar, baz` is treated as though every form maps to the first. `lenient` lets the filter skip rules it cannot parse rather than failing outright — convenient for large hand-maintained files, dangerous because a typo then disappears silently. `format` accepts the default Solr syntax or `wordnet`. Large synonym sets are a real operational cost: rules are held per index and parsed at analyzer construction, so tens of thousands of rules affect memory and shard open time. Rules that expand aggressively also inflate query complexity, since every alternative path multiplies the clauses the query generates. ## The answer in one breath Multi-word synonyms need a branch of the token stream that is longer than one position; `synonym_graph` records that with position length and the query parser walks it, while the flat `synonym` filter cannot represent it and corrupts positions instead. Graph output is query-time by nature; indexing it requires `flatten_graph` and gives up some correctness.

  • What does setting expand to false do to a rule written as "ipod, i-pod, i pod"?
    It treats the equivalence list as mapping every form to the first one, so all three inputs produce `ipod` rather than each producing all three alternatives. That keeps the token stream small and the term dictionary tidy, at the cost of only working when the canonical form is the one you also index.
  • Why can a stemmer placed before a synonym filter stop your rules from matching?
    Synonym rules are parsed through the tokenizer and the filters preceding the synonym filter, so the rules get stemmed too. A rule written in natural word forms may end up stored in stemmed form, matching different inputs than you intended. Write rules in the form the preceding chain produces, or place the synonym filter before the stemmer.
  • What is the risk of enabling lenient on a large synonym file?
    Rules that fail to parse are skipped instead of failing index creation, so a typo silently removes a rule and nobody notices until relevance complaints arrive. It is useful for hand-maintained sets that must not block deployment, but pair it with validation of the file before it ships rather than relying on the engine to complain.

saying these in an interview costs you the question

  • Says the two synonym filters differ only in performance
  • Claims graph output can be written straight into the index
  • Ignores that flatten_graph loses multi-word fidelity
  • Thinks synonym rules are matched as raw text
  • Assumes huge synonym files are free to load

context