skip to content

Should synonym expansion in Elasticsearch run at index time or at search time, and why?

level: seniorimportance: should knowfreq 50%

answer

  1. One side is baked into segments, one is not
  2. Think about how often the rule list changes
  3. Multi-word rules need a graph, not extra tokens
  4. Extra indexed terms shift document frequencies
  5. A reload API exists for one placement only

basics

~20 s

Search time is the usual choice: rules can change without reindexing, the index stays clean, and synonym_graph handles multi-word synonyms correctly. Index-time expansion avoids per-query cost but bakes rules into segments and distorts term statistics.

solid answer

~50 s

Expanding at **index time** writes the synonyms into the postings, so queries stay cheap and no synonym filter runs per request. The costs are severe for a system whose vocabulary evolves: every rule change requires a full reindex, the extra terms inflate document frequencies and therefore skew BM25 scoring, and multi-word synonyms distort positions in ways that break phrase matching — which is why `synonym_graph` is supported only in a search analyzer. Expanding at **search time** keeps the index faithful to the source text, lets `synonym_graph` build a proper token graph for multi-word rules, and makes rule changes immediate; with `updateable: true` on the filter you can refresh rules on a live index via `_reload_search_analyzers` instead of reindexing. The price is a little work per query and query expansion into more term lookups. Default to search time; consider index time only for a frozen vocabulary where query latency is critical.

code

json · 28 lines
json
{
  "settings": {
    "analysis": {
      "filter": {
        "my_synonyms": {
          "type": "synonym_graph",
          "synonyms_path": "analysis/synonyms.txt",
          "updateable": true
        }
      },
      "analyzer": {
        "text_search": {
          "tokenizer": "standard",
          "filter": ["lowercase", "my_synonyms"]
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "body": {
        "type": "text",
        "analyzer": "standard",
        "search_analyzer": "text_search"
      }
    }
  }
}

go deeper

for a junior

Know that synonyms are configured as a token filter in an analyzer, and that where you put that filter decides whether the extra words end up in the index or in the query.

for a middle

Explain both placements with a concrete rule and state the mechanical consequences: reindex to change index-time rules, per-query expansion for search-time ones.

for a senior

Argue the default from operations: rule churn, honest term statistics, graph handling of multi-word synonyms, and the reload path that avoids a reindex.

for a principal

Own the vocabulary as a product asset — who edits rules, how changes are validated against judgements before rollout, and what the rollback story is.

## The same result set, very different operations Synonym expansion can be placed on either side of the analysis boundary, and for simple one-word-to-one-word rules both placements can produce the same hits. Doing it at index time means the document `"laptop"` also stores the term `notebook`, so a plain query for `notebook` finds it. Doing it at search time means the document stores only `laptop`, and the query `notebook` is expanded to `notebook OR laptop` before lookup. The retrieval outcome looks alike; the engineering consequences do not. ## Index time: cheap queries, expensive change What index-time expansion buys is per-query simplicity — no filter runs on the query, and the query stays a single term lookup. What it costs: - **Reindexing on every rule change.** Synonyms are exactly the kind of configuration a search team iterates on weekly. Baking them into segments means each iteration is a full reindex of the corpus. - **Distorted term statistics.** BM25 scores a term partly by its inverse document frequency. Injecting synonyms raises the document frequency of the injected terms and changes the field length used for length normalization, so relevance shifts in ways unrelated to the actual content. - **Broken provenance.** Highlighting and debugging show terms that never appeared in the source text, and there is no way to tell an expanded term from a real one. - **Multi-word trouble.** Rules like `ny, new york` must map one token to two, or two to one. At index time this forces awkward position handling, and the graph-aware `synonym_graph` filter is not available for index analyzers at all — it is documented as a search-analyzer filter, precisely because a graph token stream is something a query parser can consume but an index cannot store faithfully. ## Search time: the default choice Search-time expansion keeps the inverted index a faithful representation of the source text. That single property gives you most of the advantages: term statistics stay honest, highlighting shows real terms, and reindexing is decoupled from relevance work. It also unlocks the correct machinery for multi-word rules. `synonym_graph` emits a token graph — alternative paths through the query, so `"ny pizza"` can be understood as either `ny pizza` or `new york pizza` — and full-text queries know how to turn that graph into the right boolean or phrase structure. That is far more accurate than anything achievable by stuffing extra tokens into the index. Operationally, a search-time synonym filter can be declared `updateable: true` (allowed only when the analyzer is used purely as a search analyzer). The rules then live in a file on every node, and `POST /<index>/_reload_search_analyzers` picks up an edited file on a live index — a rule change goes out in seconds without touching data. Recent 8.x versions additionally offer a managed synonyms API for storing synonym sets in the cluster rather than in files on disk; either way the point stands that the search side is the changeable side. The cost side is real but modest: each query runs the filter and expands into more term lookups, so a rule set that expands one word into ten multiplies the clause count. Very large rule sets, or rules that chain into each other, can make queries measurably slower and can also flatten relevance because a rare word and its common synonym have very different IDFs — a document matching only the common synonym may score surprisingly high or low. ## Choosing Default to search time. Choose index time only when all of the following hold: the vocabulary is genuinely frozen (standardized product codes, unit abbreviations), query latency is at a premium, and the rules are one-word-to-one-word so the position problem never arises. Even then, document the decision loudly, because the next person who wants to change a synonym will otherwise discover the reindex requirement the hard way. A hybrid also exists and is sometimes the pragmatic answer: normalize a small, stable set of equivalences at index time (unit spellings, brand casing) and keep the evolving, business-driven vocabulary in the search analyzer. ## What interviewers listen for The strong answer is not "search time" on its own — it is the pairing of reasons: mutability without reindex, honest term statistics, and graph handling of multi-word rules, against the honest admission that you pay for expansion on every query and that IDF differences between a term and its synonym still make scoring imperfect on either side.

  • Why is synonym_graph restricted to search analyzers?
    It emits a token graph — alternative paths of differing token counts, needed for rules like `ny, new york`. A query parser can turn that graph into the correct boolean or phrase structure, but the inverted index stores a flat sequence of positions, so the alternatives cannot be persisted faithfully. Index-time multi-word synonyms therefore distort positions and break phrase matching.
  • How do you change synonym rules on a live index without reindexing?
    Use a search-time synonym filter declared with `updateable: true` (permitted only when the analyzer serves as a search analyzer), keep the rules in a file present on every node, edit it, then call `POST /<index>/_reload_search_analyzers`. New searches use the new rules immediately; stored data is untouched. Recent 8.x versions also offer a managed synonyms API as an alternative to files.
  • Does search-time expansion give perfect scoring?
    No. The expanded terms have their own document frequencies, so a rare word and a common synonym contribute very differently to BM25, and a document matching only the common synonym can score unexpectedly. Elasticsearch mitigates some of this when synonyms occupy the same position, but relevance still needs checking against a judgement set after any rule change.
  • When is index-time expansion still defensible?
    When the vocabulary is genuinely frozen — unit abbreviations, standardized codes — the rules are one-word-to-one-word so positions are unaffected, and per-query latency matters more than iteration speed. Even then it should be documented, because changing a single rule later means reindexing the whole corpus.

saying these in an interview costs you the question

  • Says index-time synonyms are better because they are faster, ignoring reindex cost
  • Uses synonym_graph in an index analyzer for multi-word rules
  • Claims synonyms have no effect on scoring
  • Thinks editing the synonym file alone updates a running index
  • Assumes search-time expansion is free at query time

context