Should synonym expansion happen at index time or query time in a search system?
answer
- One side is data, the other is configuration
- Changing the list costs very differently in each case
- Injected tokens change document frequencies
- Two tokens where there was one is the hard case
- Rule direction is not a detail
basics
~20 sIndex-time expansion makes queries simple and fast but requires reindexing to change and distorts term statistics. Query-time expansion is editable immediately and keeps statistics honest, at the cost of larger queries and hard multi-word cases. Most teams expand at query time.
solid answer
~50 sBoth approaches inject alternative tokens so *tv* can find *television*. **Index-time** expansion writes the synonyms into the postings at the original token's position: queries stay small and fast, phrase matching works naturally, and the cost is paid once. But the list is frozen into the index — every edit means a reindex — and injected tokens inflate document frequencies, so the term statistics that drive ranking no longer reflect the corpus. **Query-time** expansion rewrites the query instead: edits take effect on a reload, the index stays clean, and you can experiment cheaply. Its costs are a wider query to execute and the genuinely hard problem of **multi-word** synonyms, where one token must expand to a phrase and the query parser has to produce a token graph rather than a flat list. There is also an IDF asymmetry: a rare synonym carries a huge weight and can dominate scoring. The common default is query-time, with a small stable set handled at index time when latency demands it.
code
text · 9 lines// index-time expansion: tokens written into the postings
doc "cheap tv deals" -> [cheap] [tv|television]@2 [deals]
-> editing the rule list requires reindexing
-> df(television) now counts every doc that said only "tv"
// query-time expansion: query rewritten instead
query "tv" -> (tv OR television)
-> rule edits take effect on reload; index untouched
-> multi-word case needs a graph: "ny" -> (ny) OR (new AND york, adjacent)go deeper
Know that synonyms let a search for one word find documents using another, and that the expansion can be applied either when indexing documents or when parsing the query.
Compare the two placements concretely: reindex cost versus per-query cost, and what each does to the terms stored in the index. Know that injected tokens sit at the original token's position.
Bring the failure modes: statistics distortion, the multi-word token-graph problem, IDF asymmetry between a common term and a rare synonym, and the filter's ordering relative to stemming and folding.
Treat the synonym list as a governed relevance asset: who may edit it, how a rule set is evaluated before and after, how directional rules encode the taxonomy, and how you stop years of accumulated rules from quietly degrading ranking.
## What synonym expansion does A synonym filter maps tokens to alternatives so that vocabulary mismatch between authors and searchers does not cost recall. Users type *tv*, documents say *television*; users type *laptop*, documents say *notebook*. The filter injects the alternative into the token stream, typically at the **same position** as the original, so both are candidates at that spot and phrase/proximity logic still lines up. Rules come in two shapes. **Two-way (equivalent)** rules make a set of terms interchangeable: any of them matches any other. **One-way (replacement)** rules expand in a single direction — searching a broader term finds the narrower ones, but not the reverse. Direction matters enormously in practice: making *shoe* and *sneaker* equivalent means someone searching for formal shoes gets sneakers, whereas a one-way rule from the broad to the narrow term expresses the taxonomy correctly. ## Index-time expansion The synonyms are applied during document analysis, so the injected tokens become real terms in the postings. *Advantages.* Queries remain single-token and fast; there is no per-query expansion work. Phrase and proximity behaviour is well defined because the injected tokens occupy real positions. Multi-word synonyms are much easier here, since you control the whole document token stream at write time. *Disadvantages.* The list is baked in. Any addition, removal or correction requires reindexing the corpus, which for a large index is a multi-hour operation and a deployment event rather than a config change — a serious problem for a list that merchandising or support teams want to edit weekly. The index grows with the injected tokens. Most importantly, **term statistics are corrupted**: injecting *television* into every document that says *tv* raises *television*'s document frequency above its true corpus value, lowering its IDF and changing how every query containing it is ranked, including queries that never touch the synonym rule. ## Query-time expansion The synonyms are applied while parsing the query, rewriting one token into a set of alternatives. *Advantages.* The list is configuration, not data: change it and the next query sees it, so iteration and A/B testing are cheap. The index stays a faithful record of the corpus, so IDF stays honest. Removing a bad rule is instant rather than an incident. *Disadvantages.* Every affected query becomes wider — more terms, more postings lists, more work — which matters at high query volume with large expansion sets. And the two hard problems: **Multi-word synonyms.** Mapping *ny* to *new york* means a single query token expands into a two-token sequence. A flat token list cannot represent "either this one token or those two tokens at this position", so the query parser must build a **token graph** and generate the correct alternatives. This has historically been a notorious source of wrong results, especially when combined with phrase matching or a minimum-match requirement, and it interacts badly with any filter that changes token counts. **IDF asymmetry.** If the query expands *tv* to include the much rarer *telly*, the rare term carries a far higher IDF, so a document matching only *telly* can outrank a squarely relevant document matching *tv*. Some engines let a synonym group be scored using blended statistics precisely to avoid this; where they do not, expansion sets should be kept to terms of broadly similar frequency, or the expansions down-weighted. ## Ordering inside the analysis chain Synonyms interact with the rest of the chain, and the ordering bug is extremely common: if the stemmer runs before the synonym filter, your rules must be written in **stemmed** form or they will never fire, because the token arriving at the filter is `runn`-shaped, not `running`. If synonyms run first, the injected tokens are then stemmed along with everything else, which is usually what you want. Similarly, case and accent folding must happen before matching a rule list written in lowercase. Whatever the order, the synonym list is part of the analysis contract and must be tested as such. ## Choosing The practical default is **query-time** expansion, because editability usually dominates and modern query parsing handles the common cases. Move a rule to index time when it is stable, when the query-side cost is measurable at your traffic level, or when it is a multi-word mapping the query path handles badly. Many systems do both: a small, stable, curated set at index time, plus a larger, evolving set at query time. Whatever you pick, treat synonyms as a relevance intervention, not a text-processing detail. They change recall and ranking together, they are easy to add and hard to evaluate one at a time, and a synonym list left to accumulate for two years is one of the most common causes of mysteriously bad relevance. Measure the effect of rule sets, keep them reviewable, and prefer directional rules over blanket equivalence.
- Why are multi-word synonyms the hard case for query-time expansion?One query token has to expand into a sequence of tokens, so a flat list cannot express "either this single token or these two, at this position". The parser must build a token graph and enumerate the alternative paths. Combined with phrase matching or a minimum-terms-must-match rule, naive handling produces wrong or empty results, which is why the multi-word case is the standard interview probe.
- How does index-time expansion distort relevance beyond the queries it targets?Injecting a synonym token into every document containing the original raises that term's document frequency above its true corpus value. Its IDF falls, so every query containing that word — including ones unrelated to the rule — ranks documents differently. The index stops being a faithful description of the corpus, and the effect is invisible until someone compares statistics.
- Where should the synonym filter sit relative to the stemmer?Usually before it, so injected tokens are stemmed along with everything else and the rule list can be written in ordinary surface forms. If the stemmer runs first, every entry in the list must be spelled as the stemmer's output or the rules silently never fire — a common and hard-to-spot configuration bug.
- When is a one-way rule clearly better than a two-way one?Whenever the terms stand in a broader/narrower relationship. Making shoe and sneaker equivalent means a search for formal shoes returns sneakers; expanding only from the broad term to the narrow one preserves the taxonomy. Two-way rules belong to genuine equivalents such as abbreviations, spelling variants and brand aliases.
Index-time synonyms are like writing every translation into the book itself — fast to read, painful to correct. Query-time synonyms are a phrasebook you carry: instantly editable, but you consult it on every question.
saying these in an interview costs you the question
- Treats synonym expansion as free extra recall with no ranking cost
- Thinks the index-time list can be edited without reindexing
- Ignores that injected tokens change document frequencies
- Makes every rule bidirectional regardless of meaning
- Writes rules in surface forms while stemming runs first