Why must index-time and query-time text analysis produce compatible terms in a search engine?
answer
- Matching is exact, not fuzzy
- The failure mode is silence, not an error
- Same-analyzer is the default, not the actual rule
- Some asymmetry is deliberate — autocomplete, synonyms
- Index-time changes need a reindex
basics
~20 sBoth sides must reduce text to the same term forms. The index stores analyzed terms, and a query can only match a term spelled identically after its own analysis, so mismatched pipelines silently return zero hits instead of an error.
solid answer
~50 sMatching in an inverted index is exact string equality on terms. The document side stores whatever the index-time pipeline emitted; the query side looks up whatever the query-time pipeline emitted. If the document was lowercased and stemmed to `run shoe` but the query is looked up as the literal `Running Shoes`, the two never meet and you get zero results with **no error** — which is why this is the single most common "my search is broken" bug. The rule is that the two sides must produce *compatible* terms, not that they must run identical chains. Deliberate asymmetry is normal and useful: index prefixes for autocomplete but analyze the query plainly; expand synonyms only at query time so the list stays editable. What is never acceptable is accidental asymmetry, and the fix when it exists is to inspect the token stream both sides actually produce.
code
text · 8 lines// document side (index-time: lowercase + english stemmer)
"Running Shoes" -> [run] [shoe]
// query side A (unanalyzed exact term lookup)
"Running Shoes" -> [Running Shoes] => no dictionary entry => 0 hits
// query side B (same analysis as index time)
"Running Shoes" -> [run] [shoe] => both terms found => matchgo deeper
Recall that documents are stored as analyzed terms and a query must produce the same terms to match. Know that the symptom is zero results with no error message.
Explain the mechanic — exact term lookup — and walk through a concrete mismatch such as a lowercased, stemmed field queried with an unanalyzed exact term. Know that index-time changes require reindexing.
Demonstrate the diagnostic routine: inspect emitted terms on both sides before touching configuration. Be able to justify deliberate asymmetry and to plan a safe reanalysis rollout for a populated index.
Own analysis as versioned schema across teams: what is safe to tune at query time, what forces a reindex and its cost, and the test and rollout discipline that stops a one-line analyzer edit from silently emptying result sets.
## The underlying mechanic An inverted index is a dictionary of terms, each pointing at a postings list of documents. A query term is resolved by looking it up in that dictionary — a byte-level exact match. There is no similarity, no case-insensitive comparison, no built-in tolerance at lookup time. All of that lives in analysis. Consequently the invariant that makes full-text search work is: **the transformation applied to document text and the transformation applied to query text must land on the same term forms.** This is often stated as "use the same analyzer on both sides", which is a good default but slightly too strong. The real requirement is compatibility of outputs, not identity of configuration. ## How it breaks The classic failure is mixing a full-text field with an exact-term lookup. Suppose a product name is indexed through a chain that lowercases and stems: `Running Shoes` becomes the terms `run` and `shoe`. Now a developer writes a filter that compares the field to the exact string `Running Shoes`. That lookup asks the dictionary for the single term `Running Shoes`, which does not exist. Result: zero hits, no error, no warning. The same shape appears with `Active` versus `active`, with accented text, and with any query type that deliberately skips analysis in order to be exact. The second common failure is configuration drift. Someone changes the analysis chain — adds a stemmer, swaps the stopword list, turns on accent folding — and applies it only to the search side, or applies it to the index definition without reindexing the existing documents. Older documents keep their old terms; new documents get new ones; queries match one cohort and not the other. Symptoms are partial, data-dependent weirdness rather than an outright zero. A third is the *language* mismatch: index with a German analyzer, query with an English one. Both sides analyze, both look plausible, and stems diverge. ## Symptoms and diagnosis The tell-tale signature is a query returning **exactly zero** results for text you can see in the document, while a substring or a simpler query works. The diagnostic procedure is always the same and interviewers want to hear it: 1. Ask the engine what terms this document text produced under the field's index-time analysis. 2. Ask it what terms this query string produces under the query-time analysis. 3. Compare the two lists. The mismatch is visible immediately — different case, an extra stem, a stopword dropped on one side, one side producing a single term and the other producing three. Engines expose an inspection endpoint or API for exactly this, and using it beats guessing. A secondary check is asking the engine to explain the scoring for a document you expect to match; if the explanation shows no matching term, the analysis is the culprit rather than the ranking. ## Legitimate asymmetry Deliberate asymmetry is a real technique, and knowing which cases are intentional separates a middle candidate from a junior one: - **Prefix autocomplete.** Generate edge n-grams at index time (`shoe` becomes `s`, `sh`, `sho`, `shoe`) but analyze the query as a plain token. If you n-grammed the query too, the query `shoes` would emit `s`, `sh`, ... and match nearly everything, destroying precision. - **Synonyms.** Expanding at query time only keeps the synonym list editable without a reindex, and keeps injected tokens out of the stored postings. - **Phrase-sensitive fields.** Some setups keep stopwords on the query side for exact-phrase behaviour while the index side is more aggressive, or vice versa. In every legitimate case the asymmetry is *chosen* and the outputs still intersect on purpose. Accidental asymmetry produces empty intersections. ## Changing analysis after data exists Index-time analysis is baked into the postings the moment a document is indexed. Changing the chain does not retroactively re-analyze anything, so an index-time change requires **reindexing** every affected document to take effect consistently — usually by writing into a new index and swapping traffic over. Query-time-only components (a query-time synonym list, for example) take effect without reindexing, which is precisely why teams push volatile configuration to the query side. This asymmetry in *cost* drives real design decisions: things you expect to tune weekly belong at query time even if index time would be faster to execute; things that are stable and expensive belong at index time. ## Practical guardrails Treat analysis configuration as schema: version it, review it, and test it. A small suite of assertions — "this input produces these terms" plus a handful of end-to-end "this query must find this document" cases — catches drift before users do. And when the two sides must differ, write down *why*, because the next person's instinct will be to make them identical.
- Name a case where you deliberately want different analysis on the two sides.Prefix autocomplete. Index edge n-grams so every prefix of every token is a term, but analyze the user's input as a plain token. If you also n-grammed the query, a query for "shoes" would emit "s", "sh", "sho" and match nearly the whole corpus. Query-time-only synonym expansion is the other standard case, chosen so the list stays editable without reindexing.
- You change a field's index-time analyzer on a populated index. What happens to documents already indexed?Nothing — their terms were written at index time and are frozen. Old documents keep old terms, new writes get new terms, and queries match one cohort inconsistently. You must reindex the affected documents, typically into a fresh index that is then swapped in behind an alias, so the whole corpus shares one term vocabulary.
- How do you prove analyzer asymmetry is the cause rather than a ranking problem?Ask the engine to analyze both the document text and the query string under their respective configurations and compare the emitted term lists. If they do not intersect, it is analysis. Ranking problems look different: the document matches but sits too low, and a score explanation shows matched terms with small contributions.
It is a lock and key cut by two different machines: the door does not tell you the key is the wrong shape, it just never opens.
saying these in an interview costs you the question
- Thinks the engine matches the original text, so case cannot matter
- Says analyzers must always be identical on both sides
- Expects an error rather than zero results on mismatch
- Believes changing an analyzer re-analyzes existing documents
- Debugs by adding boosts instead of inspecting emitted terms