skip to content

What steps turn a raw text field into indexed terms in a search engine's analysis pipeline?

level: juniorimportance: must knowfreq 68%

answer

  1. Three stages, in a fixed order
  2. Something happens before tokens even exist
  3. Lookup is exact string equality on terms
  4. Character filter, tokenizer, token filters
  5. The last stage is a chain, and order matters

basics

~20 s

Text analysis runs in three stages: character filtering (strip markup, map characters), tokenization (cut the stream into tokens), then token filtering (lowercase, fold accents, drop stopwords, stem, add synonyms). The surviving tokens become the index terms.

solid answer

~40 s

Analysis converts a string into the list of **terms** that actually go into the inverted index, and it has three stages. First, **character filters** rewrite the raw character stream — stripping HTML tags, mapping `&` to `and`, normalising Unicode forms. Second, a **tokenizer** cuts that stream into tokens; for space-delimited languages this is usually word-boundary splitting that also discards punctuation, and it records each token's position and offsets. Third, a chain of **token filters** transforms the token stream: lowercasing, accent/diacritic folding, stopword removal, stemming, synonym injection, n-gram generation. Whatever comes out the end is stored verbatim as terms. That matters because lookup in an inverted index is *exact string equality* on terms — all the fuzziness of "full-text search" is manufactured here, in analysis, not at match time.

code

text · 7 lines
text
input:   "The Quick-Brown Foxes!"
char filter (strip markup, map chars) -> "The Quick-Brown Foxes!"
tokenizer (word boundaries)           -> [The] [Quick] [Brown] [Foxes]
lowercase                             -> [the] [quick] [brown] [foxes]
stopword removal (english)            -> [quick] [brown] [foxes]
stemmer (english)                     -> [quick] [brown] [fox]
indexed terms: quick, brown, fox

go deeper

for a junior

Be ready to name the three stages in order and walk a short sentence through them, saying which terms end up in the index. Knowing that lookup is exact on those terms is the point.

for a middle

Explain what each stage can and cannot do, why something like markup stripping must happen before tokenization, and how filter ordering changes the resulting terms. Mention positions and offsets and what they are used for.

for a senior

Show judgment about which fields deserve analysis at all, and about tokenization choices that leak to users — hyphens, compounds, unspaced scripts. Be ready to debug a matching failure by reasoning about the emitted term stream.

for a principal

Own analysis as a schema-wide contract: how many analyzed variants a field really needs, what each costs in index size and reindex risk, and how you keep pipeline changes from silently breaking existing queries across teams.

## Why analysis exists An inverted index maps a **term** to the list of documents containing it. Lookup in that structure is an exact match on a byte string: the engine hashes or binary-searches the term dictionary for the term you hand it. Nothing about that structure is fuzzy. So every bit of tolerance users expect from search — case-insensitivity, matching *running* against *run*, ignoring the accent in *café* — has to be manufactured **before** anything is stored, by normalising both the document text and the query text into the same reduced vocabulary. That normalisation process is text analysis, and it runs in three ordered stages. ## Stage 1: character filtering The first stage sees the raw character stream and rewrites it before any token boundaries exist. Typical jobs: stripping HTML or XML markup so `<b>shoe</b>` does not index `b` as a word; character mapping so `&` becomes `and`, or so curly quotes become straight ones; Unicode normalisation, which matters because the same visible character can have several encodings (an *é* can be one code point or *e* plus a combining accent, and those are different byte strings — different terms). Anything that must happen before tokenization belongs here, because once the text is cut into tokens you can no longer merge or split across a boundary. ## Stage 2: tokenization The tokenizer cuts the character stream into tokens. For English and most European languages this is roughly word-boundary splitting: whitespace and punctuation separate tokens and are themselves discarded. Real tokenizers are more careful than `split(" ")` — they have rules for hyphens, apostrophes, URLs, email addresses and numbers, because those choices are visible to users. Whether `wi-fi` becomes one token or two decides whether a search for `wifi` can ever match. Tokenization is also where languages differ most sharply. Chinese, Japanese and Thai are written without spaces, so a tokenizer must either segment using a dictionary and statistical model, or fall back to overlapping character bigrams. German forms long compounds (*Fußballmannschaft*), so a decompounder may be needed for a search on *Mannschaft* to hit it. There is no language-neutral tokenizer, which is why engines ship per-language ones. The tokenizer also attaches metadata to each token that later stages and later queries depend on: a **position** (used by phrase and proximity matching) and character **offsets** into the original text (used by highlighting). Losing or shifting these is a real bug class. ## Stage 3: token filtering Token filters form a chain, each taking a token stream and emitting one. Order matters, because each filter sees only what the previous one produced. The common filters: - **Lowercasing** — so `Shoe` and `shoe` collapse to one term. - **Accent/diacritic folding** — `café` to `cafe`, so users typing without accents still match. - **Stopword removal** — dropping very common function words such as *the*, *of*, *and*. - **Stemming** — reducing inflected forms to a common stem, so *running*, *runs* and *ran* can meet. - **Synonym expansion** — injecting extra tokens at the same position, so *tv* also indexes or queries as *television*. - **N-gram generation** — emitting substrings or prefixes of a token to support partial matching. - **Length limits and trimming** — discarding absurdly long tokens produced by junk data. Ordering bugs are common: applying a synonym list *after* stemming means your synonym entries must be written in stemmed form, or they never fire. ## What actually lands in the index A useful mental exercise is to run a sentence through the pipeline by hand. `The Quick-Brown Foxes!` through a standard chain — tokenize, lowercase, remove English stopwords, stem — indexes roughly `quick`, `brown`, `fox`. The literal string is gone; only those three terms are searchable in that field. This is why a query that looks obviously correct can return nothing: it is being compared to `fox`, not to `Foxes`. ## Full-text fields versus exact-match fields Not every field wants this treatment. Identifiers, status codes, tags and anything used for exact filtering, sorting or grouping should be indexed **unanalyzed** — as a single term equal to the whole input value — because you want `IN_PROGRESS` to be one term, not `in` plus `progress`. Most engines therefore distinguish an analyzed full-text field type from a keyword-style exact field type, and let you index one source value both ways under different field names, choosing at query time which behaviour you want. Getting this choice wrong is the second-most common cause of "my filter matches nothing" after analyzer asymmetry. ## The takeaway Analysis is a lossy, deliberate transformation. Everything the index knows about a document's text is what the pipeline emitted, so when debugging matching problems the first move is always to ask what terms this text actually produced.

  • Why do search engines record a position and character offsets for each token?
    Positions let the engine answer phrase and proximity queries: it can check that *new* is immediately followed by *york* rather than just co-occurring in the document. Character offsets point back into the original text, which is how highlighters mark the matched span in the stored source even though the indexed term no longer resembles the original word.
  • When would you index a text field without running any analysis on it?
    For values that are identifiers rather than prose: status codes, SKUs, tags, usernames, enum values, keys used for exact filtering, sorting or grouping. There you want the entire value to be one term so equality is meaningful and grouping buckets are the real values, not word fragments.
  • Does changing the order of token filters change the index?
    Yes, materially. Filters chain, so each sees only the previous one's output. Stemming before synonym expansion means synonym entries must be written as stems or they never fire; folding accents after a language stemmer can defeat rules that depend on accented characters. Order is part of the configuration you must test, not a formality.

Analysis is like grinding beans before brewing: the index can only work with what came out of the grinder, and if the document and the query were ground differently, nothing lines up.

saying these in an interview costs you the question

  • Thinks the engine stores original text and matches it fuzzily
  • Believes tokenization is always just splitting on whitespace
  • Cannot name any stage before tokenization
  • Assumes token filter order does not matter
  • Runs full analysis on identifier and status-code fields

context