skip to content

Text Analysis & Tokenization

Raw text has to be cut into terms and normalized before it can be indexed — lowercasing, stemming, stopwords, language rules. This is where most why-doesn't-my-search-match bugs come from, so interviewers probe whether you know indexing and querying must analyze the same way.

on this pageshow

questions

6

What steps turn a raw text field into indexed terms in a search engine's analysis pipeline?

level: juniorimportance: must knowfreq 68%

answer

  1. Three stages, in a fixed order
  2. Something happens before tokens even exist
  3. Lookup is exact string equality on terms
  4. Character filter, tokenizer, token filters
  5. The last stage is a chain, and order matters

basics

~20 s

Text analysis runs in three stages: character filtering (strip markup, map characters), tokenization (cut the stream into tokens), then token filtering (lowercase, fold accents, drop stopwords, stem, add synonyms). The surviving tokens become the index terms.

solid answer

~40 s

Analysis converts a string into the list of **terms** that actually go into the inverted index, and it has three stages. First, **character filters** rewrite the raw character stream — stripping HTML tags, mapping `&` to `and`, normalising Unicode forms. Second, a **tokenizer** cuts that stream into tokens; for space-delimited languages this is usually word-boundary splitting that also discards punctuation, and it records each token's position and offsets. Third, a chain of **token filters** transforms the token stream: lowercasing, accent/diacritic folding, stopword removal, stemming, synonym injection, n-gram generation. Whatever comes out the end is stored verbatim as terms. That matters because lookup in an inverted index is *exact string equality* on terms — all the fuzziness of "full-text search" is manufactured here, in analysis, not at match time.

code

text · 7 lines
text
input:   "The Quick-Brown Foxes!"
char filter (strip markup, map chars) -> "The Quick-Brown Foxes!"
tokenizer (word boundaries)           -> [The] [Quick] [Brown] [Foxes]
lowercase                             -> [the] [quick] [brown] [foxes]
stopword removal (english)            -> [quick] [brown] [foxes]
stemmer (english)                     -> [quick] [brown] [fox]
indexed terms: quick, brown, fox

go deeper

for a junior

Be ready to name the three stages in order and walk a short sentence through them, saying which terms end up in the index. Knowing that lookup is exact on those terms is the point.

for a middle

Explain what each stage can and cannot do, why something like markup stripping must happen before tokenization, and how filter ordering changes the resulting terms. Mention positions and offsets and what they are used for.

for a senior

Show judgment about which fields deserve analysis at all, and about tokenization choices that leak to users — hyphens, compounds, unspaced scripts. Be ready to debug a matching failure by reasoning about the emitted term stream.

for a principal

Own analysis as a schema-wide contract: how many analyzed variants a field really needs, what each costs in index size and reindex risk, and how you keep pipeline changes from silently breaking existing queries across teams.

## Why analysis exists An inverted index maps a **term** to the list of documents containing it. Lookup in that structure is an exact match on a byte string: the engine hashes or binary-searches the term dictionary for the term you hand it. Nothing about that structure is fuzzy. So every bit of tolerance users expect from search — case-insensitivity, matching *running* against *run*, ignoring the accent in *café* — has to be manufactured **before** anything is stored, by normalising both the document text and the query text into the same reduced vocabulary. That normalisation process is text analysis, and it runs in three ordered stages. ## Stage 1: character filtering The first stage sees the raw character stream and rewrites it before any token boundaries exist. Typical jobs: stripping HTML or XML markup so `<b>shoe</b>` does not index `b` as a word; character mapping so `&` becomes `and`, or so curly quotes become straight ones; Unicode normalisation, which matters because the same visible character can have several encodings (an *é* can be one code point or *e* plus a combining accent, and those are different byte strings — different terms). Anything that must happen before tokenization belongs here, because once the text is cut into tokens you can no longer merge or split across a boundary. ## Stage 2: tokenization The tokenizer cuts the character stream into tokens. For English and most European languages this is roughly word-boundary splitting: whitespace and punctuation separate tokens and are themselves discarded. Real tokenizers are more careful than `split(" ")` — they have rules for hyphens, apostrophes, URLs, email addresses and numbers, because those choices are visible to users. Whether `wi-fi` becomes one token or two decides whether a search for `wifi` can ever match. Tokenization is also where languages differ most sharply. Chinese, Japanese and Thai are written without spaces, so a tokenizer must either segment using a dictionary and statistical model, or fall back to overlapping character bigrams. German forms long compounds (*Fußballmannschaft*), so a decompounder may be needed for a search on *Mannschaft* to hit it. There is no language-neutral tokenizer, which is why engines ship per-language ones. The tokenizer also attaches metadata to each token that later stages and later queries depend on: a **position** (used by phrase and proximity matching) and character **offsets** into the original text (used by highlighting). Losing or shifting these is a real bug class. ## Stage 3: token filtering Token filters form a chain, each taking a token stream and emitting one. Order matters, because each filter sees only what the previous one produced. The common filters: - **Lowercasing** — so `Shoe` and `shoe` collapse to one term. - **Accent/diacritic folding** — `café` to `cafe`, so users typing without accents still match. - **Stopword removal** — dropping very common function words such as *the*, *of*, *and*. - **Stemming** — reducing inflected forms to a common stem, so *running*, *runs* and *ran* can meet. - **Synonym expansion** — injecting extra tokens at the same position, so *tv* also indexes or queries as *television*. - **N-gram generation** — emitting substrings or prefixes of a token to support partial matching. - **Length limits and trimming** — discarding absurdly long tokens produced by junk data. Ordering bugs are common: applying a synonym list *after* stemming means your synonym entries must be written in stemmed form, or they never fire. ## What actually lands in the index A useful mental exercise is to run a sentence through the pipeline by hand. `The Quick-Brown Foxes!` through a standard chain — tokenize, lowercase, remove English stopwords, stem — indexes roughly `quick`, `brown`, `fox`. The literal string is gone; only those three terms are searchable in that field. This is why a query that looks obviously correct can return nothing: it is being compared to `fox`, not to `Foxes`. ## Full-text fields versus exact-match fields Not every field wants this treatment. Identifiers, status codes, tags and anything used for exact filtering, sorting or grouping should be indexed **unanalyzed** — as a single term equal to the whole input value — because you want `IN_PROGRESS` to be one term, not `in` plus `progress`. Most engines therefore distinguish an analyzed full-text field type from a keyword-style exact field type, and let you index one source value both ways under different field names, choosing at query time which behaviour you want. Getting this choice wrong is the second-most common cause of "my filter matches nothing" after analyzer asymmetry. ## The takeaway Analysis is a lossy, deliberate transformation. Everything the index knows about a document's text is what the pipeline emitted, so when debugging matching problems the first move is always to ask what terms this text actually produced.

  • Why do search engines record a position and character offsets for each token?
    Positions let the engine answer phrase and proximity queries: it can check that *new* is immediately followed by *york* rather than just co-occurring in the document. Character offsets point back into the original text, which is how highlighters mark the matched span in the stored source even though the indexed term no longer resembles the original word.
  • When would you index a text field without running any analysis on it?
    For values that are identifiers rather than prose: status codes, SKUs, tags, usernames, enum values, keys used for exact filtering, sorting or grouping. There you want the entire value to be one term so equality is meaningful and grouping buckets are the real values, not word fragments.
  • Does changing the order of token filters change the index?
    Yes, materially. Filters chain, so each sees only the previous one's output. Stemming before synonym expansion means synonym entries must be written as stems or they never fire; folding accents after a language stemmer can defeat rules that depend on accented characters. Order is part of the configuration you must test, not a formality.

Analysis is like grinding beans before brewing: the index can only work with what came out of the grinder, and if the document and the query were ground differently, nothing lines up.

saying these in an interview costs you the question

  • Thinks the engine stores original text and matches it fuzzily
  • Believes tokenization is always just splitting on whitespace
  • Cannot name any stage before tokenization
  • Assumes token filter order does not matter
  • Runs full analysis on identifier and status-code fields

context

open as a page

Why must index-time and query-time text analysis produce compatible terms in a search engine?

level: middleimportance: must knowfreq 75%

basics

~20 s

Both sides must reduce text to the same term forms. The index stores analyzed terms, and a query can only match a term spelled identically after its own analysis, so mismatched pipelines silently return zero hits instead of an error.

open as a page

What is the difference between stemming and lemmatization in search text analysis?

level: middleimportance: should knowfreq 62%

basics

~20 s

Stemming strips affixes with language rules, fast and dictionary-free, producing stems that may not be real words. Lemmatization uses a dictionary and word context to return the true base form. Search engines usually choose stemming for speed and recall.

open as a page

What do stopwords cost in a search index, and why do modern engines usually keep them?

level: middleimportance: should knowfreq 52%

basics

~20 s

Stopword removal shrinks postings for very common words and speeds queries, but it destroys phrases and titles built from them. Modern engines usually keep stopwords because ranking already gives near-zero weight to ubiquitous terms and compression makes them cheap.

open as a page

What do you trade away by indexing edge n-grams to power prefix autocomplete?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Edge n-grams index every prefix of every token so a prefix search becomes a plain term lookup. The cost is a multiplied index and term dictionary, distorted term statistics, and a full reindex whenever the gram length range changes.

open as a page

Should synonym expansion happen at index time or query time in a search system?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Index-time expansion makes queries simple and fast but requires reindexing to change and distorts term statistics. Query-time expansion is editable immediately and keeps statistics honest, at the cost of larger queries and hard multi-word cases. Most teams expand at query time.

open as a page