What does the search_quote_analyzer mapping parameter do in Elasticsearch?
answer
- A third analyzer slot on the field
- Only one kind of query text triggers it
- Quotation marks are the signal
- Stopwords help loose queries, hurt phrases
- Not used by match_phrase
basics
~10 sIt sets a separate analyzer for text inside quotes in query_string and simple_query_string queries. The usual purpose is to keep stopwords in quoted phrases while the ordinary search analyzer strips them.
solid answer
~40 s`search_quote_analyzer` is a third analyzer slot on a `text` field, applied only to quoted portions of `query_string` and `simple_query_string` queries. It exists because stopword removal helps ordinary term matching but wrecks phrases: if the search analyzer drops `the`, the phrase `"the quick brown fox"` becomes `quick brown fox`, which also matches a document containing `quick brown fox` without the article — not what a user asking for an exact phrase meant. The standard configuration indexes with a stopword-preserving analyzer so positions are complete, uses a stopword-removing analyzer as `search_analyzer` for unquoted terms, and points `search_quote_analyzer` back at the stopword-preserving chain so quoted phrases are matched with every token present. It is a narrow parameter — a dedicated `match_phrase` query uses the field's `search_analyzer`, not this one.
code
json · 12 lines{
"mappings": {
"properties": {
"title": {
"type": "text",
"analyzer": "keep_stopwords",
"search_analyzer": "drop_stopwords",
"search_quote_analyzer": "keep_stopwords"
}
}
}
}go deeper
It is enough to know a text field can name a separate analyzer for quoted search input, and that it changes only how the query is tokenized.
Explain the stopword-versus-phrase trade-off it resolves and show the three-slot mapping that makes quoted phrases exact while loose queries stay clean.
Add the scope detail — it fires only for quoted text in the query-string family — and the constraint that the index side must have preserved the tokens.
Judge whether the parameter is even warranted: with BM25 saturation, keeping stopwords indexed and skipping removal entirely is often the simpler system-wide choice.
## The problem it solves Stopword removal is a classic trade-off. Dropping high-frequency function words (`the`, `a`, `of`, `and`) shrinks the index a little, avoids meaningless matches, and stops those words dominating a loose keyword query. But phrase matching depends on **positions**, and positions depend on what the analyzer produced. Suppose you strip stopwords on both sides. The document `"the quick brown fox"` indexes as `quick`(pos 2) `brown`(3) `fox`(4) — positions are preserved as gaps, but the token `the` does not exist. Now the quoted query `"the quick brown fox"` analyzes to `quick brown fox`, and it matches a document that literally reads `quick brown fox` with no article at all. The user asked for an exact phrase and got a looser one. Suppose instead you keep stopwords everywhere. Phrases are exact, but every ordinary keyword query drags `the` along, adding a term that matches nearly every document and contributes noise. ## The three-analyzer configuration Elasticsearch's answer is to split the query side in two: - `analyzer` — index time, **keeps** stopwords, so every token has a real position and the index can support exact phrases. - `search_analyzer` — for unquoted query text, **removes** stopwords, so loose queries are not polluted by them. - `search_quote_analyzer` — for quoted query text, **keeps** stopwords, matching the index side exactly so a phrase means what it says. Note that the index side must keep stopwords for this to work: an analyzer cannot match a term that was never stored. The quote analyzer only decides which terms the *query* contributes. ## Scope: where it actually applies This is the detail interviewers probe. `search_quote_analyzer` governs text inside quotation marks in `query_string` and `simple_query_string` queries — the query types that parse user-entered syntax and therefore have a notion of "the user quoted this part". A `match_phrase` query is a phrase by construction, with no quoting syntax to detect, and uses the field's ordinary `search_analyzer` (or an `analyzer` named inside the clause). So the parameter is really a feature of user-facing query-string search boxes, not of programmatic phrase queries. That also tells you when to reach for it: you have an end-user search box that accepts quotes as an exact-phrase operator, and you strip stopwords for the unquoted case. If your application builds `match` and `match_phrase` clauses itself, you have full control per clause and simply set `analyzer` on the phrase clause instead — the parameter buys you nothing. ## Mutability Because it is a query-side parameter like `search_analyzer`, it affects nothing stored and can be reasoned about the same way: it changes how query input is tokenized, immediately, for documents indexed at any time. The constraint that does bind is the index side — if the field was indexed with a stopword-removing analyzer, no query-side configuration recovers the missing tokens, and getting exact phrases back means reindexing with a stopword-preserving chain. ## Modern context Some of the pressure that produced this parameter has eased. Modern practice often keeps stopwords at index time and relies on BM25's saturation and inverse-document-frequency weighting to make very common terms contribute almost nothing, rather than deleting them outright; the `common` terms handling that older stacks used for this has fallen out of favour for the same reason. Where stopwords are still removed for a specific language analyzer, `search_quote_analyzer` remains the clean way to exempt quoted phrases. ## What a good answer sounds like Name the slot, name the trigger (quoted text in query-string family queries), name the motivating case (stopwords versus phrase exactness), and add the constraint that the index-time analyzer must have preserved the tokens in the first place. Admitting that it does not apply to `match_phrase` is the detail that separates recall from real familiarity.
- Does search_quote_analyzer affect a match_phrase query?No. It applies to quoted text in `query_string` and `simple_query_string`, which parse user-entered quoting syntax. A `match_phrase` clause is already a phrase by construction and uses the field's `search_analyzer`, or an `analyzer` named inside the clause. If your application builds phrase clauses itself, set the analyzer there instead.
- Why must the index-time analyzer keep stopwords for this to work?Because a query can only match terms that exist in the index. If the index-time analyzer removed `the`, no query-side configuration can find it, and the quoted phrase still matches loosely. The pattern therefore pairs a stopword-preserving index analyzer with a stopword-removing search analyzer and a stopword-preserving quote analyzer.
- Is stripping stopwords still the usual advice?Less so. BM25 weights terms by inverse document frequency and saturates term frequency, so very common words contribute little on their own, and keeping them preserves exact phrase capability and avoids odd failures on queries like "to be or not to be". Stopword removal survives mainly inside language-specific analyzers and where index size is genuinely constrained.
saying these in an interview costs you the question
- Thinks it changes how documents are indexed
- Believes match_phrase uses the quote analyzer
- Configures it while the index analyzer already strips stopwords
- Confuses it with the slop parameter on phrase queries
- Assumes it applies to every query type on the field