Why is an Elasticsearch wildcard query like *smith slow, and what replaces it?
answer
- Terms are stored in sorted order
- Without a fixed start, ordering buys nothing
- Cost tracks distinct values, not hits
- Flip the string and the anchor comes back
- Pay at index time instead of query time
basics
~20 sA leading wildcard gives no fixed prefix, so Elasticsearch must walk every term in that field's dictionary and test each one. The fix is index-time: index_prefixes, an edge n-gram or reversed field, or the wildcard field type.
solid answer
~50 sTerms are stored in a sorted dictionary, so a pattern with a fixed prefix — `smith*` — seeks straight to that region and enumerates a small range. `*smith` has no anchor, so Lucene must visit **every** term in that field on **every** shard and test it, and the cost scales with field cardinality rather than with the number of matches. `regexp` has the same shape, plus a `max_determinized_states` guard (default 10,000) that rejects patterns whose automaton explodes. The answer is always to move work to index time: `index_prefixes` on a `text` field builds a dedicated prefix index; an `edge_ngram` analyzer or the `search_as_you_type` type handles prefix-as-you-type; a reversed copy of the field turns a leading wildcard into a trailing one; and the `wildcard` field type is purpose-built for grep-like patterns over high-cardinality machine data. Operationally, `search.allow_expensive_queries` can be set to false to reject these outright.
code
json · 5 lines// scans every term in the field on every shard
{ "query": { "wildcard": { "email": { "value": "*@example.com" } } } }
// after modelling the domain as its own keyword field
{ "query": { "term": { "email_domain": "example.com" } } }go deeper
Be ready to say that a pattern starting with a wildcard cannot use the sorted term dictionary and so has to check every term in the field.
Explain that cost scales with distinct term count rather than hits, and name at least one index-time alternative such as index_prefixes or an edge n-gram analyzer.
Show that you diagnose from the slow log, distinguish prefix from infix and suffix patterns, and pick the matching index-time structure — reversed field, wildcard type, or a modelled field — rather than tuning the query.
Own the guardrails: whether expensive queries are permitted on a shared cluster, how ad-hoc dashboard queries are contained, and the storage-versus-latency budget for index-time structures across teams.
## Why the leading wildcard is the expensive part An inverted index stores each field's terms in a **sorted term dictionary**. That ordering is what makes lookups cheap: an exact term is a seek, and a pattern with a fixed prefix is a seek plus a walk of one contiguous range. `smith*` visits only terms between `smith` and the next term that does not start with those five characters. Remove the anchor and the structure stops helping. `*smith` could match a term starting with any byte, so the only correct implementation is to enumerate the entire term dictionary for that field and test each term against the pattern. The cost is therefore proportional to the **cardinality of the field**, not to how many documents match — a query returning three hits can still visit millions of terms — and it is paid on every shard, in parallel, for every request. The same reasoning covers `regexp`. A pattern with a literal prefix can be executed against a bounded range; one that begins with `.*` cannot. Lucene compiles the regular expression into a deterministic automaton, and `max_determinized_states` (default 10,000) rejects patterns whose automaton would blow up, which is a safety valve against a single query consuming a node. Two syntax details matter: Lucene's regexp flavour is not PCRE, and the pattern is **anchored to the whole term** — you do not write `^` and `$`, and a pattern that matches only part of a term matches nothing. A third trap is that all of these are term-level and therefore unanalyzed. `wildcard` on a `text` field matches against individual analyzed tokens, so a pattern spanning a space never matches; on a `keyword` field it matches the whole value but is case-sensitive unless you pass `case_insensitive: true` or lowercase at index time with a normalizer. ## Move the work to index time Every good fix restructures what is indexed so the query regains an anchor. **`index_prefixes` on a text field.** This mapping parameter tells Elasticsearch to index prefixes of each term as a separate hidden field, with `min_chars` defaulting to 2 and `max_chars` to 5. A `prefix` query on that field is then answered by a term lookup rather than an enumeration. It costs index size and indexing time, and it only helps prefixes within the configured length range. **`edge_ngram` at index time.** A custom analyzer with an `edge_ngram` token filter indexes `e`, `el`, `ela`, `elas`… for each term, so a prefix search becomes an ordinary term match. The critical detail is to use a plain analyzer at **search** time — configure `search_analyzer` separately — otherwise the query is also n-grammed and precision collapses. The `search_as_you_type` field type packages this pattern with sensible defaults. **A reversed field.** Index a copy of the field through a `reverse` token filter. A leading wildcard `*smith` becomes a trailing wildcard `htims*` against the reversed copy, which has a fixed prefix and executes cheaply. This is the classic answer for suffix search over domain names, filenames or identifiers. **The `wildcard` field type.** Purpose-built for grep-like patterns over high-cardinality, machine-generated content such as log lines and URLs. It indexes n-grams of the value and uses them to shortlist candidates, then verifies each candidate against the full value held in doc values. It accepts arbitrary wildcard and regexp patterns, including leading ones, at a fraction of the cost of the same pattern on a plain `keyword` field. The trade-off is index size and that it is tuned for long, unstructured values rather than short enum-like ones. **Restructure the data.** Often the wildcard is a symptom of a modelling gap. If people search `*@example.com`, index the email domain as its own `keyword` field and filter on it exactly. That converts an unbounded scan into a single term lookup and is usually the best available answer. ## Operational guardrails `search.allow_expensive_queries` is a dynamic cluster setting, defaulting to true, that when set to false rejects the query classes with unbounded cost: `wildcard`, `regexp`, `prefix` on fields without `index_prefixes`, `fuzzy`, `joining` queries, `script` queries and range queries on `text` and `keyword` fields. Turning it off on a shared cluster is a blunt but effective way to stop one team's exploratory query from destabilizing a tier — with the caveat that it will also break legitimate uses, so it belongs behind a conversation, not a unilateral change. Also worth naming: these queries are cheap to *write*, so they arrive through Kibana and ad-hoc dashboards more often than through reviewed application code. Slow-log thresholds on the query phase and a search-thread-pool queue alert are the practical detection mechanism, and the fix in the moment is a `search.max_...` guard or cancelling the task, while the durable fix is the mapping change. ## What to say when asked to "just make it faster" The honest framing is that a leading wildcard has no cheap execution on a plain term dictionary — the structure that makes search fast is precisely the one it defeats. You either accept a linear term scan on a small-cardinality field, or you pay index-time cost and storage to build a structure the pattern *can* seek into. Which of the four index-time options you choose is a function of pattern shape, value length and cardinality.
- How does a reversed-field copy turn a leading wildcard into a cheap query?Index a second copy of the field through a `reverse` token filter so `smith` is stored as `htims`. A search for `*smith` is rewritten as `htims*` against that copy, which has a fixed prefix and seeks into one contiguous range of the term dictionary. The cost is roughly a doubled index for that field and a rewrite step in the application.
- When is the wildcard field type a better answer than index_prefixes?When the patterns are not prefixes and the values are long and high-cardinality — log lines, URLs, file paths — where users genuinely want grep semantics anywhere in the string. `index_prefixes` only accelerates prefix matching within a configured character range, so it cannot serve an infix or suffix pattern at all.
- What does max_determinized_states protect against in a regexp query?Lucene compiles the pattern into a deterministic automaton, and certain patterns — heavy nesting with repetition — cause a state explosion that would consume a node's memory and CPU. The limit, defaulting to 10,000, rejects the query instead. Raising it should be a deliberate, measured decision rather than a reflex when a pattern is refused.
- Why does a wildcard query on a text field fail to match a pattern containing a space?Wildcard is term-level and matches against individual indexed tokens, and the analyzer already split the value on whitespace, so no single token contains a space. Run the pattern against the keyword sub-field, which holds the whole value as one term, or restructure the query into a phrase-shaped full-text clause.
saying these in an interview costs you the question
- Suggests adding replicas or heap to fix a leading wildcard
- Assumes wildcard queries use the index like a prefix does
- Thinks cost scales with the number of matching documents
- Writes Lucene regexp patterns with anchor characters
- Applies wildcard patterns across whitespace on a text field