Which parts of an Elasticsearch bool query can the node query cache accelerate?
answer
- Only one context is eligible at all
- Yes-or-no answers are what can be stored
- Stored per segment, not per index
- Segments are immutable, so no invalidation
- Repeated date math must be rounded to hit
basics
~20 sOnly clauses running in filter context — the filter and must_not slots, plus anything inside constant_score. The cache stores per-segment doc-id sets for those clauses, so a repeated filter is answered from a bitset instead of re-executed. Scored clauses are never cached this way.
solid answer
~60 sOnly clauses that run in **filter context**: the `filter` and `must_not` slots of a `bool` query, and anything wrapped in `constant_score`. Those produce a pure yes/no answer per document, so the result is a doc-id set that can be stored; a `must` or `should` clause produces a score per document and cannot be reused that way. The cache lives **per node**, shared by every shard on it, holds one entry **per segment per query**, and is bounded by `indices.queries.cache.size`. Two heuristics decide what actually lands in it: Lucene only caches a query after it has been seen often enough to be worth it — very cheap queries such as a plain `term` are re-run rather than cached — and only on segments large enough to pay back the bitset. Entries stay valid because segments are immutable, and vanish when the segment is merged away. The practical corollary is that filters must **repeat byte-for-byte** to hit: `"gte": "now-1h"` is a new query every millisecond, while `"gte": "now-1h/m"` repeats for a whole minute.
code
json · 11 lines{
"query": {
"bool": {
"must": [ { "match": { "message": "timeout" } } ],
"filter": [
{ "term": { "service": "checkout" } },
{ "range": { "@timestamp": { "gte": "now-1h" } } }
]
}
}
}go deeper
Know only that filter clauses can be cached and scored clauses cannot, and that this is one reason exact-value predicates belong in the filter slot.
Explain what an entry actually is — a per-segment doc-id set for one query — and why a score-producing clause has nothing reusable to store.
Diagnose a low hit rate: unrounded now date math producing a new query per request, clauses sitting in must, or a policy that skips cheap and infrequent queries. Explain why immutable segments make invalidation a non-issue.
Own query-shape governance so filters recur across users and tenants instead of being unique per request, and weigh refresh rate against cache effectiveness when sizing a cluster.
## What the cache is The node query cache is a per-node structure shared by all shards allocated to that node. It maps a query to a **per-segment set of matching document ids** — in Lucene terms, a cached `DocIdSet`, often a compressed bitset. On a hit, the clause does not touch postings at all: the engine intersects the cached set with the rest of the query. Its size is bounded by the static node setting `indices.queries.cache.size`, expressed as a percentage of heap; because it is static, changing it requires a node restart. It can be disabled per index with `index.queries.cache.enabled`. ## Why only filter context A clause in query context must answer "how well?" per document, and that answer is a float that depends on the query terms, field norms and collection statistics. There is nothing compact and reusable to store. A clause in filter context answers "yes or no", which for a segment is exactly a set of doc ids — small, compressible and reusable across every request that asks the same question. That is the whole reason "move it to `filter`" is standard advice: it is not only about skipping score arithmetic, it is about making the clause *cacheable at all*. Eligible: `bool` `filter`, `bool` `must_not`, the inner query of `constant_score`, and clauses nested inside those. Not eligible: `bool` `must`, `bool` `should`, `dis_max` sub-queries, the `positive` and `negative` sides of a `boosting` query. ## The caching policy — why a filter can be eligible and still not cached Eligibility is not admission. Lucene applies a usage-tracking policy: a query has to be seen repeatedly before its result is worth materialising, and *cheap* queries are deliberately never cached because re-executing them costs less than storing and looking them up — a simple `term` query on a keyword field is the canonical example. Costlier shapes, such as multi-term queries and numeric or date range queries over points, are the real beneficiaries. There is also a segment-size heuristic: caching is skipped on small segments, because tiny segments are cheap to scan and are likely to be merged away soon, wasting the entry. So the honest answer to "why is my filter not cached?" has three branches: it is not in filter context; it is in filter context but too cheap or too infrequent for the policy; or it is never the *same* query twice. ## The date-math trap The third branch is the one that bites in production. Dashboards and alerting queries almost always carry a time window: ```json { "range": { "@timestamp": { "gte": "now-1h" } } } ``` `now` resolves to the current millisecond at execution, so every request produces a *different* range query. It is eligible, it is expensive enough to be worth caching, and it never hits, because the key is new each time. Rounding fixes it: ```json { "range": { "@timestamp": { "gte": "now-1h/m" } } } ``` Now the lower bound is identical for a whole minute and every request in that minute can reuse the entry. Rounding to the hour or day widens the reuse window further at the cost of window precision. A common production shape splits the range in two: a rounded, cacheable clause covering the bulk of the window plus a small unrounded clause covering the recent tail. ## Invalidation is a non-event Because Lucene segments are immutable, a cached entry for a segment can never go stale — a newly indexed document lands in a new segment, which simply has no entry yet and is evaluated normally. Entries are dropped when their segment disappears in a merge, and evicted under LRU pressure when the cache is full. There is no invalidation logic to reason about, which is a genuinely nice property to be able to state in an interview. The consequence for write-heavy indices is that new segments arrive constantly, so a share of every query is always uncached; heavy refresh rates therefore reduce effective hit ratios, and a very low `refresh_interval` costs more than just merge pressure. ## Not the same as the other caches Elasticsearch has more than one cache and mixing them up is a classic interview stumble. The node query cache described here caches *clause results per segment*. The shard request cache is a different mechanism at a different granularity, and field data structures for sorting and aggregations are different again. If asked, name the node query cache precisely and say it caches filter-clause doc-id sets. ## What to do with this in practice 1. Put every exact-value predicate in `filter` or `must_not`. 2. Round date math so the filter text repeats. 3. Keep filter clauses *stable* across requests — build them from a small, enumerable set of shapes rather than from freely-varying user input, so the same handful of filters recur across all users. 4. Do not expect a plain `term` filter to show cache benefit; its win is skipping scoring, not caching.
- A dashboard reruns the same filters every few seconds but the node query cache barely helps. What would you check first?Whether the time range uses unrounded date math. A range with gte of now-1h resolves to a new millisecond on every request, so it is a different query each time and can never hit. Rounding to now-1h/m or a coarser unit makes the clause identical for the whole rounding window. After that, check the clauses are actually in the filter slot rather than must.
- Why does a cached entry never need invalidating when new documents are indexed?Entries are keyed per Lucene segment, and segments are immutable once written. A newly indexed document goes into a new segment that simply has no cache entry yet and is evaluated normally. Existing entries stay correct until their segment is merged away, at which point they are dropped rather than updated.
- Is every filter-context clause cached the moment it runs?No. Eligibility and admission are different. Lucene's usage-tracking policy caches a query only after it has been seen frequently enough to justify the entry, skips queries cheap enough to just re-run — a plain term query being the usual example — and skips segments too small to repay the cost. Expensive, repeated range and multi-term filters are the real beneficiaries.
- How does a high refresh rate interact with node query cache effectiveness?Each refresh creates new segments, and a new segment starts with no cached entries for any query. On a write-heavy index with a very short refresh interval, a meaningful share of the searched data always sits in fresh, uncached segments, so hit ratios fall. That is one more cost of aggressive refreshing, alongside merge pressure.
saying these in an interview costs you the question
- Says the node query cache stores whole search responses
- Claims must and should clauses are cached too
- Thinks every filter clause is cached on first execution
- Confuses the node query cache with the shard request cache
- Ignores that unrounded now date math defeats caching entirely