skip to content

What does the Elasticsearch shard request cache store, and why do queries containing now never hit it?

level: seniorimportance: should knowfreq 38%

answer

  1. The cache key is the request body itself
  2. Only some kinds of response are cached
  3. size:0 is the default condition
  4. Refresh invalidates the entries
  5. Rounded date math makes the key stable

basics

~20 s

It caches per-shard search results keyed on the whole JSON request body, and by default only for requests with size:0 — aggregations, suggestions and the total hit count. A now value resolves to a new instant each request, producing a new key and a guaranteed miss.

solid answer

~50 s

The shard request cache lives on each node and stores the **local result a shard produced** for a search, so a repeated search can skip re-executing on that shard. Two properties define its behaviour. First, the cache key is the **entire JSON request body**: any byte that differs — a changed filter value, a reordered clause, a fresh timestamp — is a different key. Second, by default it only caches requests with `size: 0`, meaning aggregations, suggestions and `hits.total`, not the hits themselves. That makes it a dashboard and analytics cache. Entries are invalidated when the shard refreshes, so cached results are never staler than the refresh interval. A range filter such as `"gte": "now-24h"` resolves to a different millisecond on every request, so the body — and therefore the key — is new every time and the hit rate is zero. The fix is date-math rounding: `now-24h/h` produces a body that is byte-identical for a whole hour.

code

json · 2 lines
json
// never cached: the resolved range moves every request
{ "range": { "@timestamp": { "gte": "now-7d", "lt": "now" } } }

go deeper

for a junior

Know that Elasticsearch caches results of repeated searches per shard and that the cache is keyed on the request you send.

for a middle

Explain the two defining rules — the key is the whole request body, and only size:0 results are cached by default — and that refresh invalidates entries.

for a senior

Diagnose a zero hit rate from the statistics, connect it to date math resolving anew each request, and apply rounded ranges with an unrounded tail rather than turning the cache off.

for a principal

Treat query shape as a caching contract: standardise how dashboards express time windows and serialise request bodies, because per-query hit rate is what decides the search tier's CPU budget.

## What it is Each node holds a shard request cache covering the shards it hosts. When a search reaches a shard, the shard checks whether it has already produced a result for this exact request against the current view of its data; if so it returns the cached result and does no work. The cache is enabled by default. Note what is cached: a **per-shard** result, not the final merged response. The coordinating node still fans out, still merges, still reduces aggregations. What it avoids is the shard-level execution, which is where the CPU goes. ## The two rules that govern every question about it **Rule 1 — the key is the whole request body.** Not a normalised or canonicalised form of it, the body. Consequences follow immediately: - Two logically identical queries written with clauses in a different order are different keys. - A query that embeds the current time, a random seed, or a user id is a fresh key each time. - Clients that pretty-print or reorder JSON non-deterministically silently destroy the hit rate. **Rule 2 — by default only `size: 0` requests are cached.** The cache stores `hits.total`, aggregation results and suggestions. It does not store the hit documents. So a Kibana visualisation, a facet count, a dashboard panel — all `size: 0` aggregation requests — are cacheable. A user-facing search returning ten products is not. ## Invalidation Entries for a shard are invalidated when that shard **refreshes** and its data has changed. This preserves Elasticsearch's near-real-time contract: a cached result is never older than the last refresh, so the cache can never show you data that a fresh search would not also show. There is no correctness window to reason about, which is why the cache can be on by default. The practical corollary is that on a heavily indexed index with a one-second refresh interval, cache entries live about one second and the cache buys little. On an index that stopped receiving writes — yesterday's daily index, a rolled-over backing index, a static catalogue — entries live until evicted, and the hit rate on a dashboard that queries the last 30 days can be excellent. ## Controls - `index.requests.cache.enable` — per-index, defaults to enabled. - `request_cache=true|false` — per-request override on the search URL. - `indices.requests.cache.size` — node-level cap on how much heap the cache may use, expressed as a percentage of the heap. - `GET /<index>/_stats/request_cache` and the node stats — report `hit_count`, `miss_count`, `evictions` and `memory_size`. Measure before theorising; a near-zero hit rate with a high miss count is the signature of a varying key. ## The `now` trap in detail Date math is evaluated per request. A dashboard sending ``` { "range": { "@timestamp": { "gte": "now-24h" } } } ``` every ten seconds produces a request body whose resolved range moves continuously. Every execution is a miss, every execution populates a new entry, and those entries evict useful ones — so the cache is not merely useless here, it is actively counterproductive. Rounding fixes it. `now-24h/h` rounds down to the hour, so every request within the same hour produces an identical key and, on an index that is no longer being written to, an identical cached answer. The cost is precision: the window snaps to hour boundaries. When precision matters, split the query into two ranges combined in a `bool`: a large rounded range that is stable and cacheable, plus a short unrounded tail covering the most recent minutes. The expensive part of the work — scanning the bulk of the history — is served from cache, and only the small recent slice is recomputed. ## Where it sits among the other caches This is not the only cache in the search path, and mixing them up is a common interview stumble. The shard request cache stores whole shard-level responses keyed on the request body. Separately, a node-level query cache stores filter results, and the operating system page cache holds the index files themselves. They are configured, sized and invalidated independently. ## What interviewers listen for The strong answer names the `size: 0` restriction unprompted, states that the key is the request body, explains refresh-based invalidation as the reason it is safe to leave on, diagnoses `now` correctly, and reaches for date rounding — ideally with the rounded-plus-tail pattern — rather than proposing to disable the cache or to add a TTL.

  • Why is it safe for this cache to be enabled by default?
    Because invalidation is tied to the shard's refresh: when the shard refreshes and its data changed, its entries are dropped. A cached result can therefore never be older than what an uncached search would return at that moment, so the near-real-time contract holds and no staleness window has to be reasoned about.
  • How would you keep precision on a recent-data query while still getting cache hits?
    Split the time range in a bool query: one clause covering the bulk of the window with rounded date math, such as `gte: now-7d/h, lt: now/h`, which is byte-stable for an hour and cacheable, plus a second clause covering only the unrounded tail from `now/h` to `now`. The expensive historical scan is cached; only the small recent slice is recomputed.
  • How do you tell whether the cache is helping at all?
    Read the request-cache statistics per index or per node: hit count, miss count, evictions and memory size. A high miss count with almost no hits points at a varying request body; high evictions with a decent hit rate points at an undersized cache relative to the number of distinct queries.

saying these in an interview costs you the question

  • Thinks it caches the returned hits of ordinary searches
  • Says entries expire on a fixed TTL rather than on refresh
  • Confuses it with the node-level filter query cache
  • Proposes disabling the cache instead of rounding dates
  • Believes reordering JSON clauses still hits the same entry

context