In an Elasticsearch search request, what does size: 0 change and why does the shard request cache need it?
answer
- It removes an entire search phase
- Aggregations still scan every matching document
- One cache stores whole shard-level responses
- Eligibility depends on returning no documents
- Refresh decides how long an entry survives
basics
~20 ssize: 0 returns no documents, so Elasticsearch skips the fetch phase and only computes aggregations and the hit count. The shard request cache only stores shard-level results for such requests, making repeated identical dashboard queries nearly free until the next refresh.
solid answer
~50 s`"size": 0` tells Elasticsearch to return no hits. The query phase still runs — documents are still matched and scored where scoring applies, and aggregations still consume every match — but the **fetch phase is skipped**, so no `_source` is loaded and no top-document list is merged or transported. For an aggregation-only request, that is pure saved work. It also unlocks the **shard request cache**, which is enabled by default (`index.requests.cache.enable`) and caches the *shard-level* response — the aggregation results, suggestions, and `hits.total` — keyed by the request body. Requests with `size` greater than zero are not cached at all, because their result includes documents that would go stale. Entries are invalidated whenever the shard refreshes, so cached results are never staler than an uncached search, and the cache pays off most on read-mostly indices with a longer `refresh_interval`. The node-level size defaults to a small fraction of the heap (`indices.requests.cache.size`, 1%).
code
json · 10 lines{
"size": 0,
"track_total_hits": false,
"query": { "range": { "@timestamp": { "gte": "now-24h/h" } } },
"aggs": {
"per_hour": {
"date_histogram": { "field": "@timestamp", "fixed_interval": "1h" }
}
}
}go deeper
Recall that aggregation-only requests should set size to 0 so no documents come back, and that this makes the response smaller and the request faster.
Explain that the fetch phase is skipped while aggregations still process every match, and that only size: 0 requests are eligible for the shard request cache, whose entries die on refresh.
Show how you raise the hit rate in production: longer refresh intervals on read-mostly indices, rounded date math, stable request bodies, and knowing when caching cannot save you and pre-aggregation must.
Own the dashboard cost model: which tiers serve interactive analytics, whether panels are allowed to issue ad-hoc ranges, and when summary indices replace repeated live aggregation entirely.
## The two phases, and what size: 0 removes A normal Elasticsearch search runs in two phases. In the **query phase** each shard matches the query, collects the top `size` document IDs and their sort values, and computes any aggregations. The coordinating node merges those partial results and decides the global top `size` documents. In the **fetch phase** it then goes back to the relevant shards to retrieve the actual `_source` (or stored fields) for exactly those documents and assembles the response. Setting `"size": 0` removes the fetch phase entirely. There are no top documents to retrieve, so there is no second round trip, no `_source` decompression, and a much smaller response body. What it does *not* remove is the aggregation work: aggregations run over every matching document regardless of `size`, so `size: 0` is a cost reduction, not a cost avoidance. It is simply the correct shape for a request whose answer is a set of buckets rather than a set of documents — which is what almost every dashboard panel is. If you also want to avoid counting matches precisely, `"track_total_hits": false` (or a numeric bound) is a separate, complementary lever. ## The shard request cache Elasticsearch keeps a per-node cache of **shard-level** search responses. Unlike a result cache in front of the whole cluster, it stores what one shard returned for one request, so a query hitting ten shards can have nine cached shard results and one fresh one. Three properties define it: 1. **It only caches requests with `size: 0`.** The stored payload is the aggregation results, suggestions and the total hit count — deliberately, the parts of a response that do not include document bodies. A request that returns hits is not cached even if the request cache is enabled on the index. 2. **The cache key is the whole request body.** Byte-for-byte identical JSON hits the same entry; a reordered clause or a different whitespace-normalized form may not. This matters when an application builds query JSON dynamically. 3. **Entries are invalidated when the shard refreshes.** That is what makes the cache safe: results can never be older than the searcher a fresh query would have used. It is also what determines the hit rate — on an index refreshing every second, an entry lives at most a second. Controls: `index.requests.cache.enable` (index setting, on by default), `indices.requests.cache.size` (node setting, 1% of the heap by default), the `request_cache=true|false` query-string parameter for a single request, and `POST /my-index/_cache/clear?request=true` to drop entries. ## Making the cache actually work for you The two levers that matter are refresh frequency and request stability. On a **read-mostly** index — a product catalogue rebuilt nightly, yesterday's log index that will never receive another document — raising `index.refresh_interval` (or letting rollover leave old indices static) lets cache entries survive far longer, and the hit rate climbs accordingly. Indices behind an ILM warm or cold tier are excellent cache citizens for the same reason. On the request side, **round your date math**. A dashboard sending `"gte": "now-24h"` on every reload with millisecond-precision evaluation gives you a filter whose results shift constantly; `"gte": "now-24h/h"` rounds the boundary to the hour, so every reload within that hour is asking the identical question and can reuse cached work — and, separately, lets Lucene's segment-level query cache reuse the filter bitset. The cost is that the window edge moves in steps rather than smoothly, which for a 24-hour dashboard nobody notices. Also keep the request body stable: same field ordering, same explicit defaults, no injected timestamps or request IDs inside the body. ## What it does not do The request cache does not make an expensive aggregation cheap the first time, and it does nothing for a cold, one-off analytical query — the archetypal case where a user drags a new date range across a dashboard and every panel misses. It also does not help write-heavy indices, where refresh invalidation outruns reuse. When repeated aggregation cost is the real problem, the answer is upstream: pre-aggregate into a summary index with a transform, or downsample time-series data, so the dashboard queries a small pre-computed index that caches beautifully and is cheap even on a miss. ## Interview register A good answer separates *what `size: 0` saves* (the fetch phase and response size) from *what it unlocks* (cache eligibility), notes that aggregations still process every match, and names refresh invalidation as the property that makes the cache both safe and limited.
- Why does rounding date math to now-24h/h improve caching, and what does it cost?Rounding makes the effective range boundary change only once an hour, so successive dashboard reloads send an equivalent query whose shard-level results can be reused, and Lucene can cache the filter bitset too. The cost is a window edge that advances in steps rather than continuously — usually invisible on a 24-hour view.
- Does the shard request cache risk serving stale aggregation results?No. An entry is invalidated when the shard refreshes, so a cached result is never based on an older searcher than a fresh query would use. The tradeoff is hit rate, not correctness: an index refreshing every second gives entries about a second of useful life.
- How does track_total_hits interact with an aggregation-only request?They are independent. `size: 0` skips fetching documents, while `track_total_hits` controls how hard Elasticsearch works to count matches — counting exactly can force full evaluation where an early-terminating collector would otherwise stop. On aggregation-only panels that display no result count, setting it to false removes work the aggregation does not need.
saying these in an interview costs you the question
- Thinking size: 0 makes aggregations skip documents
- Believing the request cache stores individual documents
- Expecting caching to work on requests that return hits
- Assuming cached results can be stale after a refresh
- Confusing the shard request cache with Lucene's query cache