Why does Elasticsearch reject a search with from=10000 and size=10?
answer
- Skipping is not free when results are spread over shards
- Each shard must return its own top slice
- The coordinating node merges shards × (from + size)
- One index setting caps the sum at 10,000
- Its name ends in max_result_window
basics
~20 sBecause from + size exceeds index.max_result_window, which defaults to 10,000. Deep paging makes every shard build a top-(from+size) list that the coordinating node merges and then throws almost all of away, so Elasticsearch caps the depth instead.
solid answer
~40 s`from` skips hits and `size` returns the next N, but that skipping is not free in a distributed index. To answer page 1000, **every** shard must produce its own top `from + size` hits, and the coordinating node merges all of them — with 5 shards and `from: 9990, size: 10` that is 50,000 hits sorted and merged so 10 can be returned. Cost grows linearly with depth on both memory and network. Elasticsearch therefore refuses the request once `from + size` passes `index.max_result_window`, which defaults to 10,000, with an error saying the result window is too large. Raising the setting per index is possible but just moves the cliff; the real fix is `search_after` with a point in time for deep paging, or a scroll/PIT export if you actually need every hit.
code
json · 6 lines// GET /products/_search -> rejected: from + size = 10010
{
"from": 10000,
"size": 10,
"sort": [ { "price": "asc" } ]
}go deeper
Recall that from skips hits, size returns them, and that from + size cannot exceed 10,000 by default. Be able to name search_after as the alternative rather than reaching for the settings API.
Explain the query-then-fetch mechanics: every shard returns its own top from+size list and the coordinating node merges shards × (from + size) entries. Tie the cap directly to that arithmetic.
Show judgment about when raising index.max_result_window is defensible and when it converts a clean error into an intermittent heap incident. Expect to diagnose which client is actually requesting page 900.
Own the product-level answer: deep pagination is usually a symptom of a missing export path or weak relevance. Decide whether the system offers paging, cursors, or bulk export, and enforce that contract at the API edge.
## What from and size do In the Elasticsearch search API, `size` is the number of hits to return and `from` is the number of leading hits to skip. `from: 20, size: 10` gives you the third page of a ten-per-page list. Both are pure result-set slicing — they do not change which documents match, only which slice of the ordered match list is returned to the client. `from` defaults to 0 and `size` defaults to 10. ## Why depth is expensive in a distributed index A search runs in two phases. In the **query phase** the coordinating node broadcasts the request to one copy of every shard. There is no global ordering available to any single shard, so each shard cannot know which of *its* documents fall into the caller's requested page — it must assume the worst and return its own best `from + size` results (document ids plus sort values) to the coordinating node. The coordinating node builds a priority queue of `number_of_shards × (from + size)` entries, sorts them into the true global order, discards everything outside the requested window, and then runs the **fetch phase** to load `_source` for the surviving `size` hits. So the work is proportional to `from`, not to `size`. `from: 9990, size: 10` on a five-shard index means 50,000 hit entries built, serialized, transferred, and merged so that ten documents can be shown. At `from: 1000000` the same request would be tens of millions of entries — enough to put real pressure on the coordinating node's heap and to trip circuit breakers. Nothing about this gets cheaper with more hardware: adding shards makes it *worse*, because the multiplier is the shard count. ## The cap: index.max_result_window To stop that failure mode, each index carries the dynamic setting `index.max_result_window`, default **10,000**, and Elasticsearch rejects any search where `from + size` exceeds it. The error message names the limit explicitly ("Result window is too large, from + size must be less than or equal to: [10000]") and suggests `search_after`. Note that it is the *sum* that is capped: `from: 9995, size: 10` fails just as `from: 10000, size: 10` does. A companion setting, `index.max_inner_result_window` (default 100), caps the same arithmetic for `inner_hits` and the `top_hits` aggregation, which is why you cannot page deeply inside nested results either. ## Why raising the setting is usually the wrong move `index.max_result_window` is dynamic, so it is tempting to `PUT` a bigger number and move on. That converts a clear 400-level error into an intermittent memory problem: the request now succeeds most of the time and takes down a coordinating node when several deep pages arrive at once, or when someone scripts a crawl of every page. The limit is a guard rail, not a bug. Raise it only for a bounded, known workload — for example an internal tool that occasionally needs page 1,200 of a small index — and raise it modestly. ## What to do instead - **`search_after`** is the supported deep-paging mechanism: sort by something with a unique tiebreaker, and pass the previous page's `sort` values to fetch the next page. Cost per page is constant regardless of depth, because each shard only has to find the first `size` documents *after* a given sort position. Pair it with a **point in time (PIT)** so pages stay consistent while the index is being written to. - **Scroll or a PIT-based export** when the goal is not paging at all but processing every matching document once, offline. - **Rethink the UI.** Almost no human clicks to page 900. Deep paging requests in production are usually a crawler, a broken "export" button, or an infinite-scroll client that never stops. Better facets, better relevance, or an explicit export endpoint removes the requirement instead of engineering around it. ## Common variations of the question Interviewers often follow up with "so what is the cost of `from: 10, size: 10000`?" — that is legal (the sum is 10,010, which fails, so make it 9,990) and cheap on the query phase but expensive on the *fetch* phase, because thousands of `_source` documents must be read and returned. Depth and page width are two different costs; the window bounds their sum.
- Does the cost of deep paging grow with the number of shards?Yes, and that is the uncomfortable part. The coordinating node merges roughly `number_of_shards × (from + size)` hit entries, so splitting an index into more shards makes deep pages proportionally more expensive to merge, even though each individual shard does less matching work. It is one of the few operations where over-sharding hurts directly and visibly.
- Is index.max_result_window a cluster-wide or a per-index setting?Per index, and it is dynamic, so it can be updated with the update-index-settings API without a reindex. It applies to the index being searched; when you search an alias or a wildcard pattern, each concerned index enforces its own value. Setting it in an index template makes new indices inherit the change.
- Why does from: 9990, size: 10 succeed while from: 10000, size: 10 fails?The window caps the sum. 9,990 + 10 is exactly 10,000, which is allowed; 10,000 + 10 is 10,010, which is not. That is also why the last reachable page under the default window is the one ending at hit 10,000 — a detail worth checking when a paging UI mysteriously breaks on the final page rather than the first.
saying these in an interview costs you the question
- Claiming from/size cost is constant because only size hits return
- Recommending max_result_window = 1000000 as the standard fix
- Saying the limit exists to protect the client, not the cluster
- Thinking the cap counts size only, not from + size
- Believing more shards make deep paging cheaper