skip to content

Pagination, Sorting & search_after

Deep from/size paging makes every shard build an enormous top-N list, which is why index.max_result_window caps it at 10,000 by default. Interviewers expect search_after with a point-in-time for deep pages, and the scroll API only for full exports.

part ofElasticsearchoverview, primer and where to startread it →
on this pageshow

questions

6

Why does Elasticsearch reject a search with from=10000 and size=10?

level: juniorimportance: must knowfreq 70%

answer

  1. Skipping is not free when results are spread over shards
  2. Each shard must return its own top slice
  3. The coordinating node merges shards × (from + size)
  4. One index setting caps the sum at 10,000
  5. Its name ends in max_result_window

basics

~20 s

Because from + size exceeds index.max_result_window, which defaults to 10,000. Deep paging makes every shard build a top-(from+size) list that the coordinating node merges and then throws almost all of away, so Elasticsearch caps the depth instead.

solid answer

~40 s

`from` skips hits and `size` returns the next N, but that skipping is not free in a distributed index. To answer page 1000, **every** shard must produce its own top `from + size` hits, and the coordinating node merges all of them — with 5 shards and `from: 9990, size: 10` that is 50,000 hits sorted and merged so 10 can be returned. Cost grows linearly with depth on both memory and network. Elasticsearch therefore refuses the request once `from + size` passes `index.max_result_window`, which defaults to 10,000, with an error saying the result window is too large. Raising the setting per index is possible but just moves the cliff; the real fix is `search_after` with a point in time for deep paging, or a scroll/PIT export if you actually need every hit.

code

json · 6 lines
json
// GET /products/_search  -> rejected: from + size = 10010
{
  "from": 10000,
  "size": 10,
  "sort": [ { "price": "asc" } ]
}

go deeper

for a junior

Recall that from skips hits, size returns them, and that from + size cannot exceed 10,000 by default. Be able to name search_after as the alternative rather than reaching for the settings API.

for a middle

Explain the query-then-fetch mechanics: every shard returns its own top from+size list and the coordinating node merges shards × (from + size) entries. Tie the cap directly to that arithmetic.

for a senior

Show judgment about when raising index.max_result_window is defensible and when it converts a clean error into an intermittent heap incident. Expect to diagnose which client is actually requesting page 900.

for a principal

Own the product-level answer: deep pagination is usually a symptom of a missing export path or weak relevance. Decide whether the system offers paging, cursors, or bulk export, and enforce that contract at the API edge.

## What from and size do In the Elasticsearch search API, `size` is the number of hits to return and `from` is the number of leading hits to skip. `from: 20, size: 10` gives you the third page of a ten-per-page list. Both are pure result-set slicing — they do not change which documents match, only which slice of the ordered match list is returned to the client. `from` defaults to 0 and `size` defaults to 10. ## Why depth is expensive in a distributed index A search runs in two phases. In the **query phase** the coordinating node broadcasts the request to one copy of every shard. There is no global ordering available to any single shard, so each shard cannot know which of *its* documents fall into the caller's requested page — it must assume the worst and return its own best `from + size` results (document ids plus sort values) to the coordinating node. The coordinating node builds a priority queue of `number_of_shards × (from + size)` entries, sorts them into the true global order, discards everything outside the requested window, and then runs the **fetch phase** to load `_source` for the surviving `size` hits. So the work is proportional to `from`, not to `size`. `from: 9990, size: 10` on a five-shard index means 50,000 hit entries built, serialized, transferred, and merged so that ten documents can be shown. At `from: 1000000` the same request would be tens of millions of entries — enough to put real pressure on the coordinating node's heap and to trip circuit breakers. Nothing about this gets cheaper with more hardware: adding shards makes it *worse*, because the multiplier is the shard count. ## The cap: index.max_result_window To stop that failure mode, each index carries the dynamic setting `index.max_result_window`, default **10,000**, and Elasticsearch rejects any search where `from + size` exceeds it. The error message names the limit explicitly ("Result window is too large, from + size must be less than or equal to: [10000]") and suggests `search_after`. Note that it is the *sum* that is capped: `from: 9995, size: 10` fails just as `from: 10000, size: 10` does. A companion setting, `index.max_inner_result_window` (default 100), caps the same arithmetic for `inner_hits` and the `top_hits` aggregation, which is why you cannot page deeply inside nested results either. ## Why raising the setting is usually the wrong move `index.max_result_window` is dynamic, so it is tempting to `PUT` a bigger number and move on. That converts a clear 400-level error into an intermittent memory problem: the request now succeeds most of the time and takes down a coordinating node when several deep pages arrive at once, or when someone scripts a crawl of every page. The limit is a guard rail, not a bug. Raise it only for a bounded, known workload — for example an internal tool that occasionally needs page 1,200 of a small index — and raise it modestly. ## What to do instead - **`search_after`** is the supported deep-paging mechanism: sort by something with a unique tiebreaker, and pass the previous page's `sort` values to fetch the next page. Cost per page is constant regardless of depth, because each shard only has to find the first `size` documents *after* a given sort position. Pair it with a **point in time (PIT)** so pages stay consistent while the index is being written to. - **Scroll or a PIT-based export** when the goal is not paging at all but processing every matching document once, offline. - **Rethink the UI.** Almost no human clicks to page 900. Deep paging requests in production are usually a crawler, a broken "export" button, or an infinite-scroll client that never stops. Better facets, better relevance, or an explicit export endpoint removes the requirement instead of engineering around it. ## Common variations of the question Interviewers often follow up with "so what is the cost of `from: 10, size: 10000`?" — that is legal (the sum is 10,010, which fails, so make it 9,990) and cheap on the query phase but expensive on the *fetch* phase, because thousands of `_source` documents must be read and returned. Depth and page width are two different costs; the window bounds their sum.

  • Does the cost of deep paging grow with the number of shards?
    Yes, and that is the uncomfortable part. The coordinating node merges roughly `number_of_shards × (from + size)` hit entries, so splitting an index into more shards makes deep pages proportionally more expensive to merge, even though each individual shard does less matching work. It is one of the few operations where over-sharding hurts directly and visibly.
  • Is index.max_result_window a cluster-wide or a per-index setting?
    Per index, and it is dynamic, so it can be updated with the update-index-settings API without a reindex. It applies to the index being searched; when you search an alias or a wildcard pattern, each concerned index enforces its own value. Setting it in an index template makes new indices inherit the change.
  • Why does from: 9990, size: 10 succeed while from: 10000, size: 10 fails?
    The window caps the sum. 9,990 + 10 is exactly 10,000, which is allowed; 10,000 + 10 is 10,010, which is not. That is also why the last reachable page under the default window is the one ending at hit 10,000 — a detail worth checking when a paging UI mysteriously breaks on the final page rather than the first.

saying these in an interview costs you the question

  • Claiming from/size cost is constant because only size hits return
  • Recommending max_result_window = 1000000 as the standard fix
  • Saying the limit exists to protect the client, not the cluster
  • Thinking the cap counts size only, not from + size
  • Believing more shards make deep paging cheaper

context

open as a page

How do you page beyond 10,000 hits in Elasticsearch using search_after and a point in time?

level: middleimportance: must knowfreq 70%

basics

~20 s

Open a point in time, sort by a field plus a unique tiebreaker, then send each next page with search_after set to the previous page's last hit sort array. Cost per page stays constant, and the PIT keeps pages consistent while indexing continues.

open as a page

Why does sorting an Elasticsearch search on a text field fail, and what do you sort on instead?

level: middleimportance: should knowfreq 62%

basics

~20 s

Sorting needs one value per document read from a columnar doc_values structure, and analyzed text fields have doc_values disabled, so the request errors telling you fielddata is off. Sort on a keyword sub-field such as title.keyword, which has doc_values by default.

open as a page

Why does hits.total report 10000 with relation gte, and what does track_total_hits change?

level: middleimportance: should knowfreq 58%

basics

~20 s

Elasticsearch stops counting matches accurately at 10,000 by default so top-k queries can skip non-competitive documents. A relation of gte means at least that many matched. track_total_hits: true forces an exact count, false skips counting entirely, and an integer sets a different threshold.

open as a page

When should you still use Elasticsearch's scroll API instead of search_after with a point in time?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Almost never for new code. Scroll is a stateful snapshot cursor built for exporting an entire result set once, and Elasticsearch no longer recommends it for deep pagination; search_after with a point in time covers the same ground statelessly. Scroll survives mainly in existing batch jobs.

open as a page

An Elasticsearch search is fast until it returns 100 large documents per page — how do you cut the fetch cost?

level: seniorimportance: should knowfreq 45%

basics

~20 s

That cost is the fetch phase reading and decompressing _source for every returned hit. Return fewer hits per page, use _source filtering to shrink the payload, read small values from doc_values with docvalue_fields, or map the few needed fields with store true so _source is never touched.

open as a page