skip to content

How do you page beyond 10,000 hits in Elasticsearch using search_after and a point in time?

level: middleimportance: must knowfreq 70%

answer

  1. The cursor is the previous hit's sort values
  2. Cost per page no longer grows with depth
  3. Ties at a page edge duplicate or drop documents
  4. A frozen view keeps pages consistent while writes continue
  5. One implicit sort field starts with an underscore and shard

basics

~20 s

Open a point in time, sort by a field plus a unique tiebreaker, then send each next page with search_after set to the previous page's last hit sort array. Cost per page stays constant, and the PIT keeps pages consistent while indexing continues.

solid answer

~50 s

`search_after` replaces offset paging with a cursor made of sort values. You open a point in time (`POST /index/_pit?keep_alive=1m`), then search with the returned `pit.id` in the body and **no index in the path**, sorting on something meaningful plus a tiebreaker that is unique per document. Each response's last hit carries a `sort` array; you pass that array as `search_after` on the next request. Because each shard seeks directly to that sort position, per-page cost is constant no matter how deep you go, and `index.max_result_window` does not apply. The tiebreaker matters: without a total order, documents sharing a sort value can be skipped or repeated at page boundaries. With a PIT, Elasticsearch adds `_shard_doc` as an implicit tiebreaker if you did not supply one. `from` is not allowed alongside `search_after`, you cannot jump to an arbitrary page, and you should close the PIT when done.

code

bash · 8 lines
bash
# open a point in time over the target index
curl -X POST "localhost:9200/logs-2026/_pit?keep_alive=1m"
# -> { "id": "46ToAwMDaWR5B..." }

# close it when the walk is finished
curl -X DELETE "localhost:9200/_pit" \
  -H 'Content-Type: application/json' \
  -d '{ "id": "46ToAwMDaWR5B..." }'

go deeper

for a junior

Know the shape: sort, read the last hit's sort array, send it as search_after next time. Recall that this is the supported way past the 10,000-hit window.

for a middle

Explain why per-page cost stays constant, why a unique tiebreaker is mandatory, and what a point in time freezes. Be able to write the request bodies from memory.

for a senior

Discuss operating it: keep_alive sizing, pinned segments and file handles, closing contexts, cursor tokens in a public API, and what to do when a client resumes with an expired PIT.

for a principal

Own the API contract. Decide whether the product exposes page numbers or opaque cursors, what a cursor encodes, how long it stays valid, and how bulk consumers are served without competing with interactive search.

## The problem search_after solves Offset paging (`from` + `size`) forces every shard to rebuild a top-`from + size` list for each page, so cost climbs with depth and Elasticsearch caps it at `index.max_result_window` (10,000 by default). `search_after` inverts the idea: instead of saying "skip 10,000 hits", you say "give me the hits that sort *after this exact position*". A shard can seek to that position and collect the next `size` documents, so page 1 and page 10,000 cost the same. ## Sorting and the tiebreaker `search_after` requires an explicit `sort`. The cursor is literally the `sort` array Elasticsearch returns on each hit, so the sort must impose a **total order**: if two documents share the sort value at a page boundary, the engine cannot tell which side of the boundary they belong on, and you get duplicated or silently skipped documents. The fix is a final tiebreaker field that is unique per document — an id with `doc_values`, a monotonically increasing sequence, or, in a PIT search, the built-in `_shard_doc` field. When a PIT is used and no tiebreaker is present, Elasticsearch appends `_shard_doc` automatically. ``` "sort": [ { "@timestamp": "desc" }, { "_shard_doc": "asc" } ] ``` ## Point in time: why it is the other half Without a PIT, `search_after` still works, but every page is executed against whatever the index looks like at that moment. New documents landing between page 3 and page 4 shift the sequence, so a user can see the same document twice or miss one entirely — the same drift that plagues offset paging over a live feed. A PIT freezes a view of the shards' segments. You open it, receive an id, and pass that id in the body of each subsequent search. Because the PIT identifies the target indices, the search request is sent to `/_search` with no index in the path. Each response returns a `pit_id` — possibly a refreshed one — and the well-behaved client always sends back the most recent id. When the walk finishes you `DELETE /_pit` with that id. `keep_alive` bounds how long the PIT survives between requests; it is extended by each search that uses it. Keeping a PIT open is not free: it pins the segments it references so they cannot be deleted after a merge, which costs disk space and file handles on the data nodes. Use short keep-alives that you extend as you go, rather than a one-hour window "to be safe", and always close explicitly instead of waiting for expiry. ## What the API will not let you do - **No `from`.** A `search_after` request must not carry a `from` parameter; the cursor *is* the position. - **No random access.** You can only step forward from a page you have already seen. There is no "jump to page 500", and going backwards means either keeping the previous cursors client-side or reversing the sort order and re-walking. - **No changing the sort mid-walk.** The cursor's meaning is defined by the sort; change the sort and the previous `sort` array is meaningless. The same applies to changing the query — the safe contract is that a cursor is only valid for the exact query, sort, and PIT that produced it. ## Slicing for parallel consumption When the goal is to consume everything as fast as possible rather than to serve a UI, a PIT search can be sliced: each worker issues the same query with a different `slice` id and walks its own `search_after` sequence. That gives parallel, non-overlapping coverage of one consistent snapshot — the modern replacement for sliced scroll. ## Designing the API around it Because `search_after` cursors are forward-only, an HTTP API built on them should expose an **opaque cursor token** ("next page" link) rather than a page number. Encoding the sort values, the PIT id, and a hash of the query into the token keeps clients from mixing a cursor with a different filter or sort — a mismatch that would otherwise return nonsense rather than an error. Decide up front how long a cursor stays valid, because it can outlive its PIT: a client resuming an hour later must be told to start over rather than served a stale or expired context. ## Rules of thumb Use `from`/`size` for the first handful of pages of a human-facing UI, `search_after` with a PIT for deep or unbounded paging and for infinite scroll, and sliced PIT walks for exports. If the interviewer asks what to do about the 10,000 limit and the answer is "raise the setting", that is the wrong half of the tradeoff.

  • What happens if your search_after sort has no unique tiebreaker?
    Documents that share the sort value sit ambiguously on the page boundary, so the next page can repeat some of them and skip others. The symptom is a paging UI where totals never add up and a few records are invisible. Adding a unique final sort key — a document id with doc_values, or `_shard_doc` under a PIT — restores a total order.
  • Do you still need a point in time if the index is read-only?
    No. A PIT exists to freeze the view against concurrent indexing and merging. On a static index, `search_after` alone gives stable pages and avoids the cost of pinning segments. You do lose the automatic `_shard_doc` tiebreaker, so supply your own unique sort field explicitly.
  • How do you page backwards with search_after?
    There is no backward cursor. Either the client keeps the cursor of each page it has visited and re-issues that request, or you reverse the sort direction, walk forward with the current page's first-hit sort values, and reverse the returned hits client-side. Most UIs simply cache prior cursors.
  • Why must the pit id from each response be reused instead of the original one?
    A search can return an updated `pit_id`, for example when shards change, and the newest id is the one guaranteed to address the current context. Always send back the id from the latest response and close that id at the end; reusing a stale one risks a context-missing error mid-walk.

saying these in an interview costs you the question

  • Passing from together with search_after
  • Sorting on a non-unique field and calling it a cursor
  • Believing search_after needs max_result_window raised
  • Thinking a PIT is free to keep open for hours
  • Claiming search_after can jump to an arbitrary page number

context