skip to content

An Elasticsearch search is fast until it returns 100 large documents per page — how do you cut the fetch cost?

level: seniorimportance: should knowfreq 45%

answer

  1. Two phases: find the hits, then load them
  2. Cost tracks returned documents, not matched ones
  3. The document body lives in compressed blocks
  4. Filtering the body saves bandwidth, not the disk read
  5. Columnar values can be read without the body at all

basics

~20 s

That cost is the fetch phase reading and decompressing _source for every returned hit. Return fewer hits per page, use _source filtering to shrink the payload, read small values from doc_values with docvalue_fields, or map the few needed fields with store true so _source is never touched.

solid answer

~50 s

A search runs query-then-fetch. The query phase finds and ranks hit ids; the **fetch phase** then loads each returned document's `_source` from the stored-fields file, decompressing the block it lives in, and ships it back. That cost scales with `size` and with document size, which is why latency jumps only when the page gets wide or the documents get fat. Options, roughly in order: reduce `size`; apply `_source` filtering (`"_source": { "includes": [...] }`) — this cuts network and client parsing but the whole `_source` is still read and decompressed on the data node; use `docvalue_fields` for values already in doc_values, which skips `_source` entirely; or map the handful of display fields with `store: true` and request them via `stored_fields`, so only those small stored fields are read. Also check `index.codec` — `best_compression` trades fetch CPU for disk. Disabling `_source` outright is usually a mistake: it breaks reindex, update-by-query and highlighting.

code

json · 5 lines
json
// before: 100 large documents fetched in full
{
  "size": 100,
  "query": { "match": { "body": "outage" } }
}

go deeper

for a junior

Know that a search response carries each document's _source, and that you can ask for a subset of fields with _source includes to shrink the payload.

for a middle

Explain the query-then-fetch split and why fetch cost scales with size and document size. Know that docvalue_fields and stored_fields exist as alternatives to reading _source.

for a senior

Diagnose it properly: separate fetch from query cost, weigh compression codec, highlighting and filesystem cache, and choose between smaller pages, doc-value reads and genuinely stored fields.

for a principal

Own the document-shape decision: what belongs in the search index at all versus in the system of record, whether the search API returns documents or references, and how that bounds fetch cost as the corpus grows.

## Where the time goes An Elasticsearch search is two phases. In the **query phase** every involved shard matches, scores and sorts, returning only document ids and sort values to the coordinating node. Nothing is read from the documents themselves. In the **fetch phase** the coordinating node asks the relevant shards for the actual content of the surviving hits: `_source`, highlights, `fields`, and anything else the response needs. That split explains the symptom. Query-phase cost tracks the number of *matching* documents; fetch-phase cost tracks the number of *returned* documents times their size. A query that is instant at `size: 10` and sluggish at `size: 100` with 200 KB documents is not a matching problem at all, and no amount of query tuning will help it. ## What _source really is `_source` is the original JSON body, kept in Lucene's stored fields. Stored fields are compressed in **blocks** of documents (LZ4 by default; `index.codec: best_compression` switches to DEFLATE for smaller disk at higher CPU). Retrieving one document means locating and decompressing the block it sits in. Retrieving 100 documents scattered across segments means many such decompressions. The crucial nuance, and the one interviewers listen for: **`_source` filtering happens after the source is loaded and parsed.** `"_source": { "includes": ["id", "title"] }` reduces the bytes sent over the wire and the JSON the client has to parse — often a large win end to end — but the data node still read and decompressed the entire document. If the field-extraction cost on the data node is the bottleneck, filtering will not move it. ## The four levers **1. Return fewer hits.** The cheapest fix. A UI showing 100 results per page is usually showing 90 nobody looks at. Combine a smaller page with `search_after` so deep paging stays cheap too. **2. `_source` filtering.** `"_source": false` to return no body at all (ids and sort values only), a list of fields, or `includes`/`excludes` with wildcards. Excellent when the document is large but the list view needs three fields. Saves bandwidth and client CPU. **3. `docvalue_fields`.** Reads values from the columnar doc_values that already back sorting and aggregations, without touching `_source` at all. Works for fields with doc values (keyword, numeric, date, boolean, ip); does not work for `text`. Values come back in the doc_values representation, which for dates means you may want to pass a `format`. For a list view rendering an id, a price and a timestamp, this can eliminate the stored-fields read entirely. **4. `stored_fields`.** Mapping a field with `store: true` writes it as its own stored field, separate from `_source`. Requesting it via `stored_fields` reads only that small field. This is the right shape when a document has one huge field (a scraped page body, a base64 blob) and a few tiny display fields: the tiny ones are read directly and the huge one is never decompressed. The `fields` parameter is a fifth, more convenient option: it returns values formatted according to the mapping, works with runtime fields, and handles multi-fields consistently, drawing from `_source` and doc values as appropriate. ## What not to do **Do not disable `_source`.** `"_source": { "enabled": false }` saves real disk, and then: reindex cannot rebuild the index from itself, `update` and `update_by_query` stop working, highlighting on non-stored fields breaks, and debugging becomes guesswork because you can no longer see what you indexed. The usual middle ground is `excludes` on the mapping's `_source` for genuinely useless giant fields — accepting that those excluded fields will not survive a reindex. **Do not confuse this with deep paging.** `from: 9990, size: 10` is a *query-phase* merge problem bounded by `index.max_result_window`; `from: 0, size: 500` on fat documents is a *fetch-phase* problem with no such guard. They are diagnosed differently: the first shows coordinating-node merge cost, the second shows time in fetch on the data nodes. The search profile API and the node-level fetch metrics separate them. ## Diagnosing it Ask for the split. If `took` grows roughly linearly with `size` while the match count is unchanged, it is fetch. Look at whether documents are large, whether `best_compression` is set, whether highlighting is on (highlighting re-analyzes field content per hit and is frequently the real culprit hiding behind "the fetch is slow"), and whether the filesystem cache is large enough to hold the stored-fields files being touched. Fixing it is usually a product conversation — narrower list rows — dressed as a performance one.

  • Does _source filtering reduce the work the data node does?
    Only partly. The stored field is still located, decompressed and parsed before the include and exclude rules are applied, so the disk and CPU cost on the data node largely remains; what you save is serialization, network bytes and client-side parsing. To avoid the read itself you need docvalue_fields or genuinely stored fields.
  • When is store: true on individual fields worth it?
    When documents contain one very large field and the response needs only a few small ones. Storing those small fields separately lets the fetch read them directly instead of decompressing a block containing the huge body. It costs a little extra disk and only pays off when the size asymmetry is real.
  • How would you confirm the latency is in the fetch phase rather than the query phase?
    Hold the query constant and vary size: if `took` scales with the number of returned hits while the match count is unchanged, it is fetch. The search profile API and the node-level indices search fetch metrics confirm it directly. Also check whether highlighting is enabled, since it does per-hit work in the same phase.
  • Why is disabling _source rarely worth the disk it saves?
    Reindexing an index from itself, update and update_by_query, and highlighting on non-stored fields all read `_source`. Without it, a mapping change means re-ingesting from the system of record, and you can no longer inspect what was actually indexed. Excluding one oversized field from `_source` is the safer compromise.

saying these in an interview costs you the question

  • Claiming _source filtering avoids reading the document on the data node
  • Disabling _source to save disk without weighing reindex and update
  • Blaming the query phase when only the page width changed
  • Expecting docvalue_fields to return analyzed text field values
  • Treating the result window as a guard against wide pages

context