skip to content

Why does sorting an Elasticsearch search on a text field fail, and what do you sort on instead?

level: middleimportance: should knowfreq 62%

answer

  1. Sorting asks document to value, not term to documents
  2. Analyzed fields have no single value to compare
  3. A columnar on-disk structure backs sort and aggregations
  4. The error mentions an in-heap alternative you should refuse
  5. Dynamic mapping already gives you a sub-field for this

basics

~20 s

Sorting needs one value per document read from a columnar doc_values structure, and analyzed text fields have doc_values disabled, so the request errors telling you fielddata is off. Sort on a keyword sub-field such as title.keyword, which has doc_values by default.

solid answer

~50 s

To sort, Elasticsearch must read a single comparable value per document. That comes from **doc_values**, a columnar, on-disk structure written at index time. A `text` field is analyzed into many tokens, so there is no one value to sort by; doc_values are disabled on it and the request fails with an error saying fielddata is disabled and suggesting a `keyword` field instead. The standard fix is a multi-field: map `title` as `text` for searching with a `title.keyword` sub-field for sorting and aggregating. Turning on `fielddata: true` technically works by building the sort structure in heap at query time, and it is the classic route to an out-of-memory cluster — avoid it. Useful sort options are `missing` for documents without a value, `mode` for multi-valued fields, `unmapped_type` when searching across indices, and `_doc` when order does not matter. Note that sorting on a field means scores are not computed unless you set `track_scores`.

code

json · 13 lines
json
// PUT /articles
{
  "mappings": {
    "properties": {
      "title": {
        "type": "text",
        "fields": {
          "keyword": { "type": "keyword", "ignore_above": 256 }
        }
      }
    }
  }
}

go deeper

for a junior

Recall that you sort on keyword, numeric or date fields, not on analyzed text, and that the usual fix is the .keyword sub-field a dynamic mapping already created.

for a middle

Explain doc_values as a columnar document-to-value structure, why an analyzed field has no single value, and what missing, mode and unmapped_type control on a sort clause.

for a senior

Be ready to diagnose the field-level surprises in production: ignore_above silently dropping values, case-sensitive keyword ordering, sorts spanning indices with divergent mappings, and the memory story behind fielddata.

for a principal

Own the mapping strategy that makes sorting cheap at scale: which fields get keyword sub-fields at all, whether index sorting is worth committing to at index creation, and how mapping choices bound query cost.

## What sorting needs Ranking by relevance is what the inverted index is built for: postings lists tell you which documents contain a term. Sorting by a *field value* is the opposite access pattern — given a document, what is its value? That requires a **document-to-value** structure, and in Lucene that structure is `doc_values`: a columnar, compressed, on-disk column per field, written when the segment is written and read via the filesystem cache. Doc values back sorting, aggregating, and script access to field values. Doc values are enabled by default on `keyword`, numeric, date, boolean, `ip` and similar non-analyzed types. They are **not** available on `text`. ## Why text refuses A `text` field is passed through an analyzer, so a title like "The Quick Brown Fox" is stored as the terms `the`, `quick`, `brown`, `fox` — lowercased, possibly stemmed, order flattened. There is no single value that could serve as a sort key, and the original string is not kept in the index at all (only in `_source`, which is not readable per document during the sort phase without enormous cost). Elasticsearch therefore rejects the sort with a message along the lines of: *Fielddata is disabled on text fields by default. Set fielddata=true on [title] in order to load fielddata in memory by uninverting the inverted index. Note that this can however use significant memory. Alternatively use a keyword field instead.* That message contains both the trap and the fix. ## The fix: a keyword multi-field Map the field twice: ``` "title": { "type": "text", "fields": { "keyword": { "type": "keyword", "ignore_above": 256 } } } ``` Queries use `title`; sorts and aggregations use `title.keyword`. This is exactly what dynamic mapping produces for a detected string, which is why `field.keyword` appears everywhere in Kibana. Two consequences to be able to state: - The `keyword` sub-field stores the **whole, unanalyzed** string, so sorting is by raw byte order: uppercase sorts before lowercase, and "Zebra" precedes "apple". A `normalizer` (lowercase, maybe `asciifolding`) on the keyword field gives case-insensitive, accent-insensitive sorting. - `ignore_above` means strings longer than the limit are not indexed into the keyword sub-field at all, so those documents sort as *missing* rather than by their long value — a genuinely surprising bug the first time you meet it. ## The fielddata trap Setting `fielddata: true` on a text field makes the sort work by **uninverting** the index at query time: reading every term's postings and building a document-to-terms map in the JVM heap. It is unbounded relative to field cardinality, it is rebuilt per segment, and it is protected only by circuit breakers. It exists for narrow cases like aggregating on analyzed tokens (significant terms in a text body), never as a way to make a sort compile. "Just set fielddata true" is a red-flag answer. ## Sort options worth knowing - **`missing`** — `"_last"` (default for ascending in most cases), `"_first"`, or a literal substitute value. Documents lacking the field have no doc value at all, and where they land is a product decision, not a detail. - **`unmapped_type`** — when a search spans several indices and the sort field is missing from some mapping, the request fails unless you declare a type to assume. Common on rolling time-based indices where a field was added later. - **`mode`** — a field can be multi-valued. `min`, `max`, `sum`, `avg`, `median` pick which of the array's values represents the document. - **Nested sort** — sorting by a value inside `nested` objects requires a `nested` sort clause with a path and usually a filter, otherwise the wrong child's value is used. - **`_doc`** — internal storage order, the cheapest sort; use it when order is irrelevant, such as an export walk. - **`_score`** — the default. Once you sort by a field, scores are not computed at all and `_score` comes back null; set `track_scores: true` if you still need them, at the cost of scoring every hit. - **Index sorting** (`index.sort.field` at index creation) pre-sorts documents inside each segment, which can let a matching-and-sorting query terminate early — a real optimization for time-ordered data, and one that has to be decided before the index exists. ## Cost notes Sorting reads doc values for candidate documents, so it is cheap per document but touches the filesystem cache; a high-cardinality string sort over a huge index is heavier than a numeric one because the values are larger. Sorting also defeats the block-skipping optimizations that pure relevance queries enjoy, unless index sorting lets the collector stop early. On a paging endpoint, the practical rule is: sort on a numeric or date field with doc values plus a unique tiebreaker, and you are on the fast path for `search_after` as well.

  • Why does sorting on title.keyword put Zebra before apple?
    A keyword field stores the raw, unanalyzed string and sorts by its byte order, where uppercase letters precede lowercase ones. Apply a `normalizer` with a lowercase filter to the keyword sub-field for human-friendly ordering; the normalizer is applied at index time, so changing it requires reindexing.
  • What happens to documents whose value exceeds ignore_above on the keyword sub-field?
    The value is not indexed into that sub-field, so those documents have no doc value there. They are treated as missing by both sorting and aggregation, landing wherever the sort's `missing` option puts them and vanishing from terms buckets. The document itself is still searchable through the analyzed text field.
  • Why is _score null when you sort by a field?
    Sorting by a field lets the collector skip relevance computation entirely, so Elasticsearch does not calculate scores and returns null. If you need both the field order and the score — to display it, or to re-rank client-side — set `track_scores: true`, accepting that every matching hit must now be scored.

saying these in an interview costs you the question

  • Setting fielddata true on a text field to make a sort work
  • Believing sorting reads values out of _source
  • Expecting a keyword sort to be case-insensitive by default
  • Forgetting a tiebreaker so ties order differently on each run
  • Assuming missing values always sort last regardless of direction

context