When would you set index: false or doc_values: false on an Elasticsearch field?
answer
- two different access paths, two switches
- one serves queries, one serves aggregations
- neither touches the stored document
- think request IDs versus latency metrics
- disk savings on the very largest fields
basics
~20 sSet index: false on fields you only display or aggregate on, never search by; set doc_values: false on fields you only search or retrieve, never sort, aggregate or script on. Each removes one index structure to save disk and indexing time.
solid answer
~50 sThe two parameters switch off independent structures. `index: false` drops the inverted index for a field, so it stops being efficiently searchable; `doc_values: false` drops the columnar per-document values, so the field can no longer be sorted on, aggregated on, or read cheaply from scripts. Both default to on for `keyword`, numeric, date, boolean and `ip` fields, and both leave `_source` alone, so the value is still returned in hits either way. Use `index: false` for values you only ever slice by — a metric you aggregate but never filter on — and `doc_values: false` for high-cardinality identifiers you filter by and never facet on. On huge indexes these choices are real money: doc values on a high-cardinality keyword can rival the postings in size. Assume Elasticsearch 8.x; the safe framing is that turning either off means giving up that access path.
code
json · 7 lines{
"properties": {
"response_time_ms": { "type": "long", "index": false },
"trace_id": { "type": "keyword", "doc_values": false },
"status": { "type": "keyword" }
}
}go deeper
Know that both parameters default to on for keyword, numeric and date fields, and that turning either off removes a capability rather than deleting data — the value still comes back in the search hit.
Explain which structure each switch controls: the inverted index behind queries versus the columnar doc values behind sorting, aggregations and scripts, and give one field that plausibly needs only one of them.
Show you would drive it from measured query patterns on the largest indexes, size the saving, and account for the asymmetry that turning a structure off wrongly costs a reindex to undo. Note that neither can be changed in place.
Own it as a storage-cost programme: encode the decisions in index templates so every future index inherits them, validate on one index before it propagates, and weigh the saving against losing ad-hoc investigation paths for on-call engineers.
## Two structures, two switches For a typical non-analyzed field Elasticsearch writes two things per document, for two different access patterns: 1. **The inverted index** — term to document-list. Answers "which documents contain X?" This is what `index` controls. 2. **Doc values** — a columnar, per-document store, held on disk and read through the filesystem cache. Answers "what value does document 37 have?" This is what `doc_values` controls, and it is what sorting, aggregations, and `doc['field']` script access read. Both are on by default for `keyword`, numeric, `date`, `boolean` and `ip` fields. Neither has anything to do with `_source`, the stored JSON document, which is why a field with both disabled is still returned in search hits exactly as it was submitted. ## index: false Setting `index: false` tells Elasticsearch not to write the field into the inverted index. The classic use is a numeric metric in a monitoring index that you only ever aggregate: you compute averages and percentiles over `response_time`, but you never filter `response_time > 500`. The postings for it are pure overhead — dropping them shrinks the index and speeds ingest, while the doc values that aggregations need remain. What you give up is efficient search on the field. Historically a query against a non-indexed field was refused outright. In current 8.x versions some types with doc values can still be matched by scanning those doc values, which is dramatically slower than an index lookup and does not scale — so the honest way to state the rule is: once you set `index: false`, treat the field as not searchable, and if a filter on it later becomes a requirement you are looking at a mapping change and a reindex. Other good candidates are large payload-ish fields you keep only for display, and fields duplicated from another that *is* indexed. ## doc_values: false Setting `doc_values: false` removes the columnar store. The field stays searchable, but sorting on it, aggregating on it and reading it in a script all stop working — those requests fail rather than degrading. The canonical candidate is a very high-cardinality `keyword` you filter by and nothing else: a request ID, a trace ID, a session token, a full URL. You look documents up by it; you would never build a `terms` aggregation over millions of unique values because the result is meaningless. Doc values for such a field can be a substantial share of the index — for high-cardinality keywords they are often comparable to or larger than the postings — so switching them off is one of the few genuinely large storage wins available in a mapping. Note that `text` fields never have doc values in the first place; the parameter is not applicable there. ## The interaction with store and _source A third parameter, `store`, controls whether the field's value is kept separately as a stored field. It defaults to `false` because `_source` already holds every value and Elasticsearch extracts requested fields from it. `store: true` earns its keep only in narrow cases — for example a very large document where you routinely fetch one small field and want to avoid decompressing the whole `_source`. The important consequence for this discussion: disabling both `index` and `doc_values` does **not** make the data disappear. The field is still in `_source`, still returned, still reindexable later. What you have removed is only the two query-time access paths — which is exactly why the change is reversible in principle but expensive in practice, since restoring them means rewriting every document. ## How to decide, and how to be sure Drive it from measured query patterns, not intuition. For each large field ask two questions: - Does any query ever *filter or match* on this field? If genuinely no, `index: false` is on the table. - Does any request ever *sort, aggregate or script* on this field? If genuinely no, `doc_values: false` is on the table. "Genuinely no" is the hard part. Ad-hoc investigation queries, a future dashboard, or a support engineer grepping by trace ID all count as yes. The failure mode is asymmetric: leaving a structure on costs storage continuously, while turning one off wrongly costs a full reindex to undo. That asymmetry is why these switches are worth reaching for on the biggest fields of the biggest indexes — time-series and log data, where a mapping applies to billions of documents — and not worth the risk on a modest index where the saving is megabytes. Because neither parameter can be flipped on an existing field, the practical workflow is to set them in an index template so every new backing index or daily index picks up the decision, and to validate the choice on one index's worth of data before it propagates everywhere.
- If both index and doc_values are disabled on a field, is the value lost?No. Both parameters only control index structures; the value is still stored in _source and returned in search hits exactly as submitted. What is gone is the ability to search, sort or aggregate on it — and because neither parameter can be changed on an existing field, restoring those paths means reindexing into a new mapping.
- How does store: true differ from doc_values: true?They serve different readers. Doc values are columnar and exist so sorts, aggregations and scripts can read one field across many documents cheaply. A stored field is row-oriented retrieval for one document, and it is usually redundant because _source already holds every value; it pays off mainly when you fetch a small field out of a very large document.
- Can you toggle these parameters on an existing index?No — like a field's type, they are fixed once the field exists, because the structures are written at index time and cannot be materialized retroactively. Changing them means creating an index with the new mapping and reindexing, so the decision usually belongs in an index template applied to future indices.
saying these in an interview costs you the question
- Thinks index: false removes the value from search results
- Says doc_values: false only makes aggregations slower
- Believes these parameters can be flipped on a live field
- Confuses doc_values with the stored _source document
- Applies index: false to a field support engineers filter on