In Elasticsearch, what does the stats aggregation return, and how does extended_stats extend it?
answer
- one aggregation, several numbers
- saves running four separate aggregations
- the extended variant describes spread
- sigma scales the bounds around the mean
basics
~10 sThe stats aggregation returns count, min, max, avg and sum for a numeric field in a single pass. extended_stats adds sum_of_squares, variance, standard deviation and standard-deviation bounds around the mean.
solid answer
~40 s`stats` is a multi-value metric aggregation: over whatever documents are in scope it returns `count`, `min`, `max`, `avg` and `sum` in one response, reading doc_values once instead of running four separate aggregations. `extended_stats` returns all of those plus `sum_of_squares`, `variance`, `std_deviation` and a `std_deviation_bounds` object holding `upper` and `lower`, computed as the mean plus or minus `sigma` standard deviations (sigma defaults to 2). On Elasticsearch 8.x the response also splits the spread numbers into population and sampling variants — `variance_population` / `variance_sampling` and the matching `std_deviation_*` — while the plain `variance` and `std_deviation` keys are the population forms. Note that `count` counts extracted field values, not documents, and documents with no value in the field are simply skipped.
code
json · 9 lines{
"size": 0,
"aggs": {
"price_stats": { "stats": { "field": "price" } },
"price_spread": {
"extended_stats": { "field": "price", "sigma": 3 }
}
}
}go deeper
Be ready to name the five values stats returns and to say that extended_stats adds variance and standard deviation. Knowing you get them in one request rather than four is the point of the question.
Explain that count counts extracted values rather than documents, that missing values are skipped unless you set the missing parameter, and what sigma does to std_deviation_bounds.
Show judgment about when spread statistics mislead: skewed latency or price distributions need percentiles, not mean plus two sigma. Also handle the empty-bucket case where avg is null but sum is 0.
Own the reporting contract. Decide which summary statistics your platform publishes, whether population or sampling estimators are correct for the audience, and how nulls and multi-valued fields are documented so consumers cannot misread the numbers.
## Where these aggregations sit Elasticsearch aggregations divide into bucket aggregations, which group documents, and metric aggregations, which compute numbers over the documents currently in scope. `stats` and `extended_stats` are *multi-value* metric aggregations: a single aggregation clause produces several numbers in the response. They can run at the top level of a search (over everything the query matched) or nested inside a bucket aggregation, in which case they are recomputed per bucket. ## What stats returns Pointed at a numeric field, `stats` returns five keys: - `count` — how many values were extracted from the field - `min` and `max` — the smallest and largest value seen - `avg` — the arithmetic mean - `sum` — the total ```json { "aggs": { "price_stats": { "stats": { "field": "price" } } } } ``` The practical reason to reach for it is cost: the aggregation walks the field's doc_values once and derives all five numbers, instead of you writing separate `min`, `max`, `avg` and `sum` aggregations that each iterate the same data. ## The count trap `count` is a count of *values*, not of documents. If `price` is an array with three entries in one document, that document contributes three to `count`, and it also contributes three values to `sum` and to the mean. Conversely, a document that has no `price` at all contributes nothing — it is skipped rather than treated as zero. So `stats.count` inside a bucket can be larger *or* smaller than that bucket's `doc_count`, and if you want documents rather than values you should read `doc_count`. Most metric aggregations, including `stats`, accept a `missing` parameter that supplies a substitute value for documents lacking the field, which pulls those documents back into the computation: ```json { "aggs": { "s": { "stats": { "field": "price", "missing": 0 } } } } ``` ## What extended_stats adds `extended_stats` is a superset. Alongside the five `stats` numbers it reports: - `sum_of_squares` — the sum of each value squared, the raw input to the variance - `variance` and `std_deviation` — the spread of the values (population forms) - `std_deviation_bounds` — an object with `upper` and `lower`, equal to `avg ± sigma × std_deviation` On currently shipping 8.x versions the response also carries `variance_population`, `variance_sampling`, `std_deviation_population` and `std_deviation_sampling`, plus population and sampling variants inside `std_deviation_bounds`, so you can pick the estimator that matches your statistical intent. The unqualified `variance` and `std_deviation` keys remain the population figures. The `sigma` parameter controls only the bounds; it defaults to 2, so by default the bounds are two standard deviations either side of the mean, the interval that covers roughly 95% of a normal distribution. Setting `sigma` to 3 widens them. It must be non-negative. ## Empty and missing cases If nothing in scope has a value for the field, `sum` reports 0 while `min`, `max` and `avg` come back as `null`, and `count` is 0. That asymmetry surprises people who expect a uniform null: a sum over no values is legitimately zero, but a mean over no values is undefined. Downstream code that assumes a number for `avg` will break on an empty bucket, so handle the null explicitly. ## Cost and field requirements Both aggregations read doc_values, the columnar per-field storage that Elasticsearch builds by default for numeric and keyword fields. Aggregating a `text` field fails unless fielddata is explicitly enabled, which is a memory hazard and almost never the right answer — map a `keyword` sub-field or a numeric field instead. Because the aggregation is recomputed per bucket, nesting `extended_stats` under a high-cardinality bucket aggregation multiplies the work. ## Choosing between them, and when neither fits Use `stats` when you want the everyday five numbers, `extended_stats` when you genuinely need spread. Two limits are worth naming in an interview. First, there is no median in either aggregation — the median is the 50th percentile and belongs to the `percentiles` aggregation. Second, standard deviation describes a symmetric, roughly normal distribution; latency and price data are usually heavily right-skewed, so a mean-plus-two-sigma bound is close to meaningless for them and percentiles are the honest summary. Reaching for `extended_stats` on a latency field is a common way to produce a confident-looking dashboard number that no one should act on.
- Why is there no median in the stats aggregation, and where do you get one?The median is the 50th percentile, and computing it requires ordering the values rather than accumulating running totals the way count, sum, min, max and the sum of squares do. Elasticsearch exposes it through the `percentiles` aggregation, asking for the 50th percentile, which is approximate rather than exact for large data sets.
- An extended_stats aggregation on a text field fails. What is the fix?Aggregations read doc_values, which analyzed `text` fields do not have. Enabling fielddata on the text field would technically work but loads uninverted terms into heap and is a well-known way to trip circuit breakers. The correct fix is to map the value properly — a numeric field, or a `keyword` multi-field — and aggregate that.
saying these in an interview costs you the question
- Says stats returns the median along with the mean
- Reads stats count as a document count on multi-valued fields
- Thinks documents missing the field count as zero in avg
- Claims extended_stats std_deviation is the sample estimator by default
- Expects avg to return 0 rather than null for an empty bucket