Why does an Elasticsearch value_count aggregation return more than the bucket's doc_count?
answer
- documents and values are not the same thing
- arrays are not a separate mapping type
- absent fields contribute nothing at all
- one parameter pulls empty documents back in
basics
~10 svalue_count counts extracted field values, not documents. A multi-valued field contributes one count per array element, so a document with three tags adds three, while a document missing the field adds nothing.
solid answer
~40 s`doc_count` on a bucket counts documents; `value_count` counts the values a metric aggregation actually extracted from the named field. Those diverge in two directions. A multi-valued field — an array in the source, which Elasticsearch indexes as several values on one document — contributes one count per element, pushing `value_count` above `doc_count`. A document that has no value for the field is skipped altogether, pulling `value_count` below `doc_count`. The same rule governs `sum`, `avg` and `stats.count`, so an average over an array field is an average over values, not over documents. If you want absent documents included, set the `missing` parameter to substitute a value for them. If you want documents rather than values, read `doc_count` or count a field guaranteed to be single-valued.
code
json · 4 lines// three indexed documents
{ "id": 1, "tags": ["search", "lucene"] }
{ "id": 2, "tags": ["ranking"] }
{ "id": 3 }go deeper
Know that a JSON array becomes several values on one document and that Elasticsearch has no separate array type. That alone explains the inflated count.
Explain that metric aggregations iterate values rather than documents, so both multi-valued and sparse fields shift the number, and describe what the missing parameter does.
Show how you diagnose it in production: request doc_count and value_count together per bucket, check the values-per-document ratio, and decide whether to denormalize or precompute a per-document metric at ingest.
Own the semantics your reporting layer publishes. Decide whether averages are per document or per value, encode that in the index design and the ingest pipeline, and make sure sparse coverage cannot silently change a published metric's meaning.
## Two different counters Every bucket in an Elasticsearch aggregation response carries a `doc_count`: the number of documents that fell into that bucket. `value_count` is a separate, explicitly requested single-value metric aggregation that counts the values extracted from a field for the documents in scope. People treat them as synonyms and are then surprised when the two numbers disagree, in either direction. ## Why value_count can be larger Elasticsearch has no array type in the mapping. A field is simply allowed to hold zero, one or many values of its mapped type, and a JSON array in `_source` becomes several indexed values on the same document. Given ```json { "tags": ["search", "lucene", "ranking"] } ``` the `tags` field has three values on one document. `value_count` over `tags` counts three; `doc_count` counts one. Index ten such documents and `value_count` reports thirty against a `doc_count` of ten. This is not a quirk of `value_count` alone. Every metric aggregation iterates values, so `sum` over an array of numbers totals every element, and `avg` divides that total by the number of *values*. An "average score per document" computed over a multi-valued `scores` field is really an average per score. If per-document semantics matter, either denormalize so the field is single-valued, or index a precomputed per-document number at ingest time. ## Why value_count can be smaller A document that omits the field, or has it explicitly as `null`, or as an empty array, produces no values and is skipped by every metric aggregation. In a bucket of one thousand documents where only six hundred carry `price`, `value_count` on `price` is 600 while `doc_count` is 1000. That asymmetry is exactly why `avg` on a sparse field is an average over the documents that *have* the field, not over the bucket — a subtle but frequent source of wrong dashboards, because the mean silently changes meaning as data coverage changes. The `missing` parameter is the lever: it names a value to use for documents that lack one. ```json { "aggs": { "filled": { "avg": { "field": "price", "missing": 0 } } } } ``` Now every document contributes, and `value_count` with `missing` set equals `doc_count`. Whether zero is the honest substitute is a modelling decision, not a technical one; sometimes `missing` is right, sometimes excluding absent documents is right, and the interview point is that you chose deliberately. ## Diagnosing the divergence When a report's totals look inflated, the standard check is to request both numbers together and compare: ```json { "size": 0, "aggs": { "per_category": { "terms": { "field": "category" }, "aggs": { "tag_values": { "value_count": { "field": "tags" } } } } } } ``` Each bucket now shows `doc_count` next to `tag_values.value`. A ratio above one means multi-valued documents; below one means sparse coverage. It is worth doing this before trusting any average computed over a field whose cardinality per document you have not verified. ## Field type and cost notes `value_count` reads doc_values like other metric aggregations, so it works on numerics, `keyword`, dates and booleans, and fails on an analyzed `text` field unless fielddata is enabled — which you should not do. Counting values on a `keyword` field is cheap; it does not build the global ordinal structures that a `terms` aggregation needs, because it never has to identify *which* values were seen, only how many. A related trap: `value_count` is not a distinct count. It counts occurrences including repeats, so a document with `["a", "a", "b"]` contributes three. If the same term appears twice in the source array of a `keyword` field, both copies count. Distinct counting is the `cardinality` aggregation, which is approximate; `value_count` is exact but counts multiplicity. ## The useful mental model Aggregations operate on the flattened stream of (document, value) pairs that the columnar storage exposes, not on documents. `doc_count` is the only number in a bucket that counts documents. Once that model is in place, the sum, the average, the value count and the stats count all behave predictably, and the fix for a per-document metric is to make the field single-valued or to precompute the per-document figure at index time.
- Does value_count deduplicate repeated values within one document?No. It counts occurrences, so a document whose keyword array holds ["a", "a", "b"] contributes three. Distinct counting is a different aggregation — `cardinality` — and that one is approximate. Use `value_count` when you want multiplicity, `cardinality` when you want unique values.
- How does the missing parameter change an avg aggregation over a sparse field?Without it, documents lacking the field are skipped and the mean is over covered documents only, so it shifts as coverage changes. Setting `missing` supplies a substitute value, bringing those documents into both the numerator and the divisor. Whether that substitute is honest is a modelling decision — zero is not always right.
saying these in an interview costs you the question
- Treats value_count and doc_count as interchangeable
- Assumes arrays need a special array field type in the mapping
- Thinks value_count returns distinct values
- Believes documents missing the field count as zero in avg
- Reports a per-document average from a multi-valued field without checking