skip to content

How do the sampler and diversified_sampler aggregations bound the cost of expensive sub-aggregations?

level: seniorimportance: nice to knowfreq 26%

answer

  1. A single-bucket wrapper that shrinks what sub-aggregations see
  2. Per-shard cap, so shard count multiplies the sample
  3. Selection is by score, not at random
  4. One variant limits hits sharing a field value
  5. Results are approximate by construction

basics

~20 s

Both are single-bucket aggregations that restrict their sub-aggregations to a capped set of documents per shard — sampler keeps the top-scoring shard_size hits, diversified_sampler additionally limits how many hits share a given field value, so one prolific source cannot dominate the sample.

solid answer

~50 s

`sampler` is a single-bucket aggregation that collects only the highest-scoring `shard_size` documents on each shard (100 by default) and runs its sub-aggregations over just that sample. It turns an expensive analysis over millions of matches into one over a few hundred per shard. The catch is that it is only meaningful when the query actually scores — under a pure filter context every document scores the same and the "top" sample is arbitrary. `diversified_sampler` adds a `field` or `script` and `max_docs_per_value` (default 1), capping how many sampled documents may share the same value. That stops a single prolific author, domain or customer from filling the whole sample. The diversification key must be single-valued per document. The canonical use is `significant_terms` or another costly sub-aggregation over a broad match set: sampling makes it fast and often *more* useful, because it strips noise from over-represented sources. Results are approximate — never use them where exact counts matter.

code

json · 16 lines
json
{
  "size": 0,
  "query": { "match": { "body": "battery drain" } },
  "aggs": {
    "sample": {
      "diversified_sampler": {
        "shard_size": 200,
        "field": "thread_id",
        "max_docs_per_value": 2
      },
      "aggs": {
        "keywords": { "significant_terms": { "field": "body.keyword" } }
      }
    }
  }
}

go deeper

for a junior

Not expected at this level. Recall only that Elasticsearch can restrict an expensive aggregation to a sample of the best-matching documents.

for a middle

Explain that sampler is a single-bucket aggregation capping documents per shard by score, that its sub-aggregations see only that sample, and that the results are approximate.

for a senior

Show judgment about when sampling is legitimate: exploratory analysis on scored queries, never authoritative counts, and reach for diversified_sampler when one source dominates the top hits.

for a principal

Own the approximation policy — which surfaces may show sampled figures at all, how they are labelled to users, and when a pre-aggregated summary index is the honest answer instead of a sample.

## The problem they solve Some aggregations cost far more per document than a simple count. `significant_terms` compares foreground and background frequencies; scripted metrics execute code per hit; nested `terms` trees multiply state. When the query matches millions of documents, running such an aggregation over all of them is expensive and often unnecessary — the analysis you want ("what vocabulary distinguishes these results?", "what do the best matches have in common?") is answerable from a representative slice. `sampler` and `diversified_sampler` are single-bucket aggregations that define that slice. They collect a bounded set of documents per shard and expose it as one bucket; whatever sub-aggregations you nest inside see only those documents. ## sampler `sampler` takes one parameter that matters: `shard_size`, the maximum number of documents to collect **per shard**, defaulting to 100. It keeps the highest-scoring ones. Total sample size is therefore roughly `shard_size` times the number of shards queried, which is worth remembering — the same request against a 30-shard index samples thirty times as many documents as against a single shard, so tuning `shard_size` without knowing the shard count is guesswork. Because selection is by score, `sampler` is only coherent under a scoring query. Wrap a `match` query and "top 100" means the 100 best matches. Run it under a `filter` clause, a `bool` with only `filter`, or a `constant_score`, and every document has an identical score, so which 100 you get is an implementation detail — the sample becomes arbitrary rather than representative. This is the single most common misuse. The response includes the bucket's `doc_count`, which tells you how many documents the sample actually held. ## diversified_sampler A score-ordered sample has a failure mode: one source can own the top results. Search a support corpus for a product name and the top 100 hits may all be replies in one long thread; the resulting `significant_terms` then describes that thread rather than the topic. `diversified_sampler` fixes this by adding a diversification key — `field` or `script` — plus `max_docs_per_value`, which defaults to 1. At most that many sampled documents may share a given key value, so the sample spans distinct threads, authors, domains or customers. `shard_size` still caps the total per shard, and `execution_hint` (`global_ordinals`, `bytes_hash`, or `map`) chooses how the seen-values structure is held, trading memory against speed the same way the `terms` aggregation's hint does. The diversification key must be single-valued per document; a multi-valued field is not a legal choice. If the field's values are so uniform that the cap cannot be satisfied, the sample simply comes back smaller than `shard_size` — worth checking `doc_count` before drawing conclusions. ## What you give up Everything inside a sampler is an estimate over a biased-by-score subset. Counts, sums and averages computed inside one are not the counts, sums and averages of the match set, and presenting them as such is a correctness bug, not a performance tradeoff. Use samplers for *discovery* — significant terms, exploratory facets, "what characterizes these results" — and never for reporting figures a user will treat as authoritative. They also do nothing for the query itself. The query still matches every document; only the sub-aggregation work is bounded. If the query is what hurts, filter harder or narrow the time range. ## Where they sit among the alternatives Given an aggregation that is too expensive, the ladder is roughly: narrow the query first (cheapest and exact); reduce the aggregation's own size and depth; switch to `composite` paging if you need every bucket; sample if approximation is acceptable and the query scores; and pre-aggregate with a transform if the same analysis runs repeatedly. Sampling is the right rung specifically when the analysis is exploratory and the match set is large and score-ordered — a narrow but real slot. ## Interview register This is a differentiator question. Knowing that `sampler` exists is worth little; knowing that it selects by score and is therefore meaningless in a pure filter context, and that `diversified_sampler` exists precisely because score-ordered samples over-represent prolific sources, shows you have used them on a real corpus.

  • Why is a sampler aggregation nearly useless inside a pure filter-context query?
    Because it keeps the top-scoring `shard_size` documents, and in filter context every match scores identically. With no score ordering, which documents survive is arbitrary, so the sample is not representative of anything in particular. If you need bounded work under a filter, narrow the filter or pre-aggregate instead.
  • How does shard count affect the size of a sampler's sample?
    `shard_size` is per shard, so the total sample is roughly `shard_size` multiplied by the number of shards the search touched. The same request against a 30-shard index samples thirty times more documents than against a one-shard index, which changes both cost and result stability — tune it with the shard count in mind.
  • What does max_docs_per_value do, and what is its default?
    It caps how many sampled documents may share the same value of the diversification field, defaulting to 1. It prevents one thread, author or domain from monopolizing a score-ordered sample, which is exactly the bias that makes downstream significant_terms results describe a single source rather than the topic.

saying these in an interview costs you the question

  • Using sampler under a filter-only query and expecting representativeness
  • Reporting counts computed inside a sampler as exact totals
  • Forgetting shard_size is per shard, not per request
  • Thinking the sampler reduces the cost of matching documents
  • Diversifying on a multi-valued field

context