skip to content

Your billing report counts unique customers with a cardinality aggregation — how do you decide whether that approximation is acceptable?

level: principalimportance: should knowfreq 32%

answer

  1. ask who reads the number first
  2. tuning a parameter has a hard ceiling
  3. the response carries no error bound
  4. make distinctness a property of the documents
  5. exact counting can be a modelling choice

basics

~20 s

Classify the consumer first: dashboards and trends tolerate estimates, invoices and audits do not. The cardinality aggregation returns no error bound, so if the number must be defensible, restructure so an exact metric answers it rather than tuning precision.

solid answer

~50 s

Start from who reads the number and what happens if it is wrong. Trends, anomaly detection and relative comparison survive a fraction of a percent of error; an invoice, a regulatory filing or anything a customer reconciles against their own records does not. Then be honest about the ceiling: `precision_threshold` tops out at 40000, the aggregation returns a bare value with no confidence interval, and its error is a stable bias rather than noise, so retrying tells you nothing. If exactness is required, change the shape of the problem instead of the setting — pivot to one document per customer per period, using a transform or an ingest-time rollup, so the distinct count becomes an exact `value_count` or `doc_count`; or compute the figure in an offline pipeline of record and use Elasticsearch only for exploration. Finally, label estimated figures in the UI so nobody quietly promotes one into a contract.

code

json · 9 lines
json
// exploratory: approximate is fine and cheap
{
  "size": 0,
  "aggs": {
    "unique_customers": {
      "cardinality": { "field": "customer_id" }
    }
  }
}

go deeper

for a junior

Know that the number is an estimate and that estimates do not belong on invoices. Escalate rather than tuning a parameter and hoping.

for a middle

Explain the hard limits — a capped precision setting, no error bound in the response, deterministic bias — so it is clear that no configuration makes the aggregation exact.

for a senior

Propose the restructure: pivot to an entity-per-customer index with a transform or a batch rollup so the distinct count becomes an exact document count, and know why summing per-shard results is wrong.

for a principal

Own the accuracy contract across the platform: which metrics are approximate by design, which have a separate system of record, how estimates are labelled in the API and UI, and how reconciliation catches drift before a customer does.

## Frame it as a contract question, not a tuning question The wrong instinct is to reach for `precision_threshold` and hope. The right first move is to ask what the number is *for*. Distinct-count consumers fall into rough tiers: - **Exploratory and directional** — dashboards, capacity trends, "is traffic up this week". A relative error of a fraction of a percent changes no decision. Approximation is not merely tolerable, it is the correct engineering choice, because exact counting at this cadence would cost far more than the accuracy is worth. - **Operational** — alerting thresholds, anomaly detection. Approximation is usually fine, but the bias must be stable so a threshold does not flap. HyperLogLog++ obliges here: the estimate is a deterministic function of the data, not a random sample. - **Financial, contractual or regulatory** — invoices, revenue recognition, seat-count enforcement, compliance reports. These need a number that can be reproduced and defended line by line. Approximation is not acceptable at any precision setting, and the conversation should not be about tuning. ## Why raising precision does not rescue tier three `precision_threshold` defaults to 3000 and is capped at 40000; the cap exists because memory is roughly the threshold times eight bytes, per aggregation, per bucket, per shard. Above the threshold, error grows gradually rather than stopping. So a service with millions of unique customers cannot be made exact by any setting. Worse for auditability, the response carries a single value with no error bound. You cannot look at a given result and say how far off it might be, and because the estimate is deterministic, re-running the query returns the same figure. That is good for stability and bad for detection: a wrong number looks exactly as confident as a right one, and reproducibility is easily mistaken for correctness. ## Restructure so an exact metric answers the question The durable fix is to move the distinctness out of query time and into the data model. Distinct counting is only hard because many documents share a customer. If the index instead holds one document per customer per period, "how many distinct customers" becomes "how many documents" — a `doc_count` or a `value_count`, both exact and cheap. Elasticsearch offers a first-class way to build that: a **transform** pivots a source index into an entity-centric destination index, grouped by customer and period, with the metrics you need attached. Counting entities in the destination is then exact. The alternatives are equivalent in spirit — an ingest-time rollup, a nightly batch job writing a summary index, or a warehouse table that is the system of record for billing while Elasticsearch stays the exploration surface. A third option, exhaustive enumeration through a `composite` aggregation paging over every customer key, is exact but pays a full scan of the key space per report. That is defensible for a monthly job over a bounded key space and indefensible for anything interactive. ## Things that look like fixes and are not - **Summing per-shard cardinality results.** Any customer active on more than one shard is counted more than once. The whole reason Elasticsearch merges sketches instead is to avoid this. - **Cross-checking against a second approximate query.** Two runs of a deterministic estimator agree with each other and with nothing else. - **Comparing to a `terms` aggregation's bucket count.** That aggregation returns only the top buckets by size, so its bucket count is not the field's cardinality at all. ## The organizational half Even with the right pipeline, the failure mode is human: someone screenshots a dashboard tile into a customer email, and an estimate becomes a commitment. Principal-level answers cover the guardrails as well as the mechanism — estimated figures visibly marked as approximate in the UI and the API response schema, the system of record named in the documentation for every published metric, and a periodic reconciliation between the exploratory number and the billed number so drift is discovered by a job rather than by a customer. Decide once, per metric, which tier it belongs to, and let that decide the implementation. "We used cardinality because it was already there" is the answer that loses the interview; "exploration is approximate by design, billing reads an entity-per-customer index, and both are labelled" is the one that wins it.

  • How does pivoting to one document per customer per month make the count exact?
    Distinct counting is only hard because many documents share a customer. If the destination index holds exactly one document per customer per month, the distinct-customer question becomes a document count, which Elasticsearch answers exactly and cheaply. A transform, an ingest-time rollup or a nightly batch job can all produce that shape.
  • Why is re-running the cardinality aggregation a poor way to sanity-check its accuracy?
    The estimate is a deterministic function of the hashed values, so the same document set always yields the same number. Agreement between runs demonstrates determinism, not correctness. Validation requires an independent exact count over the same data, not a repetition of the same estimator.
  • When is exhaustive enumeration through a composite aggregation the right answer?
    When you need exactness over a bounded key space on a batch cadence and cannot maintain a pivoted index. It pages through every key, so cost scales with cardinality and it is unsuitable for interactive requests. For a monthly reconciliation job it is a reasonable, self-contained way to produce a defensible number.

saying these in an interview costs you the question

  • Proposes raising precision_threshold until the count is exact
  • Treats reproducibility of the estimate as proof of accuracy
  • Sums per-shard distinct counts to reach a cluster total
  • Uses a terms aggregation's bucket count as the field cardinality
  • Publishes an estimated figure with no approximation label

context