You need latency percentiles across dozens of services. Compare explicit-bucket histograms, exponential (base-2) histograms and the legacy summary data point in OpenTelemetry, and explain why attribute cardinality is a far bigger problem on a metric than on a span.
answer
- Explicit buckets: chosen up front, merge only if identical
- Exponential: scale-derived buckets, constant relative error, merge by downscaling
- Summary: pre-computed quantiles, cannot be re-aggregated — legacy only
- Metric attribute set = a stored series forever; span attribute = one field on one record
- Exemplars bridge from a histogram bucket back to a trace
basics
~20 sExplicit buckets need boundaries chosen up front and only merge across services if identical. Exponential histograms derive buckets from a scale, cover any range and merge by downscaling. Summaries carry pre-computed quantiles that cannot be re-aggregated. Metric attributes multiply stored series; span attributes do not.
solid answer
~1 min**Explicit-bucket histograms** carry a fixed boundary list plus per-bucket counts. They are universally supported and cheap to reason about, but the boundaries are a guess you make before you have the data, they cannot be changed retroactively, and two services with different boundaries cannot be merged — a fleet-wide p99 requires everyone to agree on one boundary list, which is either coarse or wasteful. **Exponential histograms** define buckets implicitly from a `scale` (base 2^(2^-scale)) with a zero bucket, so they adapt to whatever range the data has at bounded relative error and a bounded bucket count; two series at different scales merge by downscaling the finer one, which makes cross-service aggregation actually work. Backend support is newer and uneven. **Summary** points carry client-computed quantiles; percentiles cannot be re-aggregated across instances at all, so it exists for legacy interoperability and not for new instrumentation. On cardinality: a metric attribute set defines a **time series stored for the whole retention period**, and every histogram series carries its full bucket array — so cardinality multiplies both memory in the SDK and storage forever. A span attribute is one field on one record that ages out with it. High-cardinality identifiers belong on spans and logs; metrics need bounded dimensions, with exemplars as the bridge back to individual traces.
code
text · 9 linesservice A bounds: [5, 10, 25, 50, 100] ms
service B bounds: [1, 2, 4, 8, 16, 32] ms
merge? bucket "<=10" from A overlaps B's "<=8" and part of "<=16"
no exact combination exists -> not mergeable
exponential form: A at scale 3, B at scale 5
downscale B 5 -> 3 by pairwise bucket merging (exact)
then add bucket counts -> fleet-wide distributiongo deeper
Know that a histogram buckets values so percentiles can be estimated, and that putting unbounded values like user ids into metric labels is expensive.
Contrast explicit and exponential bucketing, and explain that a metric attribute combination creates a stored time series while a span attribute does not.
Argue mergeability as the deciding criterion, explain downscaling, and use exemplars to keep individual-request access without cardinality.
Set organisation-wide policy: which representation is standard given verified backend support, the permitted metric dimension vocabulary, enforced cardinality limits in the pipeline, and the explicit division of labour between metrics for aggregates and traces/logs for identity.
## The question behind the question "Percentiles across dozens of services" is really "can these distributions be **merged**?" Everything else follows from that, because a per-service p99 tells you little about the fleet, and averaging percentiles is meaningless. ## Explicit-bucket histograms A data point carries `explicit_bounds` (an ordered boundary list), `bucket_counts` (one more entry than boundaries, for the overflow), plus `count`, `sum` and optionally `min`/`max`. Percentiles are estimated by interpolating within whichever bucket contains the target rank. Strengths: universally supported, trivially understood, and you can spend resolution exactly where you care. Weaknesses, all consequences of choosing boundaries in advance: - **You must guess the range.** Boundaries tuned for a 10-500ms service are useless for one that answers in 200µs and equally useless when that service degrades to 40 seconds — everything lands in the overflow bucket and the p99 becomes "more than the last boundary". - **They cannot be fixed retroactively.** Change the boundaries and you have a new series shape; history keeps the old resolution. - **Merging requires identical boundaries.** Summing two histograms with different boundary lists is not defined. So a fleet-wide percentile forces a single organisation-wide boundary list, which is a governance problem as much as a technical one — and one list that suits a cache and a report generator is necessarily coarse. - **Resolution costs series width.** More buckets means a wider array on every series, multiplied by every attribute combination. ## Exponential histograms Buckets are not enumerated; they are derived. A `scale` parameter defines the base as 2^(2^-scale), and bucket index *i* covers (base^i, base^(i+1)]. A separate zero bucket with a zero threshold handles values at or near zero, and positive and negative ranges are tracked separately. The point carries the scale, the offsets and the bucket counts, plus count and sum. What that buys: - **Constant relative error.** Precision is proportional to the value, which is what you want for latency: 1% accuracy at 2ms and 1% at 20s, without knowing in advance which regime the service is in. - **No boundary decision.** The SDK adapts the scale to the observed data, reducing the scale automatically when the range widens beyond the configured maximum bucket count (commonly 160), which bounds memory. - **Merging just works.** Two series at different scales are merged by downscaling the finer to the coarser — buckets combine exactly two-to-one at each scale step. This is the property that makes fleet-wide percentiles feasible without fleet-wide boundary agreement. The cost is ecosystem maturity: SDK support is good but backend support is more recent than for explicit buckets, and some storage systems convert to explicit buckets on ingest, throwing away the property you chose it for. Verify end to end before standardising. ## Summary A summary data point carries a count, a sum and a list of pre-computed quantiles. It exists because some ecosystems produced this shape historically. The disqualifying property for a fleet view is that **quantiles do not aggregate**: given the p99 of ten instances there is no operation that yields the p99 of the ten combined. Use it only where a legacy source forces it; never choose it for new instrumentation. ## Percentiles are always estimates Worth saying out loud: bucketed histograms yield interpolated approximations, and how good they are depends on bucket density near the rank you asked for. A p99.9 from a histogram with sparse tail buckets is a shrug with a decimal point. Exponential histograms improve this because tail buckets stay proportionally sized, but no bucketed representation gives exact quantiles — that is the price of a mergeable, bounded-size summary. ## Why cardinality bites metrics and not spans A span is a record. Adding a `user_id` attribute adds one field to one record, which is stored once and expires with the record's retention. That is why traces are the right home for high-cardinality identity — it is exactly what makes a trace searchable for the one bad request. A metric attribute set **defines a time series**. Each distinct combination is: - an entry the SDK holds in memory for as long as the series is active (and for cumulative temporality, that means holding running state); - a stream of data points at every collection interval, whether or not anything happened; - a stored series in the backend for the full retention period, indexed and billed. And for histograms, each series carries its whole bucket array, so the multiplier is dimensions × buckets. Ten routes × five status classes × three regions is 150 modest series. Add a customer id with 50,000 values and you have 7.5 million series that no query will ever ask for individually. The failure is usually not a crash but a slow degradation: SDK memory creeps, export payloads swell, backend queries slow and bills climb, and the cause is one attribute added months earlier. The governing habits: metric attributes must be drawn from bounded, enumerable sets (route templates never raw paths, status classes rather than exact codes where possible, region, tenant only if tenants are countable); enforce it at the pipeline with cardinality limits so one bad deploy cannot melt the backend; and use **exemplars** as the bridge — a data point can carry the trace and span id of a representative measurement, so you get a fleet-level histogram plus a jump into an individual slow request without putting request identity into the series key. ## The recommendation to state Exponential histograms where the whole path supports them, because mergeability across services is the requirement; explicit buckets with a small, deliberately agreed boundary list where it does not; summaries never, except to ingest legacy sources. Bound metric dimensions hard, keep identity on spans and logs, and wire exemplars so you have not actually lost the ability to find the individual request.
- Why can't you average per-service p99s to get a fleet p99?A percentile is a rank statistic over a population, and rank statistics are not linear — the mean of ten p99s is not the p99 of the union, and it can be off by a large factor when the services have different traffic volumes or different distribution shapes. The only correct route is to merge the underlying distributions and take the percentile once, which is exactly why the histogram representation must be mergeable.
- A team wants a `customer_id` attribute on their request-duration histogram. How do you respond?Refuse it on the metric and offer the alternative. Each customer id becomes its own series carrying a full bucket array, held in SDK memory and stored for the full retention period, and the query nobody runs still costs. Put customer id on spans and logs where high cardinality is the design point, keep the metric dimensioned by route and status, and enable exemplars so a slow bucket links to a real trace that names the customer.
- What breaks if a backend converts exponential histograms to explicit buckets on ingest?You lose the property you adopted them for. The conversion pins the data to whatever boundary list the backend chose, so cross-service merging is again only valid where those boundaries coincide, and the constant-relative-error guarantee is replaced by whatever resolution that fixed list happens to give in your range. It is worth verifying the whole path end to end before standardising, since the SDK and protocol supporting a representation says nothing about how it is stored.
saying these in an interview costs you the question
- Averaging percentiles across services or instances
- Believing explicit-bucket histograms with different boundaries can be summed
- Choosing summary data points for new instrumentation because 'it already has the quantiles'
- Treating metric attributes with the same freedom as span attributes
- Assuming histogram percentiles are exact rather than bucket-interpolated estimates