skip to content

Why export histogram bucket counts when a service can compute and export its own p99?

level: seniorimportance: nice to knowfreq 27%

answer

  1. Ask what shape crosses the wire
  2. Counts add; order statistics do not
  3. Grouping deferred to query time
  4. Boundaries bought in advance are the price
  5. Geometric spacing trades a list for a resolution

basics

~20 s

Counts are additive: any set of processes and any time range can be pooled and the quantile computed once at query time. An exported quantile is a finished number for one process and one window that nothing downstream can recombine.

solid answer

~50 s

A process that computes its own quantiles keeps an estimator over the values it sees and exports finished numbers — p50, p95, p99 — alongside a count and a sum. On that one process the numbers are good and cheap: no boundary grid to choose, and accuracy that adapts to the actual value range. The problem is what crosses the wire. Quantile **values** are order statistics and do not add, so a backend holding ten of them cannot produce a fleet figure or re-answer the question over a different window. Bucket **counts** can: they add across processes and across time, so grouping is deferred to query time and the same stored data answers per-process, per-region and fleet-wide. The price is choosing boundaries in advance and storing a series per boundary. Geometric (exponential) boundaries reduce that price by deriving the grid from one resolution setting.

go deeper

for a junior

Recall that a process can either export finished percentile numbers or export counts per bucket, and that only the counts can be added together afterwards.

for a middle

Be able to explain why additivity is the deciding property: counts merge across processes and time windows, so the quantile can be computed at query time for any grouping.

for a senior

Weigh the real costs — boundaries chosen in advance, a stored series per boundary against every dimension — and say what geometric boundaries change and what support they demand end to end.

for a principal

Own the fleet-level decision about which representation every service publishes, knowing that the choice fixes for years which questions the stored telemetry can still answer.

## Two ways to ship a latency distribution There are two fundamentally different things a process can put on the wire when asked "how slow were you?". - **A client-computed quantile.** The process maintains an estimator over the values it observes and exports finished numbers — a p50, a p95, a p99 — usually with the total count and the sum of all values. - **Bucket counts.** The process counts observations into a fixed grid of boundaries and exports the counts, again with a total and a sum. Nothing is estimated inside the process at all; the quantile is derived later. The difference that matters is not accuracy on one process. It is **additivity**. | | Client-computed quantile | Fixed-boundary counts | Geometric-boundary counts | | --- | --- | --- | --- | | What crosses the wire | finished quantile values | a count per boundary | a resolution setting plus populated buckets | | Accuracy on one process | high; the estimator saw every value | limited by the grid | bounded relative error at any magnitude | | Combine across processes | no | yes | yes, after reducing to a common resolution | | Re-window after the fact | no | yes | yes | | Configuration burden | pick which quantiles to publish | pick every boundary, per instrument | pick one resolution | | Storage cost | a few values per meter | one series per boundary, per dimension | grows with the value range actually seen | ## What additivity buys you Counts over disjoint sets of observations add. That single property is what lets a backend answer questions nobody anticipated: the p99 of one process, of one region, of the whole fleet, over five minutes or over a day, all from the same stored counts, because the merge happens at query time. Deciding *how* to group is deferred until someone asks. A finished quantile forecloses all of that at the moment of export. It answers exactly one question — this process, this window — and it is the wrong shape for any other. Note carefully where the problem lies: it is **not** that no estimator can be merged. Some sketch structures are explicitly designed to be mergeable, and if the process shipped the structure itself the merge would be possible. The problem is that what crosses the wire is a handful of numbers, and numbers that are order statistics cannot be recombined no matter what the process kept internally. The count and the sum that ride alongside are additive, so a fleet-wide request rate and a fleet-wide mean survive even when the quantiles do not. That is often the only thing salvageable from a quantile-exporting fleet. ## What it costs Bucket counts are not free. - Somebody must choose the boundaries in advance, per instrument, without knowing what the latency range will be next year. - The estimate is bounded by the grid, so a quantile landing in a wide bucket is only known to that width. - Where each boundary becomes its own stored series, a grid of a dozen boundaries multiplies against every dimension already on the metric, across every process in the fleet. That third point is why fine grids are copied around less casually than they are proposed. For a district-heating billing service on a 96-hour retention floor the arithmetic is easy to check before committing; for a large fleet it is the difference between a comfortable metrics bill and an uncomfortable one. ## What geometric boundaries change A grid whose boundaries follow a geometric progression — each one a fixed multiple of the last, derived from a single resolution setting rather than listed by hand — changes the trade-off on three of the axes above at once. 1. **No boundary list to choose.** One resolution setting generates the whole grid, so the same configuration works for a service answering in microseconds and one answering in tens of seconds. 2. **Bounded relative error.** Because bucket widths grow with the value, uncertainty is a percentage of the measurement rather than a fixed number of milliseconds. A grid that resolves 40 ms usefully also resolves 4 s usefully. 3. **Only populated buckets need storing.** Representations that keep just the buckets where observations landed avoid paying for the empty span between a service's floor and its timeout. What they do not change is additivity — counts still add, which is the whole point — although merging two of them means reducing to the coarser of their resolutions first. The practical cost is support: the client, the wire format, the storage and the query layer all have to understand the representation, and a fleet is only as capable as its least capable hop. ## When a client-computed quantile is still right It is a reasonable choice when there is nothing to aggregate: a single-process tool, a batch job reporting its own run, a load generator summarising a test. It is also reasonable when the exact per-process value matters more than any fleet view and you would never merge anyway. Outside those cases, prefer counts and let the backend compute the quantile — and if you must export finished quantiles, keep exporting the count and the sum too, so at least the rate and the mean remain answerable across the fleet.

  • If a fleet already exports finished quantile values, what can you still compute across it?
    The additive companions: total request rate from the counts and a true fleet mean from the summed values over the summed count. You can also bracket the pooled quantile between the smallest and largest of the reported values, which makes the maximum a usable upper bound. What you cannot do is produce the pooled quantile itself, or re-answer it over any window other than the one each process used.
  • Two processes report geometric-boundary histograms configured at different resolutions. Can they be merged?
    Yes, but only down to the coarser of the two. Because the boundaries come from a scale, a finer grid's buckets can be combined pairwise into the coarser grid's buckets exactly, so the merge is lossless in one direction and impossible in the other. It is the same asymmetry as with fixed boundaries: you can always go coarser, never finer, so a fleet meant to be compared should agree on a resolution.

Exporting a finished p99 is like sending a baked cake: it is fine on its own plate, but you cannot combine ten of them into one bigger cake. Bucket counts are the ingredients, and ingredients always recombine.

saying these in an interview costs you the question

  • Thinks exported quantile values can be averaged across processes
  • Believes a client-side estimator is inherently inaccurate
  • Ignores that bucket boundaries must be chosen in advance
  • Assumes geometric boundaries make the quantile exact
  • Forgets that count and sum remain additive either way