skip to content

In PromQL, how do you compute a service-wide 95th percentile from histogram bucket series, and what does histogram_quantile assume?

level: seniorimportance: should knowfreq 61%

answer

  1. Buckets are counters first
  2. One label must survive the aggregation
  3. Interpolation happens inside one bucket
  4. A flat p99 means the top bucket

basics

~20 s

Rate the bucket series first, sum them while keeping the le label, then apply histogram_quantile. The function assumes observations are spread evenly inside the bucket it interpolates in, and that every aggregated series shares the same bucket boundaries.

solid answer

~40 s

Write it inside out: `histogram_quantile(0.95, sum by (le) (rate(x_bucket[5m])))`. The bucket series are counters, so they need a `rate` first, otherwise you get a lifetime figure rather than a recent one. The aggregation across replicas must keep `le`, because buckets are only additive along that dimension; grouping it away leaves the function nothing to interpolate over. `histogram_quantile` then finds the bucket the quantile falls into and **interpolates linearly inside it**, which is why the result is an estimate. It assumes every series you aggregated shares identical boundaries — mixing two services with different layouts produces a plausible-looking, meaningless number — and it cannot interpolate above the highest finite bound, so a quantile landing in the `+Inf` bucket comes back pinned to that bound. Never average per-replica percentiles; aggregate buckets, then estimate.

code

promql · 1 line
promql
histogram_quantile(0.95, sum by (le) (rate(hearing_search_duration_seconds_bucket[5m])))

go deeper

for a junior

Recall that a classic histogram is read through its bucket series and that a percentile comes from histogram_quantile rather than from any raw latency value. Recognising the shape of the standard query is enough here.

for a middle

Explain each layer of the query and why it is there: rate because buckets are counters, sum by le because buckets add only along that dimension, then the quantile estimate with its interpolation inside one bucket.

for a senior

Diagnose real results. Recognise a percentile pinned to the top finite boundary, explain why averaging per-replica percentiles is wrong, and say what an aggregation across mismatched bucket layouts silently produces.

for a principal

Own bucket boundaries as a cross-team agreement. Anything meant to be aggregated must share a layout, and the cost of that consistency has to be weighed against the series it creates across the whole estate.

## The inputs the function needs A Prometheus classic histogram is read as a family of cumulative bucket series, each carrying a label `le` — "less than or equal to" — whose value is that bucket's upper bound, plus an `le="+Inf"` series counting every observation. `histogram_quantile` takes a quantile between 0 and 1 and an instant vector of those bucket series, and returns an estimate of the value at that quantile. Everything that goes wrong with the function goes wrong in what you hand it as its second argument. ## The query, built from the inside out For a hearing-search endpoint whose latency histogram is `hearing_search_duration_seconds_bucket`, the service-wide 95th percentile over the last five minutes is: ```promql histogram_quantile(0.95, sum by (le) (rate(hearing_search_duration_seconds_bucket[5m]))) ``` Read it inside out, because each layer does something the next one requires. 1. **`rate(..._bucket[5m])`** — bucket series are counters. Without a rate, or an `increase`, you are asking for the quantile over the entire lifetime of every process since it started, a figure that drifts glacially and never reflects the last five minutes. 2. **`sum by (le) (...)`** — this is the aggregation across replicas, and `le` is the one label that must survive it. Buckets are additive along the `le` dimension: adding the `le="0.25"` counts from 41 replicas gives the fleet-wide number of requests served under 250 ms. Aggregate `le` away and the histogram is destroyed; the function then has nothing to interpolate over. 3. **`histogram_quantile(0.95, ...)`** — finds the bucket the 95th percentile falls into and interpolates within it. If you want the percentile broken out per courtroom, the grouping becomes `sum by (le, courtroom)`. `le` is always in the list. ## What the function assumes - **Observations are spread evenly inside each bucket.** The function locates the bucket containing the quantile and interpolates linearly between that bucket's lower and upper bounds. Real latencies are not uniform inside a bucket, so the answer is an estimate, not a measurement. - **The bucket boundaries are identical across everything you aggregated.** Two services with different bucket layouts summed by `le` produce a set of series that is not a coherent histogram. The output looks perfectly normal and means nothing, which makes this the most dangerous failure on this function. - **The `+Inf` bucket is present**, because that is where the total observation count comes from. - **The lowest bucket starts at zero**, which is the usual case for durations and sizes. ## How it degrades | Situation | What you see | |---|---| | Quantile falls above the highest finite bound | the highest finite bound, flat and unchanging | | One bucket holds most observations | a value pinned near that bucket's bounds, insensitive to real change | | Too few boundaries around the real latency | large quantised jumps as traffic shifts between buckets | | No observations in the window | no result rather than zero | The first row is the one to recognise on sight. A p99 graph sitting perfectly flat on a round number such as 10 is almost never a real latency plateau; it means every slow request landed in the unbounded top bucket, and since the function cannot interpolate to infinity it reports the last boundary it knows. The fix is on the instrumentation side, choosing boundaries that bracket the latencies you actually care about, not in the query. ## Why you cannot average percentiles A tempting query is an average of per-replica percentiles: compute the estimate for each replica, then `avg` them. It is wrong, and not by a rounding error. A percentile is a position in a distribution, not a quantity. The 95th percentile of the combined traffic of 41 replicas is not the average of their individual 95th percentiles, and the error grows with how unevenly the replicas are loaded: a replica serving 12 requests with a slow tail contributes exactly as much to the average as one serving 4,800 fast ones. The composition rule is: **aggregate the buckets, then estimate; never estimate, then aggregate.** The same rule explains why a quantile that arrives as a plain pre-computed number, calculated inside the application before export, cannot be re-aggregated at all. The distribution it came from is gone, and no query can reconstruct it. - Correct: `histogram_quantile(0.95, sum by (le) (rate(x_bucket[5m])))` - Wrong: an `avg` wrapped around a per-replica `histogram_quantile` - Wrong: `histogram_quantile(0.95, sum(rate(x_bucket[5m])))`, which aggregates `le` away

  • A p99 latency graph sits perfectly flat on 10 for hours. What is your first hypothesis?
    That 10 is the highest finite bucket boundary and the 99th percentile has moved above it. The function cannot interpolate into the unbounded top bucket, so it reports the last boundary it knows and the line goes flat. Confirm by checking how much of the total count sits in the top bucket; the fix is on the instrumentation side, extending the boundaries to cover the latencies that actually occur.
  • How would you produce the same percentile broken down per courtroom rather than service-wide?
    Add the dimension to the grouping while keeping le: sum by (le, courtroom) around the rate, then apply histogram_quantile to that. le is always in the list. Be aware this multiplies the output series by the number of courtrooms, and that each courtroom's estimate now rests on far fewer observations, so a low-traffic courtroom's percentile will be noisy.
  • What breaks if two services with different bucket layouts are aggregated together?
    The sum is no longer a coherent cumulative histogram, because a given le value means different things in the two sets of series and some boundaries exist on only one side. The function still returns a number, and that is the danger: nothing errors and nothing looks wrong. Consistent boundaries across anything you intend to aggregate is an instrumentation-side agreement, not something a query can repair.

Averaging each replica's 95th percentile is like averaging two classrooms' median scores and calling the result the median of all the students.

saying these in an interview costs you the question

  • Aggregating bucket series without keeping the le label
  • Averaging per-replica percentiles to get a service percentile
  • Applying histogram_quantile to raw counters without a rate
  • Treating the interpolated estimate as an exact measurement
  • Assuming buckets from different services can be summed safely