In Prometheus, why does one histogram appear as many series in the exposition format, and why is the bucket boundary a label rather than a metric name?
answer
- One histogram is not one series
- Suffixes: bucket, sum and count
- The boundary rides along as a label value
- The last bucket is always +Inf
basics
~20 sPrometheus has no multi-value sample, so one histogram is exposed as many plain series: a <name>_bucket series per boundary carrying an le label, plus <name>_sum and <name>_count. A summary instead exposes quantile-labelled series alongside its own _sum and _count.
solid answer
~50 sA Prometheus sample is one `float64` on one series, so a composite type has to be decomposed. A histogram is exposed as `<name>_bucket{le="<boundary>"}` — one series per boundary, counts **cumulative** so each includes all lower buckets — plus a mandatory `le="+Inf"` bucket equal to `<name>_count`, and a `<name>_sum` of all observed values. A summary emits no buckets: it puts client-computed estimates on the base name under a `quantile` label, with `_sum` and `_count` alongside. The boundary is a label because that keeps every boundary in one selectable family: label matchers reach all of them at once, aggregation operators can group and re-aggregate across instances with `le` as the grouping key, and adding a boundary adds a series instead of creating a new metric name that no query mentions. `le` and `quantile` are reserved on those series.
code
text · 8 lines# HELP seedorder_pack_duration_seconds Time to pack one seed order
# TYPE seedorder_pack_duration_seconds histogram
seedorder_pack_duration_seconds_bucket{region="eu-west-3",le="0.05"} 12047
seedorder_pack_duration_seconds_bucket{region="eu-west-3",le="0.25"} 38914
seedorder_pack_duration_seconds_bucket{region="eu-west-3",le="1"} 41266
seedorder_pack_duration_seconds_bucket{region="eu-west-3",le="+Inf"} 41339
seedorder_pack_duration_seconds_sum{region="eu-west-3"} 7412.883
seedorder_pack_duration_seconds_count{region="eu-west-3"} 41339go deeper
Recall the three suffixes a histogram produces — _bucket, _sum and _count — and that the boundary appears as an le label value. Knowing a summary looks different, with a quantile label, is already a good answer at this level.
Explain that bucket counts are cumulative, that the +Inf bucket equals _count, and that every suffix carries the family's other labels too, so one histogram multiplies into many series per label combination.
Demonstrate why the label layout is what makes buckets re-aggregable across instances, and be ready to spot an exporter that changed how it formats a boundary and quietly forked the family into two.
Own the fleet-level consequence: how many series a standard histogram shape costs across every service and label combination, and who decides which families get one at all.
## One sample holds one number The unit Prometheus stores is a single `float64` at a single timestamp on a single series. There is no composite value: no struct, no array, no map. A histogram is composite by nature — a set of counts, a running total of the observed values and a count of observations — so it cannot be one series. The exposition format solves this by decomposing it into several ordinary series that a naming convention ties back together. ## What a histogram looks like on the wire For a family declared `# TYPE seedorder_pack_duration_seconds histogram`, the target emits: - `seedorder_pack_duration_seconds_bucket{le="<boundary>"}` — one series per boundary. Each count is **cumulative**: it is the number of observations less than or equal to that boundary, including everything counted in the lower buckets. - `seedorder_pack_duration_seconds_bucket{le="+Inf"}` — mandatory, and numerically equal to the `_count` series. - `seedorder_pack_duration_seconds_sum` — the running total of every observed value. - `seedorder_pack_duration_seconds_count` — the running number of observations. Every one of those suffixed series also carries the family's own labels. If the family is dimensioned by `region`, then each bucket, the sum and the count exist once per region: a histogram with eight finite boundaries across four regions is 4 x (8 + 1 + 2) = 44 series from what a developer thinks of as one metric. | Declared type | Series emitted | Reserved label | |---|---|---| | `counter` | `<name>` (named `<name>_total` by convention) | none | | `gauge` | `<name>` | none | | `histogram` | `<name>_bucket`, `<name>_sum`, `<name>_count` | `le` on the bucket series | | `summary` | `<name>` with quantiles, `<name>_sum`, `<name>_count` | `quantile` on the base series | ## The summary's different shape A summary emits no `_bucket` series at all. It emits the base metric name carrying a `quantile` label — `seedorder_pack_duration_seconds{quantile="0.99"}` — and those numbers are estimates the client has already computed, expressed in the metric's own unit rather than as counts. The same `_sum` and `_count` series appear alongside them. So the two types differ in shape as well as in meaning: a histogram ships counts per boundary and leaves any estimation to the query side, a summary ships the estimates themselves. Both reserve a label name on their series — `le` and `quantile` respectively — which you may not reuse for a dimension of your own. ## Why the boundary is a label and not part of the name Four reasons, and any one of them is enough: 1. **One selector reaches the whole family.** `seedorder_pack_duration_seconds_bucket` selects every boundary at once and label matchers narrow it. If the boundaries were baked into names — `..._bucket_0_05`, `..._bucket_0_25` — there would be no way to ask for "all the buckets" without enumerating them. 2. **Aggregation is dimension-aware.** Prometheus's aggregation operators group by labels, so buckets from twenty instances can be summed dimension-wise into one fleet-wide histogram with `le` kept as the grouping key. Boundary-named metrics cannot be grouped that way; you would hand-align twenty expressions instead. 3. **Changing a boundary is a data change, not a schema change.** Adding a 2.5-second boundary adds one series to an existing family. If boundaries were names, it would create a new metric that nothing queries. 4. **The quantile-estimating function takes the family as one input.** That is only expressible because every boundary shares one metric name and differs only in a label. ## Reading the layout correctly - Bucket counts are cumulative, so the number of observations that fell *between* two boundaries is the difference of two adjacent bucket series, not the value of one bucket. - `le` values are **strings** and are matched as strings. `le="1"` and `le="1.0"` are different label values and therefore different series, so an exporter that changes how it formats a boundary silently forks the family. - The `+Inf` bucket is not decoration. Without it there is no total to normalise against and nothing tells you how many observations landed past the largest finite boundary. - `_sum` behaves like a monotonic counter only while observations are non-negative, which is a good reason to keep them so. ## The cost you are signing up for The multiplication is the thing to internalise: every boundary you add multiplies through every label combination on the family, and `_sum` and `_count` ride along with each combination too. Where the boundaries should sit, and what an estimate drawn from them is worth, are separate questions with their own answers. What the exposition format fixes is only the layout — one metric name plus one reserved label, expanded into one ordinary series per boundary, so that everything downstream can treat a histogram as nothing more exotic than a set of counters.
- What must the `+Inf` bucket of a Prometheus histogram equal, and why is it mandatory?It equals the `_count` series, because every observation is at or below infinity. Since buckets are cumulative, the top bucket is the only place the total appears in the bucket family itself. Without it there is no denominator to normalise the lower buckets against, and nothing reveals how many observations landed beyond the largest finite boundary.
- Can you attach your own label named `le` to a histogram metric in Prometheus?No. `le` is reserved on `_bucket` series, exactly as `quantile` is reserved on summary series. Reusing either name collides with the dimension the format already uses there, and the result is a family whose buckets are split or merged by a dimension that has nothing to do with the boundaries — which corrupts everything computed from it.
- How do a histogram's `_sum` and `_count` series behave over time?Both are cumulative: they only increase while the process lives and reset to zero when it restarts. `_sum` accumulates the observed values in the metric's unit, `_count` the number of observations. Their ratio over a window gives the average observation for that window, which is often more useful and far cheaper than any estimate drawn from the buckets.
Prometheus has no spreadsheet cell that can hold a chart, only ordinary cells. So a histogram is laid out across a row of ordinary cells, and the le label is the column heading that says which boundary each cell counts up to.
saying these in an interview costs you the question
- Thinks a histogram is stored as one multi-value sample
- Reads bucket counts as per-bucket rather than cumulative
- Believes le is part of the metric name
- Omits the +Inf bucket when hand-writing an exposition
- Assumes summary quantile series carry counts, not values