How do you choose histogram bucket boundaries for latency, and what breaks when nearly all requests land in one bucket?
answer
- The grid is a ruler printed in advance
- Bracket from fastest response to past timeout
- Resolution belongs at decision thresholds
- Everything in one bucket flattens every percentile
- You can coarsen later, never refine
basics
~20 sBoundaries must bracket the decision range — below the fastest realistic response, above the timeout, densest at the thresholds you act on. If one bucket holds nearly all traffic, every percentile is interpolated inside it and moves as one line.
solid answer
~50 sBucket boundaries are chosen once and are effectively permanent, so choose them against decisions rather than round numbers. The **lowest** boundary must sit below the fastest response you realistically serve, or everything piles into the first bucket. The **highest finite** boundary must sit above your request timeout, or the tail falls into the open-ended top bucket where there is no upper edge to interpolate against. In between, put boundaries where thresholds live, and space the rest geometrically so a handful of buckets covers milliseconds to seconds. When 99.6% of traffic lands in the first bucket, every percentile interpolates inside that one interval, so p50, p95 and p99 become a fixed fraction of the boundary and stop depending on the data — a tenfold regression that stays inside the bucket does not move the graph at all.
go deeper
Know that the boundaries are fixed when the histogram is created and that observations are counted into them, so the values you can distinguish are decided in advance.
Be able to explain both failure modes: a lowest boundary above most traffic flattens the low percentiles, and a highest finite boundary below the timeout hides the tail in an open-ended bucket.
Demonstrate that you choose boundaries from thresholds and measured traffic rather than round numbers, and that you know a change only helps future data and leaves a seam in the graphs.
Own the tradeoff between resolution and stored cost across many services, and decide where a shared grid is worth more than a per-service fit because it makes numbers comparable between teams.
## What the boundaries have to bracket A bucket grid is a ruler printed before you know what you will measure. Three properties matter. 1. **The lowest boundary must be below the fastest response you realistically serve.** Everything faster than it is indistinguishable, and if that is most of your traffic the lower percentiles are meaningless. 2. **The highest finite boundary must be above your request timeout.** Everything slower lands in the open-ended top bucket, which has no upper edge to interpolate against. When a quantile's rank falls there, implementations differ in what they report — some the highest finite boundary, some an infinite value — and none of them is a measurement. 3. **Resolution belongs where decisions are made.** A boundary placed exactly at a threshold turns "what fraction finished under the threshold" into a ratio of two counters with no estimation in it. Everywhere else, geometric spacing (each boundary a fixed multiple of the previous) covers a wide dynamic range with few buckets, because latency distributions are read on a logarithmic scale, not a linear one. Boundaries are not free. In formats where each boundary is exposed as its own stored time series, each one you add multiplies with every dimension already on the metric, so a grid that is twice as fine is twice as expensive across the whole fleet. The useful test for an extra boundary is whether any decision would change if it existed. ## The one-bucket dashboard Consider the billing API of a district-heating system. Its latency histogram was configured with boundaries at 0.5 s, 1 s, 2.5 s, 5 s and 10 s, on the reasonable-sounding grounds that billing calls are slow. In practice 99.62% of its 218,743 daily requests finish inside 0.5 s. Now every percentile up to the 99th has its rank inside the first bucket, and interpolation across that bucket runs from zero to 0.5 s. The dashboard reads: - p50 near 251 ms - p90 near 452 ms - p99 near 497 ms Those numbers are not derived from the data. They are the requested percentile multiplied by the boundary, and they will read the same whether typical requests take 12 ms or 480 ms. A tenfold latency regression that keeps every request under half a second moves nothing on the graph; three people sharing an on-call rotation will see a perfectly flat p99 while users are complaining. The tell is that the percentile lines are suspiciously smooth and sit in fixed proportion to each other. The mirror-image failure is a top boundary that is too low. If the service times out at 30 s but the grid stops at 10 s, then in an incident where the tail crosses 10 s the p99 stops moving — exactly when you needed it to keep moving — because the rank has moved into the open-ended bucket. ## Why boundaries are effectively permanent Changing the grid does not change data already stored. | What you want to do | Is it possible? | | --- | --- | | Merge adjacent buckets after the fact | Yes — counts add, so you can always go coarser | | Split a bucket after the fact | No — the information to divide it was never recorded | | Compare a window before and after a boundary change | Only at a grid both windows share | | Get finer history for last week's incident | No — you get it from next week onward | That asymmetry is the whole reason boundary choice is a design decision rather than a setting. A change takes effect from the moment of deployment; until the old data ages out, any query spanning the change is mixing two grids. With a 96-hour retention floor, the graph is only homogeneous again four days after the change lands — and an incident review that reaches back across that seam has to be read with the change in mind. ## A procedure that works - Measure first. Run with a deliberately wide, coarse grid for a week and look at where the counts actually accumulate. - Put explicit boundaries on the values that appear in objectives, contracts or alerts. - Add a boundary just above the timeout, so the difference between "slow" and "gave up" is visible. - Fill the rest geometrically — roughly doubling — rather than with round decimal numbers, which cluster resolution in the wrong place. - Keep one grid per class of endpoint rather than one per endpoint: comparability across services is worth more than a perfect fit for each. - Re-check after a major performance change, and accept that the new grid buys you nothing retroactively. The judgement being tested is whether you understand that a bucket grid encodes an assumption about the service's latency range, and that the assumption ages: a grid chosen when a call took two seconds is actively misleading a year later when it takes forty milliseconds.
- The grid is wrong and you cannot redeploy today. What numbers from this histogram can you still trust?Everything that is a count rather than an estimate. The number of requests at or below any existing boundary is exact, so any ratio built from two of them is exact, and the mean is exact from the running sum over the total count. You can also merge buckets to answer coarser questions. What you cannot do is recover resolution inside a bucket, so stop quoting the interpolated quantile and quote the fraction under a boundary instead.
- How would you decide that a histogram has too many boundaries?Cost is paid per boundary and multiplied by every dimension already on the metric, so the test is decision value: if removing a boundary would not change any alert, objective or diagnosis, it is paying rent for nothing. Boundaries clustered in a region where almost no observations land are the usual candidates, as are grids copied unchanged onto services whose latency ranges differ by orders of magnitude.
Bucket boundaries are the marks printed on a ruler: after it leaves the press you can always read it more coarsely, but you can never read between two marks that were never printed.
saying these in an interview costs you the question
- Picks round numbers unrelated to the service's real latency
- Adds dozens of boundaries without accounting for the cost
- Assumes new boundaries retroactively improve historical graphs
- Leaves the highest finite boundary below the request timeout
- Reads a flat percentile line as evidence of a stable service
- Believes a bucket can be split back into finer ones later