skip to content

How do you turn /gc/pauses:seconds from runtime/metrics into a per-interval distribution?

level: middleimportance: nice to knowfreq 25%

answer

  1. one slice longer than the other
  2. the numbers only ever go up
  3. this interval means subtraction
  4. the slice you kept may be overwritten
  5. copy Counts before the next read

basics

~10 s

Call Value.Float64Histogram, copy its Counts, and subtract the previous tick's Counts elementwise, because /gc/pauses:seconds is cumulative since process start. Buckets holds len(Counts)+1 boundaries, so Counts[n] covers the range from Buckets[n] to Buckets[n+1].

solid answer

~50 s

A histogram metric arrives as `KindFloat64Histogram`, and `sample.Value.Float64Histogram()` gives you a `*metrics.Float64Histogram` with two slices: `Counts`, the weight in each bucket, and `Buckets`, the boundaries. `Buckets` is one longer than `Counts` — `Counts[n]` covers `[Buckets[n], Buckets[n+1])` — and the outer edges may be negative or positive infinity. The counts are cumulative for the whole process lifetime, so a distribution "for the last minute" is the elementwise difference between the current `Counts` and the copy you kept from the previous tick. Two traps: you must copy `Counts` before the next `Read`, because reusing the same `Sample` lets the runtime overwrite the histogram in place; and you should re-read `Buckets` alongside `Counts` rather than caching boundaries indefinitely and assuming your saved layout still lines up. From the delta you get a rate by summing it, and an approximate quantile by walking the buckets — approximate, because a bucket only tells you a range.

code

go · 15 lines
go
s := []metrics.Sample{{Name: "/gc/pauses:seconds"}}
metrics.Read(s)
h := s[0].Value.Float64Histogram() // len(h.Buckets) == len(h.Counts)+1

cur := make([]uint64, len(h.Counts))
copy(cur, h.Counts) // copy before the next Read overwrites it

if len(prev) == len(cur) {
	delta := make([]uint64, len(cur))
	for i := range cur {
		delta[i] = cur[i] - prev[i] // pauses observed in this interval
	}
	report(delta, h.Buckets)
}
prev = cur

go deeper

for a junior

Know that some runtime metrics are distributions rather than single numbers, that the accessor is Value.Float64Histogram, and that the counts are running totals since the process started.

for a middle

Explain the bucket layout — one more boundary than counts, outer edges possibly infinite — and show the elementwise subtraction that turns cumulative counts into an interval. Expect a question about why your copy matters.

for a senior

Demonstrate the operational instincts: copy before the next read, guard against a length mismatch, skip the first interval after start-up, and export a bounded summary rather than every raw bucket.

for a principal

Decide what a distribution is worth carrying: raw buckets every tick is a lot of data for a number people read after an incident, so pick the summary shape and the retention deliberately rather than exporting everything the runtime offers.

## What a histogram metric looks like Most `runtime/metrics` values are a single number. A few are distributions — `/gc/pauses:seconds` (how long stop-the-world GC pauses lasted) and `/sched/latencies:seconds` (how long goroutines waited between becoming runnable and running) are the two you meet first. Their `Value.Kind()` is `KindFloat64Histogram`, and the accessor is `Value.Float64Histogram()`, returning a pointer to: ``` type Float64Histogram struct { Counts []uint64 Buckets []float64 } ``` `Counts[n]` is the weight observed in the range `[Buckets[n], Buckets[n+1])`. Hence `len(Buckets) == len(Counts)+1`, and the first and last boundaries may be `-Inf` and `+Inf` so that the buckets tile the whole real line. Calling `Value.Uint64()` on such a metric panics — always switch on `Kind()`. ## Cumulative, not per-interval The key property is the one in `Description.Cumulative`: for `/gc/pauses:seconds` it is true. The counts are totals accumulated since the process started, and there is no reset — the runtime does not offer one, and asking for one misunderstands the model. So a chart of "GC pauses in the last minute" is not something you read; it is something you compute. The computation is an elementwise subtraction between successive reads: ``` delta[i] = current.Counts[i] - previous.Counts[i] ``` Sum the delta and you have the number of pauses in that window. Multiply each delta by a representative value for its bucket and sum, and you have an approximate total pause time. Walk the delta accumulating counts until you cross 99% of the total and you have an approximate p99 — approximate because a bucket tells you an interval, not the observations inside it. If you need an exact worst-case number, a histogram is the wrong instrument; take the maximum from something that records individual values. ## The retention trap `metrics.Read` is designed for slice reuse, and a reporting daemon should reuse its `[]metrics.Sample` on every tick. But the histogram you get back is not a private copy handed to you for keeps — reading into the same `Sample` again can overwrite the counts in place. If you stash the `*Float64Histogram` from tick N and diff it against tick N+1, you may find you are subtracting a slice from itself and reporting all zeros. The fix is one line: copy `Counts` into a slice you own before the next read. ``` prev := make([]uint64, len(h.Counts)) copy(prev, h.Counts) ``` That is also the reason to compute your delta immediately after the read rather than queueing histograms for later processing. ## The bucket-layout trap The second, subtler trap is treating `Buckets` as a constant you can read once at start-up and cache forever. Read the boundaries from the same sample you read the counts from, and if you are keeping a previous `Counts` to diff against, keep the length you saw with it — if the layout you get back is not the same length as the one you saved, the honest move is to drop that interval and start again from the new baseline rather than to subtract mismatched slices. In practice a defensive `if len(cur) != len(prev)` guard costs nothing and prevents a whole class of nonsense numbers in a dashboard. ## Exporting it Whatever collects your numbers usually wants either a rate or a small set of summary values. From one delta you can produce: a count of events in the window, an approximate sum, and a handful of approximate quantiles. Exporting all of the raw buckets is possible but the bucket layout is the runtime's, not yours, and it is finer-grained than most dashboards need — for a reporting daemon that wakes on a ticker, a count plus two or three quantiles per interval is usually the right shape, and it keeps the per-tick work bounded and predictable. ## Why this is worth knowing GC pause and scheduler-latency distributions are precisely the numbers you want in a postmortem, and they are the ones most often recorded wrongly: charted as cumulative totals that only ever rise, or diffed against a histogram that was overwritten in place so the line sits flat at zero. Both mistakes look like data until someone needs to read it under pressure.

  • How do you get a p99 GC pause out of that delta?
    Sum the delta to get the total observation count, then walk the buckets accumulating counts until you pass 99% of that total; the boundary you are standing on is your p99. It is inherently approximate — a bucket reports a range, not the values inside it — so quote it as a bucket edge rather than pretending to millisecond precision.
  • What happens if you call Value.Uint64 on /gc/pauses:seconds?
    It panics. The accessors are checked against the value's kind, and this metric is `KindFloat64Histogram`, so the only legal accessor is `Float64Histogram()`. That is why a reporting loop switches on `Value.Kind()` instead of assuming a shape per name — the assumption is exactly what breaks when a name resolves to something you did not expect.
  • Can you reset the histogram so each interval starts from zero?
    No — there is no reset, and the counts are cumulative by design. The interval is entirely a client-side construct: keep the previous copy, subtract, report. That also means the first tick after start-up has no baseline to diff against and should be skipped rather than reported as a spike.

The histogram is an odometer per speed range, not a trip meter: to say what happened this hour you have to write down the readings and subtract.

saying these in an interview costs you the question

  • Charts cumulative bucket counts as if they were per-interval
  • Assumes len(Buckets) equals len(Counts)
  • Keeps the returned histogram pointer across reads without copying
  • Looks for a reset call on the histogram
  • Quotes an exact p99 from bucketed data