skip to content

Sampling Runtime Metrics

runtime/metrics is the supported, cheap way to read what the runtime knows about its heap and scheduler, and it supersedes ReadMemStats — but sampling it in a hot loop still costs.

part ofGo (Golang)overview, primer and where to startread it →
on this pageshow

questions

4

Why does runtime.ReadMemStats stop the world when runtime/metrics.Read does not?

level: middleimportance: must knowfreq 50%

answer

  1. one promises a coherent whole struct
  2. coherence has to be bought somehow
  3. allocator counters live per-P
  4. the other names only what it needs
  5. a lock is cheaper than a world stop

basics

~20 s

runtime.ReadMemStats stops the world for the whole call so that every field of the MemStats struct is one mutually consistent snapshot. runtime/metrics.Read instead aggregates the runtime's per-P statistics under an internal lock, so it adds no stop-the-world pause.

solid answer

~50 s

`runtime.ReadMemStats` promises that all of `MemStats` describes the same instant. Allocation statistics are maintained per-P for speed, so producing that single coherent picture means halting every goroutine, draining the per-P caches and filling the struct — the runtime does exactly that, and the pause is charged to your whole program, not just to the caller. `runtime/metrics` was designed to remove that cost: `metrics.Read` computes only the aggregates the metrics you asked for actually depend on, using the runtime's consistent-statistics machinery, guarded by an internal lock rather than a world stop. Practically: `ReadMemStats` on a slow ticker is fine, `ReadMemStats` per request or in a loop injects latency into every goroutine in the process, and `metrics.Read` is the default for anything you sample regularly. `metrics.Read` is not free either — it takes a global lock and concurrent readers serialise on it — so batch all your names into one call.

code

go · 7 lines
go
var ms runtime.MemStats
runtime.ReadMemStats(&ms) // stops the world for the whole call
_ = ms.HeapAlloc

samples := []metrics.Sample{{Name: "/memory/classes/heap/objects:bytes"}}
metrics.Read(samples) // no stop-the-world; internal lock only
_ = samples[0].Value.Uint64()

go deeper

for a junior

Remember the headline: ReadMemStats pauses the whole program to fill its struct, so it belongs on a slow ticker, never on a request path. runtime/metrics is the modern way to sample the same ground.

for a middle

Explain why the pause is needed at all — allocator counters are kept per-P for speed, so a mutually consistent struct means halting the mutators — and what metrics.Read does instead. Expect to be pushed on whether Read is free.

for a senior

Show the diagnosis: a service whose pause events track request volume rather than allocation volume usually has instrumentation calling ReadMemStats in a handler. Say how you would confirm it and where you would move the call.

for a principal

Own the standard: one sampling goroutine, one batched read, an interval tied to what actually consumes the numbers, and a rule that no request path reads runtime statistics. That is a review rule, not a per-incident fix.

## Why a snapshot is expensive Go's allocator keeps its counters where allocation happens: on the per-P mcache, so that the fast path needs no shared writes. That design is what makes allocation cheap, and it is also why an accurate global total is not simply sitting in a variable somewhere. Any API that promises *every* number in a struct describes the same instant has to reconcile all of those distributed counters while nothing is changing them. `runtime.ReadMemStats(&m)` makes exactly that promise, and it keeps it the blunt way: it stops the world, walks the runtime's state on the system stack to fill the struct, and starts the world again. Nothing else in your program runs during that window. The pause is typically short — the same order as the runtime's other stop-the-world operations — but it is a real latency event and it is paid by every goroutine, including the ones serving your slowest tail requests. Call it once a second and it disappears into the noise; call it per request, or inside a benchmark loop, and you have built a latency generator whose only output is a number nobody is reading that often. ## What runtime/metrics does instead `runtime/metrics` was introduced to break the tie between *observing* the runtime and *pausing* it. `metrics.Read` does not stop the world. It takes an internal runtime lock, ensures the aggregates that the requested metrics depend on are up to date — reading the runtime's consistent per-P statistics structures, which are designed to be readable without halting the mutators — and then fills each `Sample.Value`. Two properties follow from that design, and both matter when you write a reporting daemon: 1. **Cost scales with what you asked for, not with the whole world.** A metric that is a plain counter is nearly free to read. A metric that requires aggregating heap statistics across all Ps costs more. A histogram costs more again, because its bucket counts have to be produced. 2. **The aggregates are computed once per `Read` call.** Six names in one call share the underlying aggregation and one lock acquisition; six separate calls repeat both. Batching is a genuine optimisation, not style. The lock is the part people forget. `metrics.Read` is cheap relative to a world stop, but it is a process-wide serialisation point, so a goroutine reading it in a tight loop can contend with, say, a debug endpoint doing the same thing. ## Choosing between them For anything you sample on a schedule, `runtime/metrics` is the default. It covers the ground `MemStats` covers — allocation totals, heap occupancy by memory class, GC cycle counts, GC pause distributions — with names that are discoverable at runtime and a documented unit per name. `ReadMemStats` still has a place. If you have existing code or a dashboard built on `MemStats` field names, or you specifically want the struct's own shape (its fixed `PauseNs` ring of recent pause durations, its `BySize` table), you read it — just rarely, on a ticker measured in seconds, never on a request path. A nuance worth stating in an interview: the cost of `ReadMemStats` is not proportional to how much memory you are using in any simple way, and it is not a GC. It does not force a collection, it does not free anything, and it does not allocate. The cost is the stop-the-world round trip plus the work to reconcile per-P state, so it is closer to constant-with-P-count than to anything about heap size. ## The failure mode this prevents The classic incident is self-inflicted: a well-meaning instrumentation layer calls `ReadMemStats` inside a middleware to record heap usage per request. Under load, the process now stops the world thousands of times a second. Latency degrades in a way that looks like GC pressure — because it *is* pause time — and the profile of the offending call is small, because the cost lands on every other goroutine rather than on the sampler. The tell is that pause events line up with request volume rather than with allocation volume, and the fix is to move the measurement to a single reporting goroutine on a ticker and to use `runtime/metrics` for it. ## Summary One API buys consistency across a whole struct by pausing the program; the other buys cheapness by letting you name exactly the numbers you need and reading them without a pause. Sample regularly with the second, reach for the first deliberately and infrequently, and batch either one into a single call from a single goroutine.

  • So what does metrics.Read actually cost, if not a pause?
    An internal runtime lock plus the work to bring the aggregates your requested names depend on up to date. Plain counters are nearly free, heap aggregates cost more, histograms more again. Because the aggregation happens once per call, six names in one `Read` is cheaper than six calls of one name, and concurrent readers serialise on the lock — so sample from one goroutine.
  • Is there any reason left to call runtime.ReadMemStats at all?
    Yes, when you specifically want the `MemStats` shape — existing code or a dashboard keyed on its field names, or its own structures such as the fixed 256-entry `PauseNs` ring. For those you pay the stop-the-world, so read it on a slow ticker from one place rather than anywhere a request can reach.
  • How often is it safe to call runtime.ReadMemStats in production?
    Match it to whatever consumes the numbers — typically once every several seconds. The individual pause is short, but it is charged to every goroutine, so frequency is the whole risk. Per-request, per-loop-iteration, or inside a benchmark's inner loop are the three places it must never appear.
  • Does calling runtime.ReadMemStats trigger a garbage collection?
    No. It neither starts a GC cycle nor frees anything; it only reads and reconciles existing statistics. The confusion comes from the pause looking like GC pause time in a latency chart, which is exactly why the misdiagnosis is common.

ReadMemStats is a stocktake where the shop closes so every shelf is counted at the same moment. metrics.Read is asking two named shelves for their running totals while the shop stays open.

saying these in an interview costs you the question

  • Thinks ReadMemStats triggers or waits for a garbage collection
  • Calls ReadMemStats per request to record heap usage
  • Assumes metrics.Read is completely free of synchronisation
  • Believes ReadMemStats cost scales with heap size
  • Issues one metrics.Read call per metric name
open as a page

How do you read a value from Go's runtime/metrics package, such as the live goroutine count?

level: juniorimportance: should knowfreq 40%

basics

~10 s

Build a []metrics.Sample with each element's Name set to a metric name such as /sched/goroutines:goroutines, call metrics.Read on that slice, then switch on each Value's Kind to read a Uint64, Float64 or Float64Histogram.

open as a page

Your reporting goroutine samples runtime/metrics every 10 ms and p99 latency rose. How do you decide what to sample and how often?

level: seniorimportance: should knowfreq 32%

basics

~20 s

Sample no faster than something reads the numbers, usually the scrape interval in seconds rather than milliseconds, batch every name into a single metrics.Read call from one goroutine, and keep runtime.ReadMemStats off any request path because it stops the world.

open as a page

How do you turn /gc/pauses:seconds from runtime/metrics into a per-interval distribution?

level: middleimportance: nice to knowfreq 25%

basics

~10 s

Call Value.Float64Histogram, copy its Counts, and subtract the previous tick's Counts elementwise, because /gc/pauses:seconds is cumulative since process start. Buckets holds len(Counts)+1 boundaries, so Counts[n] covers the range from Buckets[n] to Buckets[n+1].

open as a page