Your reporting goroutine samples runtime/metrics every 10 ms and p99 latency rose. How do you decide what to sample and how often?
answer
- who actually reads these numbers
- the observer can outweigh the observed
- one goroutine, one batched call
- not every name costs the same
- unsampled numbers do not exist afterwards
basics
~20 sSample no faster than something reads the numbers, usually the scrape interval in seconds rather than milliseconds, batch every name into a single metrics.Read call from one goroutine, and keep runtime.ReadMemStats off any request path because it stops the world.
solid answer
~50 sStart from the consumer: numbers nobody reads more than once every fifteen seconds do not need reading every ten milliseconds, so the interval is the largest lever. Then make each tick cheap — one `metrics.Read` with every name in one reused slice, from exactly one goroutine, because `Read` takes a process-wide lock. Grade the names by cost: plain counters are nearly free, heap aggregates cost more, histograms cost more again and add per-tick copy and diff work. `runtime.ReadMemStats` comes off the hot path entirely; its stop-the-world is charged to every goroutine, which is how instrumentation surfaces as tail latency rather than as its own CPU. To confirm rather than guess, take a `go tool trace` and see where the sampler runs. And choose the always-on set deliberately: after an incident you only have what you were already collecting.
code
go · 13 linessamples := []metrics.Sample{
{Name: "/sched/goroutines:goroutines"},
{Name: "/gc/heap/allocs:bytes"},
{Name: "/gc/cycles/total:gc-cycles"},
{Name: "/gc/pauses:seconds"},
}
tick := time.NewTicker(15 * time.Second)
defer tick.Stop()
for range tick.C {
metrics.Read(samples) // one lock acquisition, shared aggregates computed once
report(samples)
}go deeper
Know that sampling has a real cost and that the interval should match whatever collects the numbers. Reading runtime statistics thousands of times a second is not free just because each read looks small.
Explain why one batched metrics.Read from a single goroutine beats several small reads: one lock acquisition, one shared aggregation. Be able to say which kinds of names are cheap and which are not.
Show how you would prove it: a trace over a window, then bisecting on interval and on metric list. Expect to be asked why the sampler's own CPU profile looks innocent while the pauses land elsewhere.
Own the always-on set as a standing decision with a cost attached, and make the split between permanent and on-demand telemetry explicit — because after an incident, the data you were not collecting is simply gone.
## The shape of the problem A sidecar goroutine that wakes on a ticker and reports the process is the standard way to expose runtime numbers. It is also one of the few pieces of code in a service whose entire cost is overhead — it produces nothing a user asked for. So its budget deserves an explicit decision, and the failure mode when it does not get one is characteristic: the observer costs more than what it observes, and the cost lands on unrelated goroutines as tail latency rather than on the sampler as CPU. ## Lever one: the interval Ask who consumes the numbers. If something collects every 15 seconds, sampling at 10 ms produces 1499 readings out of every 1500 that are thrown away, and pays the full cost for each. The interval is not a quality knob — a faster sampler does not make the numbers more accurate for a consumer that reads slowly; it only makes them more expensive. Match the ticker to the consumer, and treat sub-second sampling as something you turn on deliberately for a bounded investigation, not as a default. ## Lever two: one batched call from one goroutine `metrics.Read` computes the aggregates that the requested names depend on once per call, under a process-wide lock. That gives two rules. Put every name you want into a single `[]metrics.Sample`, built once at start-up and reused every tick, so you take the lock once and share the aggregation. And sample from exactly one goroutine, so that your reporter, a debug endpoint and a health check are not contending on the same lock at three different intervals. ## Lever three: know what each name costs The names are not interchangeable in price. A simple counter such as a GC cycle total is close to free. Metrics derived from heap or memory-class aggregates require the runtime to reconcile per-P state, so they cost more. Histograms cost the most: the runtime produces bucket counts, and your reporter then copies and diffs them every tick. If your always-on set contains a dozen histograms sampled every second, the per-tick work is no longer negligible — and none of it is visible as a slow request, which is why it survives so long. And `runtime.ReadMemStats` is in a category of its own: it stops the world. One call on a slow ticker is invisible; a call per request turns instrumentation into a latency generator whose profile looks small, because the pause is paid by everyone else. ## Confirming rather than guessing The honest way to test the hypothesis is to look at what the runtime was doing between your sample points. `go tool trace` over a short window shows the sampler goroutine's executions in context — how often it runs, how long it holds the processor, and whether its ticks line up with the latency events you are chasing. Two cheap experiments alongside it: raise the interval by an order of magnitude and see whether the p99 recovers, and cut the metric list to the counters only and see whether it recovers further. Between the trace and those two bisections you get a cause rather than a correlation. ## The always-on set is a decision, not a default The part that matters most in hindsight: the reason to keep a cheap always-on set at a slow interval is that after an incident, the runtime does not retroactively remember anything. Whatever you were not sampling simply does not exist for the postmortem, and no amount of clever analysis recovers it. So the sensible split is a small, permanent set — goroutine count, allocation totals, GC cycle counts, one or two distributions — sampled at the collection interval, plus a heavier set exposed on demand behind an endpoint for when someone is actively investigating. Write that split down. In a postmortem the most useful sentence is often "we were sampling these seven names every fifteen seconds, and nothing else", because it tells the next reader exactly which conclusions the data can support and which ones are speculation. ## Summary Budget the observer explicitly: interval matched to the consumer, one batched read from one goroutine, a metric list graded by cost, no world-stopping call anywhere near a request, and a deliberate always-on set chosen before you need it rather than after.
- How would you confirm the sampler is the cause rather than a coincidence?Take a `go tool trace` over a window and look at where the sampling goroutine runs relative to the runtime's own activity and to the slow requests. Then bisect: raise the interval tenfold and see if p99 recovers, then cut the list to plain counters and see if it recovers further. Trace plus two bisections gives a cause, not a correlation.
- What belongs in the always-on set versus behind a debug endpoint?Always-on: a handful of cheap, high-value names at the collection interval — goroutine count, allocation totals, GC cycle counts, maybe one distribution. On demand: expensive or verbose things, including anything using `runtime.ReadMemStats` and any fine-grained interval, exposed behind an endpoint someone hits while investigating. The permanent set is chosen for postmortem value, not for completeness.
- Two subsystems each sample runtime metrics on their own ticker. What is the objection?They contend on the same process-wide lock inside `metrics.Read` and duplicate the aggregation, at two unrelated intervals nobody reconciles. One sampling goroutine that reads the union of the names and fans the values out to both consumers costs less and gives both a consistent instant.
Weighing a parcel every ten milliseconds when the courier collects once a day does not make the weight more accurate; it just means you spend the day at the scales.
saying these in an interview costs you the question
- Treats a faster sampling interval as more accurate data
- Samples runtime metrics from several goroutines independently
- Assumes all runtime/metrics names cost the same to read
- Leaves ReadMemStats on a request path for heap numbers
- Expects to raise resolution retroactively during an incident