skip to content

Which aggregations across a fleet of per-process gauges are meaningful, and why is an average of averages wrong?

level: middleimportance: should knowfreq 48%

answer

  1. Ask what the number physically is
  2. Some quantities add up; ratios do not
  3. Each process carries a different weight
  4. Sum numerator and denominator, then divide

basics

~20 s

It depends on the quantity. Gauges of things that add up — queued items, bytes in use — can be summed across a fleet; ratios and percentages cannot. A mean of per-process ratios ignores that each process handled a different volume.

solid answer

~50 s

It depends on what the gauge physically measures. **Extensive** quantities — bytes of heap in use, open connections, queued items — add up, so a fleet sum is meaningful and min, max and mean are all useful for spotting an outlier process. **Intensive** quantities — utilisation percentages, hit ratios, mean sizes — do not add up, and summing them yields a number with no physical meaning. Averaging them across processes is worse than useless, because each process's value carries a different weight: a process serving 3,860 requests per second and one serving 118 contribute equally to an unweighted mean, so a fleet that is actually struggling reads as healthy. The fix is to stop publishing the derived value at all — publish the numerator and denominator as counters, sum each across the fleet, and divide once at query time.

code

pseudocode · 8 lines
pseudocode
# wrong: the ratio is computed inside each process and can never be recombined
gauge cache_hit_ratio = hits / lookups
fleet = mean_across_processes(cache_hit_ratio)      # 0.94, and wrong

# right: publish the parts, divide once at query time
counter cache_hits_total
counter cache_lookups_total
fleet = sum(rate_over(cache_hits_total, 5m)) / sum(rate_over(cache_lookups_total, 5m))

go deeper

for a junior

Recall that a gauge is a value at an instant, and that adding up bytes or queued items across machines makes sense while adding up percentages does not. Being able to spot that a summed percentage is nonsense is enough at this level.

for a middle

Explain the extensive-versus-intensive split and work through why an unweighted mean of per-process ratios is not the fleet ratio. Show the fix: publish the numerator and denominator separately and divide after summing.

for a senior

Demonstrate that you have been misled by such a panel. Talk about fleet size changing the sum during a deploy, about max versus mean for saturation, and about insisting that derived values are never published pre-divided.

for a principal

Own it as a convention across teams: services publish raw parts, dashboards state their aggregation in the panel title, and derived-value gauges are treated as a review defect. The cost of getting this wrong is an entire organisation trusting a green number.

## A gauge is a point sample of a level A gauge records where something stood at the instant it was observed. It accumulates nothing, so unlike a counter it carries no memory of what happened between one observation and the next, and that shapes every aggregation you can build on top of it. When the same gauge is published by every process in a fleet, you hold a grid: one value per process per timestamp. Collapsing that grid into a single fleet number is where most dashboards quietly go wrong, because the collapse that is arithmetically available is not always the collapse that means something. ## Which fleet aggregations mean something The test is not about the metric system; it is about the physical quantity. Ask whether the thing being measured **adds up**. - An **extensive** quantity is proportional to the size of what you measured: bytes of heap in use, open connections, items waiting in a queue, threads alive. Two processes holding 400 MB each are genuinely holding 800 MB. Summing across the fleet is meaningful, and so is the mean, the maximum and the minimum. - An **intensive** quantity is a ratio or an intensive property that does not scale with size: utilisation percentage, a cache hit ratio, a mean object size, a temperature. Summing fourteen hosts each at 46% utilisation produces 644%, a number with no physical referent at all. | Aggregation | Extensive gauge (queued items) | Intensive gauge (hit ratio) | |---|---|---| | Sum across processes | Meaningful: total work waiting | Meaningless: no physical referent | | Mean across processes | Meaningful: typical per-process level | Misleading unless every process carries equal weight | | Max / min across processes | Finds the outlier process | Finds the outlier process | | Count of processes over a threshold | Tells you how widespread it is | Tells you how widespread it is | Two further cautions apply to both columns. First, a fleet sum moves when the fleet size moves, so a rolling deploy that briefly runs eleven processes instead of fourteen produces a dip that is an artefact of the deploy, not of the workload; wherever that matters, publish the process count beside the sum so the two can be read together. Second, a max over a window is the maximum of the samples you happened to take, not the true maximum of the underlying level — a spike between two observations leaves no trace at all. ## Why an average of averages is not the average The most common intensive-quantity mistake has its own name. When each process publishes a value that is *already* a ratio or an average, taking the unweighted mean of those values across the fleet gives every process an equal vote regardless of how much work it did. That is only the fleet figure in the special case where the workload is perfectly even, which is precisely the case that stops holding during an incident. The correct fleet ratio recombines the parts: 1. Have each process publish the **numerator** and the **denominator** separately, as counters — cache hits and cache lookups, not the ratio. 2. Sum each of those across the fleet for the window you care about. 3. Divide once, at query time, at the level you want to report. That way the arithmetic weights each process by its actual volume, and the same two series still support a per-process breakdown when you want one. ## A worked example A seed-catalogue ordering service ran fourteen processes behind a load balancer and each published a `cache_hit_ratio` gauge. During a 5,400-request-per-second peak, a sticky-session bug pinned one process to 3,860 requests per second while the other thirteen shared the remaining 1,540. The busy process, thrashing its cache, sat at a 0.61 hit ratio; the thirteen quiet ones sat at 0.96 each. The fleet dashboard averaged the fourteen gauges and showed **0.94** — a healthy-looking green number that stayed green for the whole incident and that nobody could reconcile with what users were experiencing. The true fleet ratio weights by volume: `(3860 x 0.61 + 1540 x 0.96) / 5400`, which is about **0.71**. More than a quarter of all lookups were missing, and the dashboard was off by twenty-three percentage points. Nothing was broken in the metrics pipeline; the ratio had been divided too early, inside each process, and once divided it could not be recombined. ## What to publish instead - **Publish the raw parts, divide at query time.** Any value you compute inside a process and then publish as a single number can no longer be re-aggregated. Ratios, percentages and means all have this property. - **Keep an identifying dimension on the series** so the per-process values are still reachable when you want to find the outlier — the fleet number and the outlier hunt need different aggregations. - **Say the aggregation out loud on the dashboard.** "Mean of per-process hit ratio" and "fleet hit ratio" are different numbers, and a panel titled just "hit ratio" invites the reader to assume the one that would have been correct. - **Prefer max over mean for saturation.** One process at 98% utilisation is an incident even when the fleet mean is 40%, and the mean is designed to hide exactly that.

  • Which fleet aggregation would you choose to catch one bad process rather than a fleet-wide problem?
    The maximum across processes, plus a count of processes past a threshold. A mean is designed to dilute a single outlier: one process at 98% utilisation inside a fleet of fourteen barely moves the mean, while the max shows it immediately and the count tells you whether it is one process or spreading.
  • A fleet sum of a memory gauge dips during a rolling deploy. Is that a real drop in memory use?
    Partly an artefact. The sum is over however many processes were publishing at that timestamp, so replacing processes one at a time removes terms from the sum for as long as each gap lasts. Publishing the process count alongside the sum, or reading the mean per process, separates a genuine change in memory use from a change in fleet size.
  • Is a mean across processes ever the right fleet number for an intensive gauge?
    Only when the processes genuinely carry equal weight and you say so — an even shard assignment, or a fleet where each process handles a fixed slice. Even then it is fragile, because the assumption stops holding exactly when traffic skews, so a weighted computation from published parts is the safer default.

Averaging fourteen processes' hit ratios is like averaging fourteen players' batting averages to get the team's: it is only the team figure if every player had the same number of at-bats.

saying these in an interview costs you the question

  • Averages per-process ratios and calls the result the fleet ratio
  • Sums a utilisation percentage across hosts and reports 644%
  • Assumes a fleet sum is valid for every gauge regardless of the quantity
  • Ignores that a rolling deploy changes how many terms the sum has
  • Publishes a pre-divided ratio with no way to recombine it later
  • Uses a fleet mean to look for a single saturated process