What does a metric aggregate lose that a per-request record keeps, and when is losing it the right trade?
answer
- Aggregation is one-way
- It keeps counts, not identities
- New questions need new instrumentation
- Quantiles do not merge across processes
- Flat query cost is the payoff
basics
~20 sAggregation discards identity, every attribute nobody made a dimension, and the ability to re-slice history: it keeps how many, never which ones. It stays the right trade when the question is a rate or trend answered at fixed cost.
solid answer
~50 sFolding an event into an aggregate is a one-way door. Three things go: the **identity** of the contributing events, every **attribute that was not made a dimension in advance**, and so the ability to slice history along a dimension that did not exist when it was written. A fourth loss is mathematical: a quantile computed inside one process cannot be combined with another process's quantile, because quantiles are not additive. It is nonetheless the right trade whenever the question is about the population rather than an instance. Aggregates give a complete denominator covering traffic whose per-event telemetry was thinned, a query cost that stays flat as the window grows, an evaluation fast enough to alert on, and a shape that carries no per-user fields. Fold what you will ask every minute; keep events for the questions you have not thought of yet.
code
pseudocode · 11 linesone renewal, kept as a record:
{plate: "KX21-ZTM", office: "north", channel: "kiosk",
vehicles: 2, ms: 412, outcome: "ok"}
the same renewal, folded into aggregates:
renewals_total{office: "north", outcome: "ok"} += 1
renewal_ms_sum{office: "north"} += 412
still answerable next month: how many succeeded at the north office
lost at the moment of the fold: that this plate, with two vehicles,
took 412 ms through the kiosk channelgo deeper
Recall that an aggregate keeps how many, not which ones, and that the detail is dropped when the number is emitted rather than later. Knowing that much prevents the most common wrong assumption at this level.
Explain the three losses precisely: identity, attributes that were never made dimensions, and therefore the ability to re-cut old data. Be able to say why a dimension added today produces no history.
Show judgement about where the line sits. Argue for the aggregate on the grounds of a complete denominator and flat query cost, and give a case where you deliberately kept per-event fidelity because the question was not known in advance.
Own the boundary as a design decision with a cost attached. Be ready to defend a policy on what gets folded and what stays as events, and to explain why widening dimensions is the wrong answer to a fidelity complaint.
## Aggregation is a one-way door When a request is folded into an aggregate, the fold happens **at emission**, inside the process, before anything reaches storage. What is written is a number attached to a name and a set of dimension values. The request that contributed to it is not written anywhere, is not compressed, is not archived — it simply never existed as a stored thing. That is why no retention setting, no re-indexing job and no cleverer query recovers it. Four distinct things are lost, and it is worth separating them because they have different consequences. 1. **Identity.** The aggregate says 4,180 renewals completed; it cannot say that permit KX21-ZTM was one of them. Any question of the form "which one" is unanswerable by construction. 2. **Unchosen attributes.** Only what was made a dimension survives. If payment channel was not a dimension, then for every past sample the channel is not merely hard to query — it was never captured. 3. **Re-slicing history.** This follows from the second and is the one that bites in practice. Adding a dimension today starts producing data today. Last month cannot be re-cut, so the answer to a question first asked on a Tuesday is available from Tuesday onwards. 4. **Combinability of quantiles.** A quantile computed inside one process describes that process's own distribution. Averaging the p99 of twelve processes does not give the fleet p99, because a quantile is not additive: the true value depends on the merged distribution, and one heavily loaded process can dominate it. Counts and sums merge across a fleet; quantiles computed at the edge do not. | Question | Answerable from an aggregate | Answerable from per-event records | |---|---|---| | How many renewals failed last Tuesday? | yes, cheaply, over any window | yes, by scanning everything | | Which renewals failed, and what did they carry? | no | yes | | Split last month by a dimension added today | no | yes, if the field was in the record | | Fleet-wide p99 from per-process quantiles | no | yes, from the merged events | | A number to alert on within seconds | yes | expensive and slow | ## When folding is nonetheless right The loss is real and the trade is still usually correct, for reasons that have nothing to do with laziness. - **The denominator is complete.** An aggregate observes every event, including those whose per-event telemetry was thinned or never kept. A rate computed from records that survived a keep decision is a rate over survivors, which is not the same number. - **Query cost is flat in the window.** Reading a year of an aggregate is a bounded amount of work. Deriving the same answer from records means reading a year of records. - **It is fast enough to act on.** An alert has to evaluate every interval, forever. That is only affordable against pre-aggregated numbers. - **It is stable to look at.** Trends, seasonality and slow regressions are visible in an aggregate and invisible in a firehose of individual events. - **It carries no per-subject detail.** An aggregate over a bounded dimension set is one of the few telemetry shapes you can retain for a long time without holding anything about an identifiable person. ## A concrete case The municipal parking-permit service runs to a 320 ms p99 budget with an on-call rotation of three people. Every renewal is folded into counters and a latency distribution keyed by office, outcome and channel — 210 combinations, cheap, and enough to know within thirty seconds that the budget is being missed. On a Thursday the p99 sits at 604 ms for two hours. The aggregate says which office and which channel, which is enough to page correctly and to know the blast radius. It cannot say which renewals were slow, and when someone asks whether slow renewals correlated with permits that had more than one vehicle attached, the answer is that no such dimension exists and last Thursday cannot be re-cut. That question is answerable only from the records for the renewals whose per-event telemetry was retained, and only if the vehicle count happened to be a field on them. The lesson is not "add more dimensions". Adding vehicle count as a dimension would have multiplied the series count for a question asked once. The lesson is that the two shapes are answering different classes of question and the boundary between them should be chosen deliberately. ## Deciding what to fold - **Fold what you will ask on a schedule** — the numbers that go on an alert or a status view, asked every interval, over long windows. - **Keep as events what you will ask once** — the exploratory question, the one-off correlation, the field you cannot enumerate in advance. - **Never widen a dimension set to preserve per-event fidelity.** A dimension whose value set is large is a fidelity dream and a cost catastrophe; that detail belongs on a record. - **Emit the aggregate and the record from the same code path** so the number you alert on and the events you investigate cannot disagree about what happened. - **Treat a new dimension as a forward-only change** and say so when someone asks for history: they are asking for data that was never created.
- Someone asks on Tuesday for last month's latency split by payment channel, and channel was never a dimension. What do you tell them?That the aggregate cannot answer it and never will, because the channel was not captured when those samples were written. Add the dimension now if its value set is small, and say plainly that data starts today. If per-event records for that window were retained and carried the channel as a field, the answer is there instead — at the cost of scanning them.
- Why can you not average the p99 of twelve processes to get the fleet p99?Because a quantile is a property of a distribution, not a quantity that adds. The fleet p99 is the 99th percentile of the merged population, which depends on each process's whole distribution and on how much traffic each served. A single busy or degraded process can move it far from the mean of the twelve. Merging requires a structure that can be combined, or the events themselves.
An aggregate is a census tally: it will tell you how many households own three vehicles, and no amount of re-reading it will ever tell you whose.
saying these in an interview costs you the question
- Believes stored aggregates can be re-sliced by a dimension added later
- Averages per-process percentiles to obtain a fleet percentile
- Thinks keeping every event makes aggregates unnecessary
- Adds an unbounded dimension to preserve per-event fidelity
- Assumes lost detail can be recovered by re-processing storage