skip to content

A per-request question answered out of aggregate metrics: how do you recognise the trap, and what should have been emitted?

level: seniorimportance: nice to knowfreq 26%

answer

  1. Notice when the evidence is a shape
  2. Co-movement in time is not identity
  3. Aggregation is a one-way transformation
  4. New dimensions only collect going forward

basics

~20 s

You are in the trap when your evidence is a shape rather than a record: inferring one request's fate from a bump in a curve. The fix is at emit time, with records that carry the identifiers you will later filter by.

solid answer

~50 s

Three tells. First, your evidence is a **shape, not a record** — you say the request "probably timed out" because a latency boundary grew, not because you found it and read it. Second, you are **joining by timestamp**: two curves moved together at 09:14, so you assume they describe the same requests, which no aggregate can establish. Third, you are **adding a dimension mid-incident** to slice, which rebuilds a per-request record one field at a time and only collects from the deploy onward. The cost is an answer of the wrong kind: an inference presented as a measurement, arriving after a code change rather than from data you already hold. What should have been emitted is one record per request carrying the fields you would want to filter by — outcome, route, tenant, duration — with the identifier intact even where volume forces you to keep only a fraction.

go deeper

for a junior

Recall that a counter tells you how many, never which. If a question names one request or one customer, an aggregate cannot answer it however you slice the query.

for a middle

Explain why aggregation is one-way and what that costs: identity and any un-promoted field are gone for that time range, so adding a dimension later fixes the next incident and not this one.

for a senior

Show the self-awareness during a live investigation: notice when you have started inferring, say so explicitly, and name the emit-time change that would have made the answer a fact instead of a story.

for a principal

Own the trade honestly. Per-request detail is not free, so decide which questions are worth paying to be able to answer, and make sure the volume strategy keeps whole identifiable requests rather than stripping them to save space.

## The signal you have is the signal you use Investigations drift toward whatever store the responder is fluent in. If a team's aggregates are excellent and its per-request records are thin, every question gets answered out of aggregates — including the questions aggregates structurally cannot answer. The output is not a wrong number. It is a *plausible story*, delivered with the confidence of a measurement, and it survives into the incident review as though it were a finding. Aggregation is a one-way transformation. Folding ten thousand requests into a count, a sum and a set of bucket boundaries deliberately discards which requests they were. Every property of an individual request that was not promoted to a dimension is gone at that moment, permanently, for that time range. ## Three tells that you are in the trap 1. **Your evidence is a shape, not a record.** You are saying "it probably hit the client timeout" because a boundary in a latency distribution grew, not because you found a request that timed out and read it. The word *probably* is the tell. 2. **You are joining by timestamp.** Two curves moved in the same minute, so you conclude they describe the same requests. No aggregate can establish that. Co-movement at one-minute resolution is equally consistent with the same requests, with two independent effects of one cause, and with coincidence — and at high volume, coincidence is routine. 3. **You are adding a dimension mid-incident so you can slice.** That is rebuilding a per-request record one field at a time, under time pressure, and it buys nothing for the incident in progress: the new dimension starts collecting at the deploy, so the history you actually need is still missing. A softer fourth tell: you catch yourself asking "how many of those 6,180 failures were this one tour's?" and no query you are able to write will say. ## What an aggregate genuinely cannot do | The question | Answerable from aggregates? | Why | |---|---|---| | "How many failed in that window?" | Yes | This is exactly what a counter is for | | "Did failures rise after 09:14?" | Yes | A trend over a population is the aggregate's job | | "Which requests failed?" | No | Identity is what aggregation throws away | | "Was this customer among them?" | No | Not unless the customer was already a dimension | | "What were the inputs on the slow ones?" | No | Per-event detail never survives into a summary | | "Is the slow set the same as the failing set?" | No | Two summaries cannot be joined on identity | ## The cost, priced honestly The cost is not that you get no answer. It is that you get an answer of the wrong *kind*: - It is an inference, so it is wrong at a rate nobody measures, and everything downstream inherits the error. - It arrives late — after a code change and a wait — instead of coming out of data you already hold. - It cannot be narrowed. The moment someone asks "does it only affect tours routed through a second jurisdiction?", you are back at the same wall. - It quietly encourages the worst fix available: hanging a high-cardinality identifier onto a metric so that it can be sliced, which is how teams melt a metrics store. ## What should have been emitted The correction lives at emit time and is unglamorous: 1. **One record per request carrying the fields you would want to filter by** — outcome, route, tenant or tour identifier, the upstream called, duration, client version. Choosing that list is a design act, and the honest heuristic is "what would I have wanted to group by during the last three incidents?" 2. **A shared identifier on every record belonging to the same request**, so one hop leads to the rest instead of to a timestamp search. 3. **Keep that identifier off the metric.** The tour or customer identifier belongs on the per-request record, not as a dimension on a counter. 4. **A volume strategy that preserves identity.** If you cannot keep every record, keep a fraction — but keep whole requests rather than stripped ones, and keep the failures unconditionally. A record with its identifying fields removed to save space answers nothing later. ## When answering from aggregates is correct Not every question deserves a per-request record, and over-instrumenting is its own failure. If the decision in front of you is "roll back or not", the blast radius from an aggregate is often the whole answer and hunting for one example request is a detour. The trap is not using aggregates. It is using them for a question about identity and then reporting the result as if it were established fact. Saying which one you are doing is what a senior responder does: "I have not found the request; the shape is consistent with a timeout" is a stronger sentence in an incident channel than a confident guess.

  • Two curves moved in the same minute during an incident. What would it take to show they describe the same requests?
    Records of the individual requests carrying both properties, so you can count the ones that have each and the ones that have both. That is a join on identity, and it is only possible where identity survived. From two summaries you can show co-movement and nothing more, which at high volume happens by coincidence often enough that acting on it is a real risk.
  • You add a new dimension to a metric mid-incident. What are you buying, and what are you not?
    You are buying the ability to answer the same question the next time it happens. You are not buying anything for the incident in front of you, because collection starts at the deploy and the history you need is already lost. You may also be buying a long-term cost, since a dimension with many distinct values multiplies the series a metric produces.
  • When is answering from aggregates alone genuinely the right call?
    When the decision only needs blast radius. If the question is whether to roll back, knowing that failures are confined to one route and one region is often sufficient, and hunting for a single example request is a detour. The discipline is to say which kind of answer you have, rather than to always dig further.

saying these in an interview costs you the question

  • Presents co-movement in time as proof of causation
  • Believes a percentile can be filtered to one caller afterwards
  • Adds a high-cardinality identifier as a metric dimension
  • Calls an incident explained without finding one example
  • Assumes retention covers a question asked days later
  • Reports an inference in the same register as a measurement