skip to content

What must an exemplar on an aggregated metric carry, and when is it worthless?

level: seniorimportance: should knowfreq 34%

answer

  1. aggregates lose identity
  2. a pointer from a data point to one request
  3. identifier, timestamp, observed value
  4. keep the interesting example, not any example
  5. dead if the trace was dropped or expired

basics

~20 s

An exemplar points from one aggregated metric data point to a single concrete request behind it. It must carry a trace identifier, the timestamp and the observed value. It is worthless if that trace was never sampled or has aged out.

solid answer

~50 s

Aggregation destroys identity: a data point saying `41 requests fell in the 2s-to-4s bucket` names no request you can go and read. An **exemplar** is the deliberate escape hatch — one representative observation kept beside the aggregate, carrying the identifier of the trace it came from, the timestamp it happened at, and the raw value observed. Only a few are kept per data point, so which one is retained matters: an exemplar for a fast request tells you nothing about a slow tail. The link is half the mechanism. Following it needs the referenced trace to still exist, and two independent policies decide that: **trace sampling**, which may have discarded it at creation, and **trace retention**, which may have expired it since. An exemplar that outlives its trace is a dead link that looks live.

code

pseudocode · 9 lines
pseudocode
data_point:
  window        = 09:14:00 .. 09:14:15
  dimensions    = { service: checkout, route: /basket }
  bucket_upper  = 4.0 seconds
  count         = 41
  exemplar:
    trace        = a3f9c17b5d2e4088     # must still be stored to be followable
    observed_at  = 09:14:07.412
    value        = 3.87 seconds         # pick the interesting one, not the last one

go deeper

for a junior

Recall the shape of the idea: an aggregated number cannot name a specific request, so an exemplar is a small pointer attached to that number that leads to one real example behind it.

for a middle

Explain what has to travel with the pointer for it to work — an identifier for a stored trace, the moment it happened, and the value observed — and why only a handful are kept per aggregated data point.

for a senior

Demonstrate the operating knowledge: trace sampling and trace retention independently decide whether the target still exists, so exemplars should be emitted only for kept traces, and drill-down beyond the trace retention window is expected to fail.

for a principal

Own the cost argument. Exemplars are the only bridge from aggregate to instance, so whatever the telemetry budget, protect enough biased trace retention to keep that bridge alive for recent data rather than spreading a uniform rate too thin to be followable.

Metrics are cheap because they are aggregates: many observations collapse into a count, a sum, or a set of bucket counters, and the cost stops growing with traffic. The price of that collapse is identity. A data point telling you that 41 observations landed above two seconds knows nothing about which 41 requests they were. An **exemplar** is the deliberate repair for that: a small record attached to an aggregated data point that points at one concrete observation which contributed to it. ## What an exemplar has to carry For an exemplar to be followable at all, three things must travel with it: - **An identifier that can locate a stored trace.** Without it there is nothing to jump to. This is why exemplars are almost always emitted from code that is inside an active trace — the identifier is read from the request context at the moment the observation is recorded. - **A timestamp.** The aggregate covers an interval, so the exemplar must say when within that interval the observation happened, or the trace lookup has no window to search and the value cannot be placed on a chart. - **The observed value itself.** A pointer to a trace that ran in 1.9 seconds, attached to a spike you are investigating, is a wasted click. The value is what lets a reader decide whether this example is the one worth opening. Beyond the mechanics there is a selection question that decides whether the feature is useful. An aggregate covers an interval and a dimension set, and only a small number of exemplars are kept per data point — otherwise the aggregate stops being an aggregate. So **which** observation is retained is a design decision, and the only useful policy for latency work is to bias toward the interesting end: the slowest observation in the interval, or an errored one, rather than whichever arrived last. An estate that keeps an arbitrary example ends up with a link that is technically present on every point and useful on almost none. ## The two independent reasons an exemplar leads nowhere The link is a reference, not a copy. Following it requires the referenced trace still to exist, and two separate policies decide that: | Policy | When it acts | What it does to the exemplar | |---|---|---| | **Trace sampling** | at or shortly after the request runs | the trace was never stored, so the reference is dead from the moment it was written | | **Trace retention** | as the trace ages | the trace existed and has since expired, so the reference rots after a fixed age | Sampling is the sharper of the two, and the interaction is counter-intuitive. If a low, uniform, up-front sampling rate is in force, the overwhelming majority of requests are not stored — including almost all of the slow ones, because the decision was made before anyone knew they would be slow. Attaching an exemplar to every data point in that estate produces a wall of dead links. The two mechanisms that fix it are to **restrict exemplar emission to observations whose trace was actually sampled**, so a link is only written when there is something behind it, or to make the sampling decision late enough that slow and errored requests are kept preferentially. Retention is the quieter failure, and it is what bites when budgets move. Consider a vinyl-record marketplace whose telemetry budget was just halved, leaving a **96-hour floor on trace retention** while metric data points are kept for thirteen months. Every dashboard older than four days now carries exemplars whose targets are gone. Nothing is broken in a way that raises an alert; the links simply stop resolving, and engineers quietly learn not to click them. The rule of thumb worth stating in an interview is that **an exemplar is only meaningful within the retention window of the thing it points at**, so metric retention exceeding trace retention is normal and expected — what is not acceptable is presenting the older region as though the drill-down still works. ## Why this is the highest-leverage link in the stack The three signals form a chain: an aggregate shows you that something changed, a trace shows you where the time went for one request, and logs show you what that request's code actually said. The metric-to-trace step is the hardest of the three because it is the only one where the source has no identity of its own — logs and traces both carry identifiers naturally, and aggregates do not. The exemplar is the entire mechanism for crossing that gap. That also explains the honest limitation. An exemplar gives you **an** example, not a representative sample and not a population. It is a way in, not evidence. A candidate who says "the exemplar showed a slow database call so the database is the cause" has over-read a single observation; the correct reading is that the exemplar gives you one confirmed instance of the shape of the problem, which you then check against the aggregate to see whether it generalises.

  • Why does a low uniform up-front trace sampling rate make exemplars mostly useless?
    Because the decision is taken before the request's outcome is known, so the stored traces are a random few per cent and almost every slow or errored request is discarded. Exemplars written for unstored traces are dead links. The fixes are to emit an exemplar only when the observation's trace was actually kept, or to move the keep decision later so that slow and failing requests are preferentially retained.
  • Should exemplars carry the request's dimensions as well, or just the trace identifier?
    Just enough to locate and judge the example. The dimensions already live on the data point the exemplar hangs from, so repeating them adds volume without adding information. What is worth carrying is anything the trace lookup needs that the data point cannot supply — most importantly the precise timestamp, since the aggregate only knows the interval.
  • Metrics are kept for a year and traces for four days. How should a drill-down present that?
    Make the boundary visible rather than letting links silently fail. Inside the trace retention window the example is followable; outside it the aggregate is still a valid measurement and the pointer is not, so the reasonable behaviour is to stop offering the jump beyond that age. Presenting a dead link as a live one teaches people the whole feature is unreliable.

It is the photograph pinned to a monthly sales chart: the chart proves the month was unusual, the photo shows you one actual sale — and it is useless if the file it points to was deleted last week.

saying these in an interview costs you the question

  • Thinks an exemplar stores the trace rather than referencing it
  • Believes one exemplar is statistically representative of the aggregate
  • Attaches exemplars without checking the trace was sampled and kept
  • Assumes metric retention and trace retention can be reasoned about together
  • Keeps an arbitrary observation instead of the slow or errored one