skip to content

Explain what an exemplar attached to a Prometheus or Mimir time series is, everything that has to be enabled end to end for one to appear as a clickable dot on a Grafana time series panel, and why those clicks sometimes lead to a trace that no longer exists.

level: seniorimportance: should knowfreq 28%

answer

  1. Exemplar = one example observation attached to an aggregate
  2. OpenMetrics `# {trace_id=…} value ts` / OTLP transport
  3. Prometheus: opt-in feature flag + bounded circular buffer, not durable
  4. Datasource exemplars: internal link + labelName trace_id; time series panel, range query
  5. Dead link = head sampling dropped the trace, or trace retention < metric retention

basics

~30 s

An exemplar is a single sampled observation attached to a metric point, carrying labels such as a trace id. The instrumentation must record it, the exposition format must carry it, the metrics store must have exemplar storage enabled, and the Grafana data source needs an exemplar link to the trace store with the panel's exemplars option on. Clicks dead-end when the trace was never sampled or has already been dropped.

solid answer

~1 min

An **exemplar** is a concrete example observation attached to an aggregated metric sample — typically a histogram bucket increment — carrying its own label set, most usefully a `trace_id`. It is the bridge from *aggregate* to *individual*: the latency histogram says the 99th percentile blew up, the exemplar hands you one actual slow request. The chain has four links, all of which must hold: 1. **Instrumentation** records the exemplar with the active trace id when the observation happens. 2. **Transport**: the OpenMetrics exposition format carries exemplars after the sample (`# {trace_id="…"} 0.67 1600000000`), or they ride in OTLP. 3. **Storage**: Prometheus needs exemplar storage explicitly enabled (a feature flag plus a bounded in-memory circular buffer sized by a max-exemplars setting); Mimir has its own per-tenant exemplar limits. This buffer is **not durable** — it wraps. 4. **Grafana**: the Prometheus/Mimir data source has an exemplars entry with an internal link to the trace data source UID and the label name to use (`trace_id`), and the panel is a time series panel with the query's *exemplars* option turned on. Exemplars come back only for range queries. Dead links have two dominant causes: the request was **not sampled**, so no trace was ever exported even though the exemplar records its id; and **retention asymmetry**, where the trace has aged out of the trace store while the metric point remains.

code

text · 4 lines
text
http_request_duration_seconds_bucket{le="0.5",route="/checkout"} 1027 # {trace_id="4bf92f3577b34da6a3ce929d0e0e4736"} 0.47 1700000000.123

read as: this bucket has 1027 observations; here is ONE of them,
value 0.47s, from the trace with that id.

go deeper

for a junior

Define an exemplar as one example observation with a trace id attached to a metric point, and know that it makes a metric spike clickable.

for a middle

Walk the chain: instrumentation, OpenMetrics/OTLP transport, opt-in bounded storage, data source exemplar link plus the panel's exemplars option on a range query.

for a senior

Diagnose why dots are missing or dead — per-tenant limits, buffer eviction, head sampling, retention asymmetry, id encoding.

for a principal

Argue the policy coupling: tail sampling keyed on latency/errors plus retention alignment is what makes exemplars a reliable entry point rather than decoration.

## What an exemplar is Metrics are aggregates: a histogram tells you how many observations fell in each bucket, not which ones. That is exactly what makes metrics cheap and exactly what makes them useless for root cause. An **exemplar** is a small escape hatch built into the metric: alongside a sample, the instrumentation may attach one example observation with its own labels and its own timestamp. The canonical labels are `trace_id` (and often `span_id`). Semantically: *this bucket got incremented, and here is one request that did it.* The payoff is the classic RED-metric workflow. A panel shows p99 latency spiking at 14:32. Without exemplars you now go hunting in the trace store for a slow request near 14:32 for that service, which is guesswork. With exemplars you click the dot sitting on the spike and land in a trace that *provably* contributed to it. ## The four links of the chain **1. Instrumentation.** The library recording the observation must support exemplars and must have access to the current trace context at record time. Typically only histogram observations and counter increments carry them, and only when a trace is active. **2. Exposition.** In the pull model, exemplars require the **OpenMetrics** exposition format: the exemplar follows the sample on the same line, after a `#`, as a label set plus a value and optional timestamp. A scraper must negotiate OpenMetrics for them to survive. In the push model they travel in OTLP alongside the data point. **3. Storage.** In Prometheus, exemplar storage is **opt-in** — a feature flag enables it, and a configuration value sets the size of a **circular in-memory buffer**. Two consequences follow and both matter in interviews: exemplars are **not persisted to the TSDB blocks the way samples are**, and the buffer **overwrites oldest-first**. High-churn services can evict exemplars within minutes, so historical panels show none. Mimir applies its own per-tenant exemplar limits, which default to zero in some deployments — a very common reason nothing ever shows up. **4. Grafana.** The Prometheus/Mimir data source has an **Exemplars** section: add an entry, choose *internal link*, select the trace data source (stored as a **UID**), and set the **label name** that holds the id, usually `trace_id`. Then the panel: exemplars render only on the **time series** panel, and the query option **Exemplars** must be enabled. Grafana requests exemplars only for range queries; an instant query returns none. They appear as small diamonds on the x-axis band; hovering shows the labels, clicking follows the link. ## Why clicks dead-end - **Sampling.** This is the big one. If the tracing pipeline uses head sampling at, say, 1%, then 99 of every 100 exemplars point at a trace that was never exported. The exemplar records the id regardless — the metric path does not know the trace was dropped. Mitigations: sample tail-based on latency or errors so the interesting traces (exactly the ones exemplars point at) are always kept, or raise sampling for the services where this workflow matters. - **Retention asymmetry.** Metrics are cheap and kept for months; traces are expensive and kept for days. Every exemplar older than trace retention is a dead link by construction. Say so explicitly rather than treating it as a bug. - **Buffer eviction on the metrics side** — the mirror-image failure: the trace exists but no exemplar survived to point at it. - **Encoding mismatch** — the label carries the id in a different form than the trace store expects (dashes, wrong length, non-hex), so the lookup misses. ## How to pitch it The framing that separates a strong answer: exemplars are a **sampled, best-effort, bounded** pointer from an aggregate to an individual. They are not an index and were never intended to be complete — you cannot enumerate all requests in a bucket. Treat coverage as probabilistic, and design the sampling policy so that the traces exemplars point at are exactly the traces the tail sampler keeps. When those two policies are aligned the workflow feels magical; when they are not, the dots are decorative and users stop clicking them.

  • Panels show exemplar dots for one service and never for another. What do you check?
    Walk the chain for the silent service: does its instrumentation record exemplars at all, is it scraped with OpenMetrics negotiated (or pushed via OTLP), is the per-tenant exemplar limit non-zero for its tenant, and is any trace active when the observation is recorded — a code path outside a traced request has no id to attach. Also check volume: a high-throughput service can evict its exemplars from the bounded buffer before you look.
  • How would you make the exemplar links actually resolve most of the time?
    Align the sampling policy with the workflow: use tail-based sampling keyed on latency and errors so the slow and failed requests — precisely those an exemplar on a p99 spike points at — are always exported. Then align retention: either shorten the metric window you expect links to work in, or lengthen trace retention for the services that matter. Finally verify the id encoding matches between the exemplar label and the trace store's lookup format.

saying these in an interview costs you the question

  • Believing exemplars are stored durably alongside samples, rather than in a bounded in-memory ring
  • Expecting exemplars on an instant query or on a non-time-series panel
  • Treating exemplar coverage as complete, as though every request in a bucket is reachable
  • Blaming Grafana when the real cause is head sampling dropping the traces the exemplars name
  • Ignoring that trace retention shorter than metric retention guarantees dead links for older points

context