skip to content

Explain delta versus cumulative aggregation temporality in OpenTelemetry metrics: what each data point carries, how a consumer detects a counter reset, and how the choice interacts with a scrape-based versus a push-based backend.

level: seniorimportance: must knowfreq 36%

answer

  1. Cumulative: total since a fixed start_time; delta: this window only
  2. Delta windows tile — each start_time = previous time
  3. Lost cumulative point self-heals; lost delta point is gone
  4. Reset = value decreased or start_time changed
  5. Scrape backends want cumulative; push backends often want delta

basics

~20 s

Cumulative points report the running total since a start time; delta points report only what happened in the interval since the previous point. Consumers detect resets from a decreased value or a changed start timestamp. Scrape-based backends want cumulative; many push backends prefer delta.

solid answer

~60 s

Every sum and histogram data point carries a start timestamp, an observation timestamp and a value; **temporality** says how to read the value. Under **cumulative**, the value is the total accumulated since `start_time_unix_nano`, which stays fixed for the life of the series — so a lost point is self-healing, because the next point still carries everything. Under **delta**, the value covers only the window between the two timestamps and the windows tile without overlap — so a lost point is lost data forever, but the producer holds no long-lived state and can forget a series the moment it goes idle. Reset detection is the mirror image: a cumulative consumer infers a reset from a value that decreased or a start timestamp that changed, and must compensate when computing rates; delta has nothing to reset. Scrape-based systems are built on cumulative and require it. Push-based backends that sum server-side often prefer delta because it avoids unbounded producer-side state. The OTLP exporter has a temporality preference setting, and mixing temporalities within one series is the mistake to avoid.

code

text · 13 lines
text
requests: 5 in [0,10), 7 in [10,20), 4 in [20,30)

CUMULATIVE   start=0
  t=10 value=5    t=20 value=12   t=30 value=16
  drop the t=20 point -> t=30 still reports 16, nothing lost

DELTA
  [0,10) 5        [10,20) 7       [20,30) 4
  drop the [10,20) point -> those 7 are gone forever

RESET (cumulative, process restart at t=25, new start=25)
  t=20 value=12   t=30 value=4  and start_time changed 0 -> 25
  consumer must read this as a reset, not as -8

go deeper

for a junior

Define the two: cumulative is a running total since a start time, delta is only the last interval, and know that a counter reset shows up as the value dropping.

for a middle

Explain the timestamp semantics, the loss/duplication asymmetry, and which backend model each suits.

for a senior

Own reset detection via start timestamps, the producer-state versus delivery-guarantee trade-off, conversion cost, and the failure signatures on a graph.

for a principal

Set the fleet-wide default and justify it against workload churn, backend contract and memory budget, and state the mixing rule as a hard invariant with a plan for where conversion happens.

## The field and what it governs In OTLP, Sum and Histogram data points carry an `aggregation_temporality` field with values `DELTA` or `CUMULATIVE` (plus an unspecified value that consumers should treat as an error). Gauges have no temporality at all — a gauge reports a current value, and there is nothing to accumulate. Every point also carries two timestamps, `start_time_unix_nano` and `time_unix_nano`, and their meaning changes with temporality: - **Cumulative**: the value is everything accumulated from `start_time` to `time`. `start_time` stays constant across the whole life of the series. - **Delta**: the value covers only `start_time`→`time`, and each point's `start_time` equals the previous point's `time`. The windows tile end-to-end without gaps or overlap. ## Cumulative: stateful producer, resilient transport A cumulative series is a running total. Its great property is **idempotent recovery**: drop a point in transit and the next one still carries the missing increments, so a network blip costs resolution, not data. That makes cumulative naturally suited to a pull model, where the collector may scrape at irregular intervals or miss a scrape entirely. Its costs are producer-side state and reset ambiguity. The producer must remember every series' running total for as long as the series exists, so a series with many attribute combinations pins memory — and a series that stops being reported must eventually be forgotten, at which point resuming it looks like something new. When the process restarts, the total drops back toward zero: a consumer computing a rate must **detect the reset** and not report a huge negative rate. Detection uses two signals — the value decreased relative to the previous point, or `start_time` changed — and the second is the reliable one, which is why start timestamps are not optional decoration. ## Delta: stateless producer, fragile transport A delta series reports what happened in an interval and then forgets. The producer needs no long-lived accumulation, so memory is bounded by the series active in the current window and idle series simply stop being reported without any "is it a reset?" question. This suits environments with churn — short-lived workloads, functions, or very high attribute turnover — and it suits push backends that sum on the server side. The cost is that **every point is load-bearing**. Lose one and the total is permanently short by that amount; there is no later point that repairs it. Duplicate one — a retry that the backend does not de-duplicate — and the total is permanently over. That puts real weight on delivery guarantees and on the backend's de-duplication behaviour. Delta also makes a consumer's job of answering "what is the total since deploy?" a summation over time rather than a subtraction of two points. ## Conversion, and where it happens Cumulative → delta is straightforward: subtract consecutive points, handling resets. Delta → cumulative requires the converter to hold the accumulated state — exactly the state the producer avoided — so it moves the memory cost rather than removing it. This conversion is a normal function of a telemetry pipeline component when a producer and a backend disagree, and it is worth knowing that whoever converts pays. ## Backend fit **Scrape/pull-based systems** are built around cumulative counters: a scrape is a point observation of a running total, resets are detected by the query engine, and a missed scrape is harmless. Feeding them delta data requires an intermediate that accumulates. **Push-based hosted backends** frequently prefer delta, because the server aggregates across many senders and does not want each sender holding per-series totals — and because with ephemeral senders, cumulative series that vanish and reappear produce constant false resets. OpenTelemetry lets you choose: the OTLP metric exporter exposes a temporality preference (commonly configured as `cumulative`, `delta`, or a `lowmemory` option that mixes — typically delta for synchronous instruments while leaving asynchronous ones cumulative). The consequential rule is **do not mix temporalities within one series**: half the points meaning "since start" and half meaning "since last" produces numbers that are not wrong so much as meaningless, and backends generally cannot detect it. ## Practical failure signatures - Sawtooth graphs where a total repeatedly falls to zero: cumulative series restarting, either from real process restarts or from series being forgotten and re-created. - Enormous negative or spike rates at deploy time: a consumer failing to detect resets, or start timestamps not being set correctly. - Under-counted totals correlated with delivery problems: delta with lost points. - Steadily growing producer memory with flat traffic: cumulative with unbounded attribute cardinality — the temporality is not the bug, it is the amplifier. ## How to answer State the two definitions with their timestamp semantics, then the trade-off in one line each — cumulative buys transport resilience with producer state and reset handling; delta buys statelessness with sensitivity to loss and duplication — then match to backend model, and finish on the mixing rule and the memory-versus-cardinality interaction. That sequence is what separates a memorised definition from having operated it.

  • How exactly does a consumer distinguish a counter reset from a genuine decrease?
    A cumulative counter must never decrease, so any decrease is by definition a reset — but relying on that alone breaks when the counter resets and climbs past its old value between two observations. The reliable signal is the start timestamp: when it changes, the accumulation window restarted, so the new value must be treated as counting from zero rather than compared with the previous point.
  • You must send delta data to a backend that only understands cumulative. Where does that conversion cost land?
    On whichever component converts. Delta-to-cumulative requires holding a running total per series for as long as the series is alive, which is precisely the state the delta producer was avoiding. Doing it in a central pipeline component concentrates that memory in one place — sized by total series across all senders — so it must be capacity-planned against cardinality, and restarts of that component look like resets downstream.
  • Why is a gauge exempt from temporality?
    Because a gauge reports the value at an instant rather than an amount accumulated over a window. There is nothing to sum across points and nothing to reset — the last observation simply wins. That is also why gauges cannot answer questions like 'how much in total', and why anything you might want to rate should be a counter instead.

saying these in an interview costs you the question

  • Believing cumulative counters lose data when a data point is dropped
  • Thinking a decreasing cumulative value is always a bug rather than a reset
  • Ignoring start timestamps and reset detection when computing rates
  • Mixing delta and cumulative points within one series
  • Assuming delta-to-cumulative conversion is free — it just relocates the state

context