OpenTelemetry metrics offer counters, up-down counters, gauges and histograms, each in a synchronous and an asynchronous (observable) form. How do you choose between them, and what does "asynchronous" actually change?
answer
- Monotonic total → Counter; can decrease → UpDownCounter
- Need percentiles → Histogram; current value → Gauge
- Async callback reports the ABSOLUTE value, never a delta
- Sync gives per-event attributes and exemplars
- Every attribute combination = one time series
basics
~20 sCounter for monotonic totals, up-down counter for values that rise and fall, histogram for distributions you want percentiles from, gauge for a current value. Synchronous instruments are called where the event happens; asynchronous ones are read by a callback at collection time.
solid answer
~60 sPick by the shape of the quantity. **Counter** — a total that only increases (requests handled, bytes sent); you record increments and the backend derives rates. **UpDownCounter** — a quantity that goes both ways (queue depth, active connections, items in a cache). **Histogram** — you need the distribution, not just the total: latency, payload sizes; the SDK buckets the values so percentiles are computable. **Gauge** — a current value that is measured rather than accumulated (temperature, heap in use, a configured limit). The synchronous/asynchronous split is about *when* the value is produced. Synchronous instruments are invoked inline at the event, so they can attach attributes from the request context and can feed exemplars linking a data point back to a trace. Asynchronous (observable) instruments register a callback that the SDK invokes once per collection interval and which reports the **absolute current value**; the SDK computes any needed deltas. Use asynchronous when the value already exists somewhere to be read and it would be wasteful or impossible to instrument every change — the classic examples being process and OS statistics.
code
text · 9 linesWRONG (observable counter callback):
observe(bytes_since_last_call) # SDK will accumulate again -> double counting
RIGHT:
observe(total_bytes_since_start) # SDK derives the delta if the exporter wants one
Rule of thumb:
sync add(+1) increments
async observe(1_204_338) absolute readinggo deeper
Map the four kinds to plain examples — requests as a counter, queue depth as an up-down counter, latency as a histogram, memory as a gauge.
Explain the synchronous/asynchronous distinction correctly, especially that async callbacks report absolute values, and know that histograms carry count and sum.
Add exemplars, callback constraints, the redundancy of a separate counter next to a histogram, and cardinality as the real cost driver.
Treat instrument choice as a cost and interpretability contract across teams: which dimensions are permitted on metrics, what belongs in traces instead, and how instrument choices constrain future queries you cannot retro-fit.
## Two independent choices Choosing an instrument is two decisions, not one: **what kind of quantity is this** (counter / up-down counter / histogram / gauge), and **how is the value produced** (synchronously at the event, or asynchronously by a callback). Candidates who conflate them end up with an observable histogram-shaped idea that does not exist, or a synchronous gauge where a counter belongs. ## The kinds **Counter** — monotonically increasing. You call `add(n)` with a non-negative `n`. Requests served, errors, bytes written, messages published. You never record the running total yourself; the SDK accumulates. Backends turn counters into rates, which is what you almost always want to graph. Recording a total that can decrease into a counter corrupts rate calculations, because consumers interpret a decrease as a restart. **UpDownCounter** — same additive idea, but `add(n)` accepts negatives. Queue depth, in-flight requests, open connections, cache entries. The distinguishing test is whether the quantity can go down for reasons other than a process restart. **Histogram** — you call `record(value)` and the SDK aggregates values into a distribution. Use it when the interesting questions are "what does the tail look like" — request duration, response size, batch size. A counter of total duration divided by a counter of requests gives you a mean, and means hide exactly the behaviour you are looking for. **Gauge** — a value that is *sampled*, not accumulated: the last observation wins, and there is no meaningful sum across time. Current heap bytes, temperature, a configured queue capacity. A gauge's data point tells you the value at an instant; it says nothing about what happened between observations, which is why a gauge is the wrong instrument for anything you might want to rate or sum. ## Synchronous versus asynchronous **Synchronous** instruments (Counter, UpDownCounter, Histogram, and a synchronous Gauge in recent specification versions) are called from application code at the moment the thing happens. That gives them two properties nothing else has: they can attach attributes derived from the current operation (route, status code, tenant), and they can produce **exemplars** — a recorded measurement can capture the trace and span id active at that moment, letting a dashboard jump from a slow bucket straight into a matching trace. **Asynchronous** (observable) instruments — ObservableCounter, ObservableUpDownCounter, ObservableGauge — are registered with a callback. The SDK invokes it once per collection cycle, and the callback **observes the absolute current value**, not a delta. This is the point people get wrong: an ObservableCounter callback reports the cumulative total so far (say, total bytes read since start), and the SDK derives deltas if the exporter needs them. Reporting an increment from an async callback double-counts or under-counts depending on temporality. When to prefer async: - The value already exists to be read cheaply and continuously changes (process CPU time, resident memory, garbage-collection totals, file-descriptor count). - Instrumenting every mutation would be invasive or hot-path expensive. - The value lives outside your code — a library's internal counter, an OS statistic. When to prefer sync: - The measurement is tied to an event you already handle (a request completing). - You need per-event attributes or exemplars. - Missing changes between collections would be misleading. Asynchronous callbacks have rules of their own: they must be fast and must not block, because they run on the collection path; they should report each distinct attribute set at most once per collection, since duplicate observations for the same attributes in one cycle are undefined; and they must not have side effects, because how often they run depends on collection configuration, not on your code. ## Data points, the shipped form Whatever instrument you choose, what leaves the process is a **data point**: an attribute set, a time window (a start timestamp and an observation timestamp), and a value. Sums (from counters and up-down counters) carry `is_monotonic` and an aggregation temporality; gauges carry neither, since a current value has no accumulation window. Histograms carry a count, a sum, and either explicit bucket boundaries with per-bucket counts or the exponential-histogram form. Exemplars may attach to any of them. ## Attributes are where the cost lives Every distinct attribute combination is a separate time series that the SDK holds in memory and the backend stores forever. A counter with a `user_id` attribute is not a metric, it is a very expensive log. Keep metric attributes to bounded sets — route template, status class, region, tenant if tenants are countable — and push high-cardinality identifiers to spans and logs, which are built for them. ## A worked selection HTTP server instrumentation: a **histogram** of request duration attributed by route template and status class; the count and sum come free with the histogram, so a separate request counter is redundant. In-flight requests: an **UpDownCounter**, incremented on entry and decremented on exit. Process memory: an **ObservableGauge** read from the runtime once per collection. Total bytes written by a library that already tracks them: an **ObservableCounter** reporting that library's cumulative total.
- You need request throughput and p99 latency. How many instruments do you create?One histogram of request duration. A histogram data point already carries a count and a sum, so request rate comes from the count and the mean from sum/count, while the buckets give you the tail. Adding a separate request counter duplicates a series for no new information — the exception being if you need counts at attribute dimensions you deliberately excluded from the histogram to control cost.
- What goes wrong if an observable counter's callback reports the increment since the last collection?The SDK treats an observable counter's observation as the cumulative value and computes deltas itself, so reporting an increment makes the series look like a small number that keeps resetting rather than a rising total. Rates become nonsense, and under cumulative temporality the exported value will not be monotonic, which some backends interpret as repeated counter resets.
- Why is queue depth an up-down counter rather than a gauge?Either can work, and the choice hinges on how the value is produced. If your code sees every enqueue and dequeue, an UpDownCounter recorded synchronously is exact and never misses a change between collections. If the depth lives inside a third-party structure you can only poll, an ObservableGauge (or ObservableUpDownCounter) reading it per collection is the honest representation, and you accept that spikes between observations are invisible.
saying these in an interview costs you the question
- Using a counter for a value that can decrease
- Reporting a delta from an asynchronous callback instead of the absolute value
- Computing latency as a sum divided by a count and calling it performance monitoring
- Putting user ids, request ids or raw URLs into metric attributes
- Doing expensive or blocking work inside an observable instrument's callback