You need OpenTelemetry metrics to land in a Prometheus-compatible metrics store. Walk through the push-versus-pull options, and the two data-model mismatches — aggregation temporality and metric naming — that you have to resolve on the way.
answer
- pull /metrics · remote-write · OTLP receive endpoint
- Prometheus = cumulative; rate() reads a drop as a reset
- delta→cumulative processor is stateful, per-series memory
- dots→underscores, unit suffix, _total on monotonic sums
- resource attrs → target_info gauge, join on job/instance
basics
~20 sEither expose a scrape endpoint (a Prometheus exporter that Prometheus pulls) or push — OTLP into Prometheus's OTLP receive endpoint, or remote-write from a collector. Prometheus needs cumulative counters, so delta temporality must be converted; and OTel metric names must be normalised (dots to underscores, unit suffixes, _total on monotonic sums), while resource attributes land on a separate target_info series.
solid answer
~60 s**Transport options.** (1) *Pull*: a Prometheus exporter exposes `/metrics` on the process or on a collector, and Prometheus scrapes it — natural fit, but the scraper must be able to reach every instance. (2) *Push via remote write*: a collector converts OTLP to Prometheus remote-write. (3) *Push via OTLP*: recent Prometheus versions accept OTLP metrics directly on a dedicated receive path. **Temporality.** OTel supports delta and cumulative; Prometheus's model is cumulative counters that reset to zero on restart, and `rate()` detects those resets. Delta series cannot be stored meaningfully as-is, so either configure the SDK's temporality preference to cumulative or run a delta-to-cumulative conversion in the collector — which is stateful, memory-hungry and loses accuracy across restarts. Cumulative-at-source is preferable for Prometheus; delta is preferred by some commercial backends, which is why the preference is configurable. **Naming.** Dots become underscores, illegal characters are replaced, the unit is appended as a suffix, monotonic sums get `_total`. Resource attributes are not copied onto every series — they become a `target_info` metric you join against. Exemplars carry trace ids through if the exposition format supports them.
go deeper
Know that Prometheus pulls by default, that OTel can expose a scrape endpoint, and that names get normalised with underscores and suffixes.
Explain cumulative versus delta and why Prometheus needs cumulative, plus the naming rules including _total and unit suffixes.
Compare all three transports operationally, cost out delta-to-cumulative conversion, explain target_info and the join, and cover histograms, exemplars and staleness.
Own the decision across destinations: temporality per backend, where translation lives, cardinality budgets enforced at the pipeline, and the migration cost when metric names change.
## Two metric models that nearly agree OpenTelemetry's metric model and Prometheus's are close enough to translate and different enough to hurt. OTel has instruments (counter, up-down counter, histogram, gauge, plus asynchronous variants), dimensional attributes, a resource describing the producing entity, and a configurable *aggregation temporality*. Prometheus has metric families with label sets, a pull-based exposition format, monotonic counters that only ever go up (until they reset to zero), and no separate notion of resource — everything is a label. The translation is specified (there is an official OTel↔Prometheus compatibility specification), but the specification's job is to tell you what is lost. ## Getting the data across: three routes **Pull — a Prometheus exporter.** The SDK (or a collector) exposes an HTTP endpoint in Prometheus exposition format and Prometheus scrapes it on its own schedule. Advantages: staleness and up/down detection are native, the scrape interval is the store's decision, and there is no queueing in the application. Disadvantages: the scraper must have network reachability and discovery for every instance, which is awkward for short-lived, serverless or NAT'd workloads. Note the model shift — a pull endpoint is stateful in the process: the exporter holds the current cumulative values and answers whenever asked. **Push — remote write from a collector.** Applications push OTLP to a collector, which converts to Prometheus remote-write and ships to the store. This suits short-lived workloads and centralised egress, at the cost of a stateful conversion tier you now operate. **Push — OTLP straight into Prometheus.** Recent Prometheus versions expose an OTLP receive path so OTLP metrics can be written without remote-write translation. It has to be explicitly enabled, and the translation rules (especially name normalisation) are configurable. Check the version before promising it. ## Mismatch 1: aggregation temporality *Cumulative* means each data point reports the total since the start of the process (or the start of the metric stream). *Delta* means each point reports only what happened since the previous point. Prometheus is cumulative by design: `rate()` and `increase()` compute over a monotonically increasing series and treat a drop as a counter reset — that is how a restart is handled without losing correctness. Feed it deltas and every point looks like a reset. So either: - **Set the SDK preference to cumulative** (the exporter-level temporality preference is configurable, e.g. via the OTLP metrics temporality preference setting). Cumulative is the default for the Prometheus path and costs the SDK a little memory to hold running totals. - **Convert in the collector** with a delta-to-cumulative processor. This is genuinely expensive: the processor must keep the running total *per series* in memory, so its footprint scales with active cardinality, and it must guess at stream restarts. State is lost when the collector restarts, producing artificial resets across the fleet. The reason delta exists at all is that some commercial backends prefer it — deltas are stateless to produce, survive instance churn better, and let the backend do the aggregation. If you export to both a Prometheus store and such a backend, you either pick cumulative and let the other side convert, or run two pipelines. Say that out loud: temporality is a *destination-driven* choice, not a universal best practice. ## Mismatch 2: naming and identity OTel metric names use dots (`http.server.request.duration`) and carry a unit as separate metadata. Prometheus historically required a restricted character set and encodes the unit in the name. The compatibility rules: - non-alphanumeric characters, dots included, become `_`; - the unit is appended as a suffix (`_seconds`, `_bytes`), using the base unit; - a monotonic sum gets `_total`; - a name may not start with a digit, so a prefix is added. Newer Prometheus versions support UTF-8 metric names, which makes normalisation optional and configurable — helpful for round-tripping, but it changes the names your dashboards and alert rules reference, so it is a migration, not a toggle. Attributes become labels, with the same character normalisation. The subtler point is the **resource**: OTel's resource attributes (service name, instance id, k8s pod, cloud region…) describe the producing entity, and copying all of them onto every series would multiply cardinality. Instead, a small identifying subset maps to `job` and `instance` (derived from service namespace/name and instance id) and the remainder is published as a separate `target_info` gauge carrying those attributes as labels, with a value of 1. To use them in a query you *join* against `target_info` on job and instance. Candidates who have never done that join usually have never actually shipped OTel metrics to Prometheus. ## Histograms and exemplars OTel explicit-bucket histograms map onto classic Prometheus histograms (`_bucket`/`_sum`/`_count`). OTel *exponential* histograms map onto Prometheus native histograms where supported — otherwise they must be converted to fixed buckets, losing the adaptive resolution that made them attractive. Exemplars — sampled data points carrying a trace id — survive if the exposition path supports them (OpenMetrics exposition, native histograms, or remote-write with exemplar support enabled), and they are what makes 'click a latency spike, land on a trace' work end to end. ## Operational cautions - **Staleness.** Prometheus infers staleness from scrapes. Pushed series do not get that for free, so a disappeared instance can leave a series that looks alive. - **Timestamp handling.** Pushed points carry their own timestamps; out-of-order or too-old samples may be rejected depending on the store's configuration. - **Cardinality.** OTel attributes are cheap to add at the SDK and expensive to store. The translation makes each attribute a label, so an unbounded attribute becomes an unbounded label set — the single most common way a migration blows up storage cost. ## Interview framing Give the three transports in one breath, then spend the answer on temporality (why Prometheus needs cumulative, what delta-to-cumulative conversion costs) and naming/identity (suffix rules, and `target_info` as the resource carrier requiring a join). Close with exemplars and cardinality.
- Why is running a delta-to-cumulative conversion in a collector risky?It is stateful: the processor must hold a running total for every active series, so memory scales with cardinality and a cardinality spike can OOM the collector. It also has to infer stream restarts, and when the collector itself restarts the accumulated state is gone, producing artificial counter resets across everything flowing through it. Choosing cumulative at the SDK avoids all of that for a Prometheus destination.
- Where do OpenTelemetry resource attributes end up in Prometheus, and how do you query them?A small identifying subset becomes the `job` and `instance` labels; the rest are published as a `target_info` series with value 1 carrying those attributes as labels. To filter or group by, say, a cloud region you join your metric against `target_info` on job and instance. They are deliberately not copied onto every series because that would multiply cardinality.
- What breaks if the same metrics must also go to a backend that prefers delta temporality?You cannot satisfy both from one cumulative stream without conversion somewhere. Either export cumulative and let that backend convert on ingest, or run separate pipelines with different temporality preferences. It is a destination-driven choice, and the honest answer names the extra cost of whichever conversion tier you take on.
saying these in an interview costs you the question
- Assuming Prometheus can store delta counters directly.
- Expecting all resource attributes to appear as labels on every series instead of via `target_info`.
- Treating delta-to-cumulative conversion as free rather than stateful and cardinality-bound.
- Adding high-cardinality attributes at the SDK because they are cheap there, ignoring that each becomes a label.
- Believing exponential histograms always survive the translation with full resolution.