skip to content

For a fleet of Spring Boot services — some long-lived, some serverless/ephemeral — how would you decide between Prometheus scraping and OTLP push, and what failure modes shape that choice?

level: principalimportance: nice to knowfreq 25%

answer

  1. axis = reachability + lifecycle
  2. pull: free up-signal, monitor owns cadence, discoverable/long-lived
  3. push: ephemeral/serverless/NAT, one OTel pipeline
  4. OTel Collector tier = choke point (auth, relabel, fan-out)
  5. cardinality is the real scaling limiter; run both via composite during migration

basics

~10 s

Use Prometheus scraping for stable, discoverable services (free liveness via failed scrapes, monitor controls timing). Use OTLP push for ephemeral/serverless or unreachable apps. Often run both through an OpenTelemetry Collector, migrating incrementally.

solid answer

~50 s

The deciding axis is reachability and lifecycle. Long-lived, discoverable services fit Prometheus **pull**: no app-side target config, the monitor owns cadence, and a failed scrape yields a free `up`/liveness signal. Ephemeral or serverless workloads may die between scrapes and often sit behind NAT/firewalls, so **OTLP push** fits — they emit before exiting, needing no inbound reachability. At fleet scale I'd standardize on an **OpenTelemetry Collector** tier as the ingestion choke point: services push OTLP (or the collector scrapes Prometheus endpoints), and the collector centralizes auth, batching, relabeling, and fan-out to storage — and can re-export to Prometheus for existing dashboards. Micrometer lets a service run both registries during migration. Key failure modes: push loses a step's data if the collector is down (mitigate with sidecar collectors), pull loses undiscovered/unreachable targets, and both are dominated by **cardinality** blow-ups from unbounded tags. Naming (snake_case vs dotted) must be reconciled at the collector, not per service.

go deeper

for a junior

Not expected; grasp only that stable services get scraped, ephemeral ones push.

for a middle

Name the reachability/lifecycle trade-off and that both can run together.

for a senior

Discuss collector tiering, free up-signal, and push data-loss mitigation.

for a principal

Own the fleet-wide topology, cardinality governance, migration via composite registry, and where naming/temporality reconcile — bridging to (not re-teaching) the downstream storage/query stack.

**Frame the decision by reachability + lifecycle, then by operational ownership.** **Pull (Prometheus scrape) — best when:** - Services are long-lived and **discoverable** (Kubernetes/Consul/DNS service discovery feeds Prometheus targets). - You value the monitor owning cadence and getting **failed-scrape liveness** for free: Prometheus synthesizes an `up` series per target, so 'app unreachable' is detectable without extra plumbing. - Simplicity in the app: no outbound target, no auth token, no buffering — the app just serves `/actuator/prometheus`. - **Weaknesses:** short-lived jobs may never be scraped (their whole life fits between scrapes); targets the monitor can't reach or discover simply disappear; scaling to huge target counts pushes you toward federation/sharding or agent-based scraping. **Push (OTLP) — best when:** - Workloads are **ephemeral/serverless** (FaaS, batch, autoscaled-to-zero) that must emit before exiting. - Apps are **unreachable inbound** (NAT, firewalls, egress-only networks). - You want a single **OpenTelemetry pipeline** for metrics, traces, and logs, with routing/auth/multitenancy centralized. - **Weaknesses:** the app owns the target, auth, and cadence; a collector outage during a step can drop that window (limited in-app buffering, no durable queue); no free `up` signal — gaps must be interpreted; too-small a step floods the collector and inflates cost. **The pragmatic fleet architecture — an OpenTelemetry Collector tier:** - Deploy collectors as **sidecars** (per-pod resilience, local buffering) and/or a **gateway** tier (central relabeling, tail-based routing, backend fan-out). - Services push OTLP to a nearby collector (resilient, low-latency), or the collector's Prometheus receiver **scrapes** service `/actuator/prometheus` endpoints — you can mix per service. - The collector normalizes **naming** (dotted OTel vs Prometheus snake_case), enforces **cardinality** limits via processors, adds resource attributes, and re-exports to whatever storage each team needs (including a Prometheus-compatible store to preserve existing PromQL dashboards and alerts). This is where the pull/push naming and temporality differences get reconciled — not in each service. **Migration path:** Micrometer's `CompositeMeterRegistry` lets a service run **both** `micrometer-registry-prometheus` and `micrometer-registry-otlp` at once. So you can introduce OTLP push while existing Prometheus scraping continues, validate parity, then retire one side — no big-bang cutover. **Cross-cutting failure modes (dominant at scale):** 1. **Cardinality explosions** — high-cardinality tags (user ids, raw URIs, request ids) multiply time series/attributes and can OOM the store; govern with `MeterFilter`s and collector processors. This dwarfs the pull-vs-push distinction. 2. **Push data loss** on collector outage within a step → sidecar collectors + reasonable step sizing. 3. **Pull blind spots** for undiscovered/short-lived targets → OTLP push or Pushgateway for batch jobs (edge case). 4. **Temporality/naming mismatches** across pipelines → standardize at the collector. 5. **Cost/latency of step vs scrape interval** — tune per SLO; don't assume push is 'realtime'. **Bridge, don't re-teach:** the choice of storage/query stack (Prometheus server, PromQL, Grafana, or a vendor OTLP backend) is a system-design concern downstream of these registries; the app-side decision is purely 'how do metrics leave the JVM and reach that tier'. **Summary heuristic:** stable+discoverable → pull; ephemeral/unreachable or one-OTel-pipeline goal → push; large heterogeneous fleet → collector tier ingesting both, with cardinality governance as the first-order scaling concern.

  • How do you monitor 'app is down' when using OTLP push, given there's no free up signal?
    You add explicit liveness detection: alert on absence of expected metrics/heartbeat within N steps, or complement with health-check probes (Kubernetes liveness, a synthetic/blackbox check). Push gives a data gap, not a signal, so you must define the gap-to-alert rule yourself.
  • Why is cardinality a bigger concern than the pull-vs-push choice?
    Both backends cost scales with the number of unique series/attribute-sets. Unbounded tags (user ids, raw URIs) create millions of series and can OOM storage regardless of transport, so tag governance via MeterFilter/collector processors is the first-order scaling lever.

saying these in an interview costs you the question

  • Treating push as strictly superior/real-time rather than a reachability/lifecycle trade-off.
  • Ignoring cardinality as the dominant scaling risk.
  • Assuming you must pick exactly one transport instead of composing both during migration.
  • Trying to reconcile naming/temporality per service instead of at the collector.

context