For a fleet of Spring Boot services — some long-lived, some serverless/ephemeral — how would you decide between Prometheus scraping and OTLP push, and what failure modes shape that choice?
answer
- axis = reachability + lifecycle
- pull: free up-signal, monitor owns cadence, discoverable/long-lived
- push: ephemeral/serverless/NAT, one OTel pipeline
- OTel Collector tier = choke point (auth, relabel, fan-out)
- cardinality is the real scaling limiter; run both via composite during migration
basics
~10 sUse Prometheus scraping for stable, discoverable services (free liveness via failed scrapes, monitor controls timing). Use OTLP push for ephemeral/serverless or unreachable apps. Often run both through an OpenTelemetry Collector, migrating incrementally.
solid answer
~50 sThe deciding axis is reachability and lifecycle. Long-lived, discoverable services fit Prometheus **pull**: no app-side target config, the monitor owns cadence, and a failed scrape yields a free `up`/liveness signal. Ephemeral or serverless workloads may die between scrapes and often sit behind NAT/firewalls, so **OTLP push** fits — they emit before exiting, needing no inbound reachability. At fleet scale I'd standardize on an **OpenTelemetry Collector** tier as the ingestion choke point: services push OTLP (or the collector scrapes Prometheus endpoints), and the collector centralizes auth, batching, relabeling, and fan-out to storage — and can re-export to Prometheus for existing dashboards. Micrometer lets a service run both registries during migration. Key failure modes: push loses a step's data if the collector is down (mitigate with sidecar collectors), pull loses undiscovered/unreachable targets, and both are dominated by **cardinality** blow-ups from unbounded tags. Naming (snake_case vs dotted) must be reconciled at the collector, not per service.
go deeper
Not expected; grasp only that stable services get scraped, ephemeral ones push.
Name the reachability/lifecycle trade-off and that both can run together.
Discuss collector tiering, free up-signal, and push data-loss mitigation.
Own the fleet-wide topology, cardinality governance, migration via composite registry, and where naming/temporality reconcile — bridging to (not re-teaching) the downstream storage/query stack.
**Frame the decision by reachability + lifecycle, then by operational ownership.** **Pull (Prometheus scrape) — best when:** - Services are long-lived and **discoverable** (Kubernetes/Consul/DNS service discovery feeds Prometheus targets). - You value the monitor owning cadence and getting **failed-scrape liveness** for free: Prometheus synthesizes an `up` series per target, so 'app unreachable' is detectable without extra plumbing. - Simplicity in the app: no outbound target, no auth token, no buffering — the app just serves `/actuator/prometheus`. - **Weaknesses:** short-lived jobs may never be scraped (their whole life fits between scrapes); targets the monitor can't reach or discover simply disappear; scaling to huge target counts pushes you toward federation/sharding or agent-based scraping. **Push (OTLP) — best when:** - Workloads are **ephemeral/serverless** (FaaS, batch, autoscaled-to-zero) that must emit before exiting. - Apps are **unreachable inbound** (NAT, firewalls, egress-only networks). - You want a single **OpenTelemetry pipeline** for metrics, traces, and logs, with routing/auth/multitenancy centralized. - **Weaknesses:** the app owns the target, auth, and cadence; a collector outage during a step can drop that window (limited in-app buffering, no durable queue); no free `up` signal — gaps must be interpreted; too-small a step floods the collector and inflates cost. **The pragmatic fleet architecture — an OpenTelemetry Collector tier:** - Deploy collectors as **sidecars** (per-pod resilience, local buffering) and/or a **gateway** tier (central relabeling, tail-based routing, backend fan-out). - Services push OTLP to a nearby collector (resilient, low-latency), or the collector's Prometheus receiver **scrapes** service `/actuator/prometheus` endpoints — you can mix per service. - The collector normalizes **naming** (dotted OTel vs Prometheus snake_case), enforces **cardinality** limits via processors, adds resource attributes, and re-exports to whatever storage each team needs (including a Prometheus-compatible store to preserve existing PromQL dashboards and alerts). This is where the pull/push naming and temporality differences get reconciled — not in each service. **Migration path:** Micrometer's `CompositeMeterRegistry` lets a service run **both** `micrometer-registry-prometheus` and `micrometer-registry-otlp` at once. So you can introduce OTLP push while existing Prometheus scraping continues, validate parity, then retire one side — no big-bang cutover. **Cross-cutting failure modes (dominant at scale):** 1. **Cardinality explosions** — high-cardinality tags (user ids, raw URIs, request ids) multiply time series/attributes and can OOM the store; govern with `MeterFilter`s and collector processors. This dwarfs the pull-vs-push distinction. 2. **Push data loss** on collector outage within a step → sidecar collectors + reasonable step sizing. 3. **Pull blind spots** for undiscovered/short-lived targets → OTLP push or Pushgateway for batch jobs (edge case). 4. **Temporality/naming mismatches** across pipelines → standardize at the collector. 5. **Cost/latency of step vs scrape interval** — tune per SLO; don't assume push is 'realtime'. **Bridge, don't re-teach:** the choice of storage/query stack (Prometheus server, PromQL, Grafana, or a vendor OTLP backend) is a system-design concern downstream of these registries; the app-side decision is purely 'how do metrics leave the JVM and reach that tier'. **Summary heuristic:** stable+discoverable → pull; ephemeral/unreachable or one-OTel-pipeline goal → push; large heterogeneous fleet → collector tier ingesting both, with cardinality governance as the first-order scaling concern.
- How do you monitor 'app is down' when using OTLP push, given there's no free up signal?You add explicit liveness detection: alert on absence of expected metrics/heartbeat within N steps, or complement with health-check probes (Kubernetes liveness, a synthetic/blackbox check). Push gives a data gap, not a signal, so you must define the gap-to-alert rule yourself.
- Why is cardinality a bigger concern than the pull-vs-push choice?Both backends cost scales with the number of unique series/attribute-sets. Unbounded tags (user ids, raw URIs) create millions of series and can OOM storage regardless of transport, so tag governance via MeterFilter/collector processors is the first-order scaling lever.
saying these in an interview costs you the question
- Treating push as strictly superior/real-time rather than a reachability/lifecycle trade-off.
- Ignoring cardinality as the dominant scaling risk.
- Assuming you must pick exactly one transport instead of composing both during migration.
- Trying to reconcile naming/temporality per service instead of at the collector.