skip to content

You must ship metrics to Prometheus (local scrape) and a hosted APM simultaneously, with per-backend cost/cardinality controls and clean shutdown. How do you design this with Micrometer registries?

level: principalimportance: should knowfreq 25%

answer

  1. two starters -> one composite, instrument once
  2. per-child MeterFilter for cost/cardinality
  3. Prometheus pull vs APM push + step
  4. graceful shutdown flushes push registries
  5. common tags at composite, denies at child

basics

~20 s

Add both registry starters so Boot fans out via one CompositeMeterRegistry. Instrument once against the injected MeterRegistry. Apply per-child MeterFilters for cost/cardinality, tune each backend's step, and ensure push registries close on shutdown to flush.

solid answer

~50 s

Rely on the registry-per-backend model: add `micrometer-registry-prometheus` and the APM's registry starter; Boot auto-configures both and adds them to the shared `CompositeMeterRegistry`, which becomes the injected `MeterRegistry`. Instrument **once** against that interface. Control cost independently by registering `MeterFilter`s on the *specific* child registry via type-targeted `MeterRegistryCustomizer` — e.g. deny high-cardinality or debug meters on the paid APM child while keeping them on Prometheus, plus `maximumAllowableTags` guards. Prometheus is pull-based (no step publishing) while the APM is push-based with a configurable `step` (publish interval) and batching — tune each separately. For clean shutdown, push registries implement lifecycle so Boot closes them and they flush a final batch; ensure the app context shuts down gracefully rather than being killed. Add common tags (application, region, instance) once via a composite-level customizer, and consider `maximumAllowableMetrics` as a global backstop.

code

java · 41 lines
java
import io.micrometer.core.instrument.MeterRegistry;
import io.micrometer.core.instrument.config.MeterFilter;
import io.micrometer.datadog.DatadogMeterRegistry;
import io.micrometer.prometheusmetrics.PrometheusMeterRegistry;
import org.springframework.boot.actuate.autoconfigure.metrics.MeterRegistryCustomizer;
import org.springframework.context.annotation.Bean;
import org.springframework.context.annotation.Configuration;

@Configuration
class MultiBackendMetricsConfig {

    // 1) Identical dimensions on BOTH backends (composite-level)
    @Bean
    MeterRegistryCustomizer<MeterRegistry> commonTags() {
        return r -> r.config().commonTags("application", "katajob", "region", "eu-west-1");
    }

    // 2) Cost control: keep noisy/high-cardinality series OFF the paid APM only
    @Bean
    MeterRegistryCustomizer<DatadogMeterRegistry> datadogCostGuards() {
        return r -> r.config()
                .meterFilter(MeterFilter.denyNameStartsWith("debug"))
                .meterFilter(MeterFilter.maximumAllowableTags(
                        "http.server.requests", "uri", 50, MeterFilter.deny()));
    }

    // 3) Prometheus keeps rich data locally (histogram buckets for server-side quantiles)
    @Bean
    MeterRegistryCustomizer<PrometheusMeterRegistry> prometheusHistograms() {
        return r -> r.config().meterFilter(new MeterFilter() {
            @Override
            public io.micrometer.core.instrument.distribution.DistributionStatisticConfig configure(
                    io.micrometer.core.instrument.Meter.Id id,
                    io.micrometer.core.instrument.distribution.DistributionStatisticConfig config) {
                return io.micrometer.core.instrument.distribution.DistributionStatisticConfig.builder()
                        .percentilesHistogram(true).build().merge(config);
            }
        });
    }
}
// application.yml: server.shutdown=graceful, management.datadog.metrics.export.step=60s

go deeper

for a junior

Know two starters give you two backends automatically.

for a middle

Explain the composite fans out and you instrument once.

for a senior

Apply per-child filters, tune step, and handle common tags.

for a principal

Architect the full solution: composite-vs-child filter placement, per-backend cost/cardinality/distribution tuning, pull-vs-push semantics, graceful shutdown flush, and failure isolation.

This is a design question testing whether you can exploit Micrometer's architecture rather than reinvent it. **Core architecture — fan-out for free:** Micrometer's *registry-per-backend* model plus `CompositeMeterRegistry` means multi-backend is a wiring concern, not a code concern. Add both dependencies: - `micrometer-registry-prometheus` → `PrometheusMeterRegistry` (pull-based: exposes `/actuator/prometheus`; Prometheus server scrapes it; there is no push 'step' for counters — values accumulate and are read on scrape). - the APM's starter (e.g. Datadog/New Relic/Dynatrace) → a push-based registry that ships batches on a **step** interval. Boot auto-configures both and registers them into the primary `CompositeMeterRegistry`. Your code injects `MeterRegistry` (the composite) and instruments **once**; every meter fans out to both. **Per-backend cost & cardinality control** — the crux. Because each child has its own `config()`, apply rules where they belong: - **Suppress expensive series on the paid backend only:** a `MeterRegistryCustomizer<DatadogMeterRegistry>` that adds `MeterFilter.denyNameStartsWith("debug")` or a predicate deny for high-cardinality names — Prometheus (cheap/local) still gets them. - **Cardinality guards:** `MeterFilter.maximumAllowableTags(name, tagKey, limit, MeterFilter.deny())` caps a tag's distinct values; `MeterFilter.maximumAllowableMetrics(n)` is a global backstop against runaway series. - **Sampling/aggregation differences:** distribution config (percentiles vs histogram buckets) can differ per backend — client-side percentiles for the APM, histogram buckets for Prometheus so Grafana computes quantiles server-side. **Tuning publishing:** push registries expose a `step` (e.g. `management.<system>.metrics.export.step=60s`) governing publish cadence and the window for rate/count aggregation; batch size and connect/read timeouts matter for reliability. Prometheus has no step in the same sense — scrape interval is set on the Prometheus server side, so align your reset semantics accordingly. **Common tags once:** attach `application`, `region`, `instance`/`host` via a composite-level `MeterRegistryCustomizer<MeterRegistry>` `commonTags(...)` so both backends carry identical dimensions for cross-correlation. **Clean shutdown / no lost tail:** push registries implement `Closeable`/Spring lifecycle. On graceful context shutdown Boot calls `close()`, which triggers a **final publish/flush** so the last window isn't lost. Design implications: - Enable graceful shutdown (`server.shutdown=graceful`) and give the platform (Kubernetes `terminationGracePeriodSeconds`) enough time for the flush. - A `kill -9` skips flush — the final step's data is lost; accept or mitigate with a shorter step. - If you manually manage registries (rare in Boot), you're responsible for calling `close()`. **Failure isolation:** one backend being down (e.g. APM endpoint unreachable) shouldn't break the other — each child publishes independently; the composite doesn't couple their success. Watch that a slow/failing push registry doesn't back up threads (registries publish on their own scheduler). **Testing & rollout:** validate with `SimpleMeterRegistry` in tests; stage the APM behind an enable flag (`management.<system>.metrics.export.enabled`) so you can dark-launch. Document where each filter lives so the cost/observability tradeoff is auditable. **Summary of the design:** one instrumentation surface (the composite), two independently-tuned children, filters/tags placed at composite-vs-child deliberately, step tuned per backend, graceful shutdown to flush push registries.

  • Why does graceful shutdown matter for a push-based registry but less so for Prometheus?
    Push registries buffer a window and flush on close(); a graceful shutdown lets Boot call close() to publish the final batch, avoiding lost tail data. Prometheus is pull-based — the server scrapes the endpoint — so there's no client-side buffer to flush; you only lose data not yet scraped.
  • If the APM endpoint is unreachable, does it affect Prometheus export?
    No. Each child registry publishes independently on its own scheduler; the composite doesn't couple their success. A failing push registry logs/retries per its config without blocking the pull-based Prometheus endpoint, though you should ensure its failures don't exhaust threads.
  • Where would you place a filter that must apply to both backends versus one backend?
    Both: register it on the composite (MeterRegistryCustomizer<MeterRegistry>) so it governs the fan-out. One backend: type-target the child (e.g. MeterRegistryCustomizer<DatadogMeterRegistry>) so only that registry's copy is affected.

saying these in an interview costs you the question

  • Instrumenting the same metric twice, once per backend, instead of relying on the composite.
  • Putting cost/cardinality filters only on the composite when they should differ per backend.
  • Ignoring graceful shutdown, silently losing the last publish window of push registries.
  • Assuming Prometheus has a client-side publish 'step' like push registries.
  • Believing one backend's outage breaks the other's export.

context