You must ship metrics to Prometheus (local scrape) and a hosted APM simultaneously, with per-backend cost/cardinality controls and clean shutdown. How do you design this with Micrometer registries?
answer
- two starters -> one composite, instrument once
- per-child MeterFilter for cost/cardinality
- Prometheus pull vs APM push + step
- graceful shutdown flushes push registries
- common tags at composite, denies at child
basics
~20 sAdd both registry starters so Boot fans out via one CompositeMeterRegistry. Instrument once against the injected MeterRegistry. Apply per-child MeterFilters for cost/cardinality, tune each backend's step, and ensure push registries close on shutdown to flush.
solid answer
~50 sRely on the registry-per-backend model: add `micrometer-registry-prometheus` and the APM's registry starter; Boot auto-configures both and adds them to the shared `CompositeMeterRegistry`, which becomes the injected `MeterRegistry`. Instrument **once** against that interface. Control cost independently by registering `MeterFilter`s on the *specific* child registry via type-targeted `MeterRegistryCustomizer` — e.g. deny high-cardinality or debug meters on the paid APM child while keeping them on Prometheus, plus `maximumAllowableTags` guards. Prometheus is pull-based (no step publishing) while the APM is push-based with a configurable `step` (publish interval) and batching — tune each separately. For clean shutdown, push registries implement lifecycle so Boot closes them and they flush a final batch; ensure the app context shuts down gracefully rather than being killed. Add common tags (application, region, instance) once via a composite-level customizer, and consider `maximumAllowableMetrics` as a global backstop.
code
java · 41 linesimport io.micrometer.core.instrument.MeterRegistry;
import io.micrometer.core.instrument.config.MeterFilter;
import io.micrometer.datadog.DatadogMeterRegistry;
import io.micrometer.prometheusmetrics.PrometheusMeterRegistry;
import org.springframework.boot.actuate.autoconfigure.metrics.MeterRegistryCustomizer;
import org.springframework.context.annotation.Bean;
import org.springframework.context.annotation.Configuration;
@Configuration
class MultiBackendMetricsConfig {
// 1) Identical dimensions on BOTH backends (composite-level)
@Bean
MeterRegistryCustomizer<MeterRegistry> commonTags() {
return r -> r.config().commonTags("application", "katajob", "region", "eu-west-1");
}
// 2) Cost control: keep noisy/high-cardinality series OFF the paid APM only
@Bean
MeterRegistryCustomizer<DatadogMeterRegistry> datadogCostGuards() {
return r -> r.config()
.meterFilter(MeterFilter.denyNameStartsWith("debug"))
.meterFilter(MeterFilter.maximumAllowableTags(
"http.server.requests", "uri", 50, MeterFilter.deny()));
}
// 3) Prometheus keeps rich data locally (histogram buckets for server-side quantiles)
@Bean
MeterRegistryCustomizer<PrometheusMeterRegistry> prometheusHistograms() {
return r -> r.config().meterFilter(new MeterFilter() {
@Override
public io.micrometer.core.instrument.distribution.DistributionStatisticConfig configure(
io.micrometer.core.instrument.Meter.Id id,
io.micrometer.core.instrument.distribution.DistributionStatisticConfig config) {
return io.micrometer.core.instrument.distribution.DistributionStatisticConfig.builder()
.percentilesHistogram(true).build().merge(config);
}
});
}
}
// application.yml: server.shutdown=graceful, management.datadog.metrics.export.step=60sgo deeper
Know two starters give you two backends automatically.
Explain the composite fans out and you instrument once.
Apply per-child filters, tune step, and handle common tags.
Architect the full solution: composite-vs-child filter placement, per-backend cost/cardinality/distribution tuning, pull-vs-push semantics, graceful shutdown flush, and failure isolation.
This is a design question testing whether you can exploit Micrometer's architecture rather than reinvent it. **Core architecture — fan-out for free:** Micrometer's *registry-per-backend* model plus `CompositeMeterRegistry` means multi-backend is a wiring concern, not a code concern. Add both dependencies: - `micrometer-registry-prometheus` → `PrometheusMeterRegistry` (pull-based: exposes `/actuator/prometheus`; Prometheus server scrapes it; there is no push 'step' for counters — values accumulate and are read on scrape). - the APM's starter (e.g. Datadog/New Relic/Dynatrace) → a push-based registry that ships batches on a **step** interval. Boot auto-configures both and registers them into the primary `CompositeMeterRegistry`. Your code injects `MeterRegistry` (the composite) and instruments **once**; every meter fans out to both. **Per-backend cost & cardinality control** — the crux. Because each child has its own `config()`, apply rules where they belong: - **Suppress expensive series on the paid backend only:** a `MeterRegistryCustomizer<DatadogMeterRegistry>` that adds `MeterFilter.denyNameStartsWith("debug")` or a predicate deny for high-cardinality names — Prometheus (cheap/local) still gets them. - **Cardinality guards:** `MeterFilter.maximumAllowableTags(name, tagKey, limit, MeterFilter.deny())` caps a tag's distinct values; `MeterFilter.maximumAllowableMetrics(n)` is a global backstop against runaway series. - **Sampling/aggregation differences:** distribution config (percentiles vs histogram buckets) can differ per backend — client-side percentiles for the APM, histogram buckets for Prometheus so Grafana computes quantiles server-side. **Tuning publishing:** push registries expose a `step` (e.g. `management.<system>.metrics.export.step=60s`) governing publish cadence and the window for rate/count aggregation; batch size and connect/read timeouts matter for reliability. Prometheus has no step in the same sense — scrape interval is set on the Prometheus server side, so align your reset semantics accordingly. **Common tags once:** attach `application`, `region`, `instance`/`host` via a composite-level `MeterRegistryCustomizer<MeterRegistry>` `commonTags(...)` so both backends carry identical dimensions for cross-correlation. **Clean shutdown / no lost tail:** push registries implement `Closeable`/Spring lifecycle. On graceful context shutdown Boot calls `close()`, which triggers a **final publish/flush** so the last window isn't lost. Design implications: - Enable graceful shutdown (`server.shutdown=graceful`) and give the platform (Kubernetes `terminationGracePeriodSeconds`) enough time for the flush. - A `kill -9` skips flush — the final step's data is lost; accept or mitigate with a shorter step. - If you manually manage registries (rare in Boot), you're responsible for calling `close()`. **Failure isolation:** one backend being down (e.g. APM endpoint unreachable) shouldn't break the other — each child publishes independently; the composite doesn't couple their success. Watch that a slow/failing push registry doesn't back up threads (registries publish on their own scheduler). **Testing & rollout:** validate with `SimpleMeterRegistry` in tests; stage the APM behind an enable flag (`management.<system>.metrics.export.enabled`) so you can dark-launch. Document where each filter lives so the cost/observability tradeoff is auditable. **Summary of the design:** one instrumentation surface (the composite), two independently-tuned children, filters/tags placed at composite-vs-child deliberately, step tuned per backend, graceful shutdown to flush push registries.
- Why does graceful shutdown matter for a push-based registry but less so for Prometheus?Push registries buffer a window and flush on close(); a graceful shutdown lets Boot call close() to publish the final batch, avoiding lost tail data. Prometheus is pull-based — the server scrapes the endpoint — so there's no client-side buffer to flush; you only lose data not yet scraped.
- If the APM endpoint is unreachable, does it affect Prometheus export?No. Each child registry publishes independently on its own scheduler; the composite doesn't couple their success. A failing push registry logs/retries per its config without blocking the pull-based Prometheus endpoint, though you should ensure its failures don't exhaust threads.
- Where would you place a filter that must apply to both backends versus one backend?Both: register it on the composite (MeterRegistryCustomizer<MeterRegistry>) so it governs the fan-out. One backend: type-target the child (e.g. MeterRegistryCustomizer<DatadogMeterRegistry>) so only that registry's copy is affected.
saying these in an interview costs you the question
- Instrumenting the same metric twice, once per backend, instead of relying on the composite.
- Putting cost/cardinality filters only on the composite when they should differ per backend.
- Ignoring graceful shutdown, silently losing the last publish window of push registries.
- Assuming Prometheus has a client-side publish 'step' like push registries.
- Believing one backend's outage breaks the other's export.