skip to content

Explain the difference between publishPercentiles, publishPercentileHistogram, and serviceLevelObjectives on a Timer/DistributionSummary, and why it matters for aggregation across instances.

level: principalimportance: should knowfreq 40%

answer

  1. publishPercentiles = client-side, NOT aggregatable
  2. publishPercentileHistogram = buckets, backend aggregates (histogram_quantile)
  3. serviceLevelObjectives = explicit threshold buckets (fraction under X)
  4. avg of p99 across pods is meaningless
  5. min/maxExpectedValue bounds accuracy+memory

basics

~10 s

publishPercentiles computes percentiles inside each instance (not mergeable across instances). publishPercentileHistogram exports histogram buckets the backend aggregates to compute percentiles. serviceLevelObjectives adds explicit boundary buckets so you can measure the fraction meeting a target.

solid answer

~40 s

There are two fundamentally different ways to get percentiles. publishPercentiles(0.95, 0.99) makes each application instance estimate its own p95/p99 client-side and export them as gauges — cheap, but you CANNOT average p99s across instances, so in a multi-replica service these numbers are misleading. publishPercentileHistogram() instead exports a set of cumulative histogram buckets (a distribution, e.g. Prometheus le buckets); the monitoring backend sums buckets across all instances and computes an accurate aggregate percentile with histogram_quantile — the correct approach for horizontally scaled services, at the cost of more series. serviceLevelObjectives(Duration...) adds explicit bucket boundaries at your SLO thresholds (e.g. 100ms, 500ms) so you can directly query 'what fraction of requests were under 500ms.' You bound accuracy/memory with minimumExpectedValue/maximumExpectedValue. Rule of thumb: histograms/SLOs for aggregatable server-side percentiles; raw publishPercentiles only for single-instance or quick local insight.

code

java · 21 lines
java
import io.micrometer.core.instrument.MeterRegistry;
import io.micrometer.core.instrument.Timer;
import java.time.Duration;

class TimerConfig {
    Timer requestTimer(MeterRegistry registry) {
        return Timer.builder("orders.request")
            // client-side estimate: cheap, single-instance only
            .publishPercentiles(0.95, 0.99)
            // aggregatable buckets: backend computes cross-instance p99
            .publishPercentileHistogram()
            // explicit SLO boundaries: fraction under 100ms / 500ms
            .serviceLevelObjectives(
                Duration.ofMillis(100),
                Duration.ofMillis(500))
            // bound histogram range -> accuracy + memory control
            .minimumExpectedValue(Duration.ofMillis(1))
            .maximumExpectedValue(Duration.ofSeconds(10))
            .register(registry);
    }
}

go deeper

for a junior

Aware percentiles must be enabled; may not distinguish the three mechanisms.

for a middle

Can name publishPercentiles vs publishPercentileHistogram and that one is client-side.

for a senior

Explains aggregatability, SLO buckets, and configures via Spring Boot properties with sensible bounds.

for a principal

Reasons about cross-instance correctness, TSDB cardinality/memory tradeoffs, backend capabilities, and picks the right mechanism per SLI.

## Why percentiles are hard A plain Timer gives you count/total/max → you can compute a **mean**, but means hide tail latency. To see p95/p99 you must configure a **distribution**. Micrometer offers three related-but-distinct knobs, and confusing them causes real production incidents (wrong SLO dashboards). ### 1. `publishPercentiles(0.5, 0.95, 0.99)` — client-side percentiles - Each **application instance** maintains an internal estimator (a ring buffer / HdrHistogram-style structure) and computes its **own** percentile values. - These are exported as **gauge** time series (e.g. `..._seconds{quantile="0.99"}`). - **Cheap to query, but NOT aggregatable.** You cannot average or sum percentiles: `avg(p99 across 10 pods)` is mathematically meaningless. In a horizontally-scaled service these numbers mislead. - Accuracy is bounded by `percentilePrecision` and the expected-value range. - Good for: a single instance, local dev, or when you truly only run one replica. ### 2. `publishPercentileHistogram()` — histogram buckets (server-side percentiles) - Instead of computing percentiles locally, each instance exports a set of **cumulative histogram buckets** — counts of observations `<= bucket boundary`. In Prometheus these are the `le="..."` buckets; the backend picks a set of boundaries covering a wide range. - The **monitoring system aggregates** buckets across all instances (sum by le) and computes the percentile with a function like Prometheus `histogram_quantile(0.99, ...)`. - **Aggregatable and correct across replicas** — this is the right choice for distributed/scaled services and for computing service-wide SLIs. - Cost: many more time series (one per bucket per tag combination) → higher storage/cardinality. - Only backends that support histogram-based percentile queries benefit (Prometheus, etc.). ### 3. `serviceLevelObjectives(Duration.ofMillis(100), Duration.ofMillis(500))` — explicit SLO boundaries - Adds **specific bucket boundaries at your SLO thresholds** on top of (or instead of) the auto histogram. - Lets you answer directly: *'what fraction of requests completed under 500ms?'* by reading the cumulative count at that bucket vs total. - Also aggregatable server-side. Ideal for SLO/error-budget dashboards where you care about a **threshold**, not an arbitrary percentile. - For `DistributionSummary` the SLOs are magnitudes (e.g. `serviceLevelObjectives(1024, 1_000_000)` bytes) rather than durations. ### Bounding accuracy & cost - **`minimumExpectedValue` / `maximumExpectedValue`** set the value range the histogram covers. Tight, realistic bounds → fewer buckets, less memory, better precision inside the range. Values outside are clamped into the edge buckets. - Histograms trade **memory/series count** for **aggregatability + accuracy**; percentiles trade accuracy/aggregatability for **cheapness**. ## Spring Boot conveniences Rather than code, you can configure these via `application.properties` for the built-in `http.server.requests`: ```properties management.metrics.distribution.percentiles-histogram.http.server.requests=true management.metrics.distribution.slo.http.server.requests=100ms,500ms management.metrics.distribution.percentiles.http.server.requests=0.95,0.99 management.metrics.distribution.minimum-expected-value.http.server.requests=1ms management.metrics.distribution.maximum-expected-value.http.server.requests=10s ``` The property keys map one-to-one to the builder methods and accept per-meter-name overrides. ## Decision guide | Need | Use | |---|---| | Percentiles on a single instance / quick local view | `publishPercentiles` | | Accurate percentiles across many replicas | `publishPercentileHistogram` + backend `histogram_quantile` | | 'Fraction under threshold' / SLO dashboards | `serviceLevelObjectives` | | Control accuracy & memory | `minimum/maximumExpectedValue` | ## Common gotchas - **Averaging p99 gauges across pods** — the classic wrong dashboard. Switch to histograms. - **Enabling percentile histograms with unbounded expected-value range** — bucket/series explosion. Always set realistic min/max. - **Assuming histogram percentiles are exact** — they're interpolated within a bucket; boundary choice affects accuracy. - **Cardinality**: histograms multiply series per tag combination; combine with high-cardinality tags and you can overwhelm the TSDB. - Not every backend supports server-side histogram quantiles; some (e.g. certain hosted systems) prefer client-side percentiles.

  • Your service runs on 12 pods and the dashboard averages the per-pod p99 gauge. Why is it wrong and what's the fix?
    Percentiles aren't linearly aggregatable — averaging p99s across pods has no statistical meaning. Fix: use publishPercentileHistogram so each pod exports buckets, and compute the service-wide p99 in the backend (e.g. Prometheus histogram_quantile over summed buckets).
  • Why set minimumExpectedValue and maximumExpectedValue?
    They bound the histogram's value range so Micrometer allocates buckets efficiently — tight realistic bounds give better precision and fewer series/less memory; out-of-range values clamp into edge buckets.
  • When would serviceLevelObjectives be better than percentiles?
    When you care about a threshold, not an arbitrary quantile — e.g. 'what fraction of requests were under 500ms' for an SLO/error-budget dashboard. SLO buckets answer that directly and aggregate across instances.

saying these in an interview costs you the question

  • Averaging or summing per-instance p99 gauges across replicas
  • Claiming publishPercentiles produces aggregatable percentiles
  • Enabling percentile histograms without bounding the expected-value range
  • Thinking histogram_quantile gives exact (not interpolated) percentiles
  • Believing all backends support server-side histogram percentiles

context