Explain the longTask, percentiles, and histogram attributes of @Timed and when you'd use each.
answer
- longTask -> LongTaskTimer -> in-flight/active count
- percentiles = client-side, cheap, NOT aggregatable
- histogram = buckets, backend computes, aggregatable across instances
- @Timed is @Repeatable: normal + longTask on same method
- watch time-series cardinality
basics
~10 slongTask=true measures long-running, still-in-progress calls with a LongTaskTimer (active count + in-flight duration). percentiles computes client-side percentile values (e.g. p95, p99). histogram=true publishes bucket counts so backends like Prometheus can aggregate percentiles across instances.
solid answer
~50 sBy default @Timed creates a regular `Timer` that records completed calls. `longTask = true` switches it to a `LongTaskTimer`: instead of recording after completion, it tracks tasks that are **currently running** — how many are active and their in-flight duration — ideal for long jobs (batch, imports) where you want to see 'it's been running 40 minutes right now'. `percentiles = {0.95, 0.99}` publishes **client-side pre-computed** percentiles as extra gauges (`.percentile` tagged series); they're cheap but **not aggregatable** across instances or re-computable for other quantiles. `histogram = true` (backed by `percentileHistogram`) publishes **bucket counts** the monitoring backend uses to compute percentiles server-side — aggregatable across instances and flexible on quantile, at the cost of more time series. Use client percentiles for quick per-instance insight; use histogram buckets when you aggregate across a fleet in Prometheus/Grafana.
code
java · 13 lines@Service
public class ReportService {
// In-flight tracking for a long batch job
@Timed(value = "report.generate.active", longTask = true)
// Plus completion timing with aggregatable buckets and cheap client percentiles
@Timed(value = "report.generate", histogram = true,
percentiles = {0.95, 0.99})
public Report generate(ReportSpec spec) {
// long-running work...
return new Report();
}
}go deeper
Just know these attributes exist to add percentiles/histograms and to track long-running calls.
Explain longTask=LongTaskTimer and the basic difference between percentiles and histogram.
Articulate aggregation semantics (why client percentiles don't aggregate) and cardinality trade-offs.
Set org conventions: histogram for SLO/fleet dashboards, bounded buckets, and cardinality budgets.
## Regular Timer vs LongTaskTimer (`longTask`) A normal `Timer` records the duration of a call **after it finishes**. That's useless for a task that runs for minutes/hours, because you learn nothing until it's over, and if the app crashes mid-task you never record it. `@Timed(value = "import.job", longTask = true)` creates a **`LongTaskTimer`** instead. A LongTaskTimer measures **in-flight** executions: - **active count** — how many invocations are running right now. - **duration** — total/max time of the currently-running invocations. So you can alert on 'a job has been active longer than N minutes' or 'more than K concurrent long tasks'. Typical uses: scheduled batch jobs, file imports, report generation, long polling. You can apply both a normal timer and a long-task timer to the same method by using the array form: `@Timed` twice? No — instead set `longTask` on one annotation; to get **both** completion timing and in-flight tracking you place two `@Timed` annotations (the annotation is `@Repeatable`), one with `longTask=false` and one with `longTask=true`. ## Client-side percentiles (`percentiles`) `@Timed(value = "orders.place", percentiles = {0.5, 0.95, 0.99})` tells Micrometer to **pre-compute** those quantiles inside the process using a rolling window, and publish each as a separate meter (for Prometheus, series suffixed and tagged with `quantile`). Trade-offs: - **Cheap and immediate** per instance. - **NOT aggregatable**: you cannot average p99 across 10 pods to get a fleet p99 — percentiles don't sum/average. - **Fixed quantiles**: you only get the ones you declared; you can't ask for p90 later. ## Histogram buckets (`histogram`) `@Timed(value = "orders.place", histogram = true)` enables a **percentile histogram**: Micrometer publishes a set of **cumulative bucket counts** (`_bucket{le=...}`). The backend (Prometheus `histogram_quantile`, Grafana) computes percentiles **server-side** from buckets. Trade-offs: - **Aggregatable across instances** — buckets sum, so a fleet-wide p99 is correct. - **Flexible** — any quantile computable after the fact. - **More time series** (one per bucket) → higher storage/cardinality cost. You can also tune bucket ranges with `serviceLevelObjectives` (SLO boundaries) via a `MeterFilter` or the `@Timed` distribution settings, and set `minimumExpectedValue` / `maximumExpectedValue` to bound bucket generation. ## Choosing | Need | Use | |---|---| | Duration of completed calls | default Timer | | Watch still-running long jobs | `longTask = true` | | Quick per-instance p95/p99, low cost | `percentiles = {...}` | | Correct percentiles aggregated across many instances | `histogram = true` | Often you combine `histogram = true` for aggregation with a couple of `percentiles` for cheap per-instance dashboards. Just remember client percentiles from different instances can't be combined. ## Gotcha Declaring many quantiles or histogram on high-throughput methods multiplies time-series count; watch cardinality, especially combined with `extraTags`.
- Why can't you average p99 values reported by 10 instances to get the cluster p99?Percentiles are non-linear order statistics; the average of per-instance p99s is not the p99 of the combined distribution. To get a correct fleet percentile you need histogram buckets that sum across instances, then compute the quantile from the aggregate.
- What does a LongTaskTimer record that a normal Timer cannot?The duration and count of invocations that are still in progress. A normal Timer only records after a call completes, so an ongoing (or crashed) long task is invisible until/unless it finishes.
saying these in an interview costs you the question
- Saying client-side percentiles can be aggregated across instances.
- Thinking histogram=true and percentiles={...} are the same thing.
- Believing longTask changes only the metric name rather than the meter type/semantics.