skip to content

What is a LongTaskTimer and how does it differ from a regular Timer?

level: middleimportance: should knowfreq 50%

answer

  1. measures IN-FLIGHT running tasks
  2. active count + duration of still-running
  3. long jobs / detect stuck work mid-run
  4. @Timed(longTask = true)
  5. must stop() or task leaks as active

basics

~10 s

A LongTaskTimer measures tasks that are still running. It reports how many are in flight and their current accumulated duration, updated while they run. A regular Timer only records after an operation finishes.

solid answer

~50 s

A regular Timer records a duration only once an operation completes, so a job that runs for 10 minutes shows nothing until minute 10 — useless for alerting mid-flight. A LongTaskTimer measures **in-progress** work: it exposes the number of currently active tasks (active count) and the total/max duration of those still-running tasks, sampled continuously. You use it for long-running things like scheduled batch jobs, imports, or Kafka consumers where you want to know 'a job has been stuck running for 8 minutes' right now. API-wise you get one via LongTaskTimer.builder(...).register(registry) (or the @Timed(longTask = true) annotation), then wrap work with longTaskTimer.record(Runnable) or start()/stop() a Sample. Each active execution is tracked independently, so concurrent runs are all visible. It is not a replacement for Timer — you often use both: LongTaskTimer for live in-flight visibility, Timer for completed-latency distributions.

code

java · 24 lines
java
import io.micrometer.core.instrument.LongTaskTimer;
import io.micrometer.core.instrument.MeterRegistry;
import org.springframework.scheduling.annotation.Scheduled;
import org.springframework.stereotype.Component;

@Component
class NightlyImportJob {
    private final LongTaskTimer inFlight;

    NightlyImportJob(MeterRegistry registry) {
        this.inFlight = LongTaskTimer.builder("batch.import.running")
            .description("Nightly imports currently executing")
            .register(registry);
    }

    @Scheduled(cron = "0 0 2 * * *")
    void run() {
        // While this runs, active count = 1 and duration climbs live,
        // so you can alert if it overruns before it ever finishes.
        inFlight.record(this::doImport);
    }

    private void doImport() { /* long-running work */ }
}

go deeper

for a junior

Know it exists for long jobs and reports in-flight duration.

for a middle

Explain active count vs Timer's completion-only recording and the @Timed(longTask=true) hook.

for a senior

Design alerting on overruns, know you often register both meters, and the leak risk of missing stop().

for a principal

Reason about when to prefer the Observation API, cardinality of active-task series, and monitoring semantics for parallel executions.

## The gap a LongTaskTimer fills A plain `Timer` only emits a measurement **when the operation ends** — it records the final duration. That's fine for millisecond requests, but consider a nightly batch that runs 30 minutes. With a `Timer`, for those 30 minutes your metrics show *nothing new*; you only learn the duration after it completes. If it hangs forever, the Timer never records and you can't alert on 'it's been running too long.' **`LongTaskTimer`** solves this by measuring **currently-executing** tasks. While tasks run it can report: - **active count** — how many executions are in flight right now (gauge-like). - **duration** — the total (and max) elapsed time of the tasks that are *still running*, sampled live. So at any scrape you can see 'there are 2 imports active, the longest has been running 8m12s' and alert on that threshold — before completion. ## API ```java LongTaskTimer scrapeTimer = LongTaskTimer.builder("batch.import.active") .description("in-flight nightly imports") .register(registry); // Option A: wrap scrapeTimer.record(() -> runImport()); // Option B: explicit sample (start when task begins, stop when it ends) LongTaskTimer.Sample sample = scrapeTimer.start(); try { runImport(); } finally { sample.stop(); // returns duration; removes it from active set } ``` Each `start()` registers an independent active task; `stop()` removes it. Concurrent executions are all tracked, so active count reflects true parallelism. ## Spring integration Annotate a long method with `@Timed(value = "...", longTask = true)` and Spring's `TimedAspect` wraps it in a LongTaskTimer instead of a Timer. Scheduled jobs and message listeners are typical candidates. ## Timer vs LongTaskTimer | | Timer | LongTaskTimer | |---|---|---| | Records when | task **completes** | **while task runs** (sampled) | | Key stats | count, total, max, (percentiles) | active count, duration of running tasks, max | | Good for | request/short-op latency + distribution | long jobs, detecting stuck/overrunning work | | A 30-min hung job | invisible until done/never | visible immediately as active + rising duration | ## Gotchas - **They answer different questions** — don't replace Timer with LongTaskTimer or vice versa. A completed LongTaskTimer task is *removed* from the active set, so it does **not** give you a historical latency distribution of finished tasks; if you want both live visibility and completed-duration percentiles, register **both** meters (or use the Observation API which can emit both). - **Must call stop()** (or use record(Runnable)) or the task stays 'active' forever, inflating active count and duration — a leak. - Percentiles/histograms are supported on LongTaskTimer too (distribution of currently-running durations), but the common use is active count + max duration for alerting. - Base unit and clock behavior mirror Timer (nanos internally, sampled against the registry Clock). ## When to use Use a LongTaskTimer for scheduled batch jobs, long imports/exports, migrations, streaming consumers, or any workload where you must observe it *while it runs* — especially to alert on overruns/stuck tasks. Keep a normal Timer alongside when you also care about the latency distribution of completed runs.

  • You have a 20-minute job and want to alert if it exceeds 15 minutes. Which meter, and why not a plain Timer?
    LongTaskTimer — it exposes the duration of the still-running task live, so you can alert at minute 15. A plain Timer records only on completion, so it can't fire while the job is still (over)running.
  • Does a LongTaskTimer give you the latency distribution of completed jobs?
    No — once a task stops it's removed from the active set. For completed-duration percentiles you need a regular Timer (register both if you want live visibility and history).

saying these in an interview costs you the question

  • Saying LongTaskTimer records the final duration like Timer does
  • Believing it replaces Timer for latency distributions
  • Forgetting stop() causes a leaked ever-active task
  • Thinking active count is just the completed count

context