skip to content

How would you implement a build-wide per-task timing report, and what pitfalls arise under parallel execution?

level: seniorimportance: should knowfreq 30%

answer

  1. TaskFinishEvent.result.startTime/endTime
  2. BuildService holds the aggregate
  3. ConcurrentHashMap — parallel onFinish
  4. report in close()/AutoCloseable
  5. legacy afterTask: cache-incompatible + manual timing

basics

~10 s

Implement an OperationCompletionListener whose onFinish handles TaskFinishEvent — each event already carries start and end times, so you record duration per task path. Aggregate in a thread-safe BuildService since tasks run in parallel.

solid answer

~40 s

The clean way is a `BuildService` that implements `OperationCompletionListener`, registered via `BuildEventsListenerRegistry.onTaskCompletion(...)`. In `onFinish`, guard `if (event is TaskFinishEvent)` and read `event.result.startTime` / `endTime` — Gradle already measures the duration, so you don't need `beforeExecute`/`afterExecute` to stash start times yourself. You accumulate `taskPath -> durationMs` in a **thread-safe** structure (e.g. `ConcurrentHashMap` or synchronized aggregation) because with `--parallel` and the worker API, multiple `onFinish` callbacks can arrive concurrently. At build end you emit a sorted report (in the service's `close()` / `AutoCloseable`). The legacy alternative — `taskGraph.afterTask` storing nanoTime on extra properties — works but is cache-incompatible and forces you to manage timing manually. Either way, the central pitfall is shared-state mutation under concurrency; the `BuildService` model gives a single managed instance with controlled lifecycle to hold that aggregate safely.

code

kotlin · 6 lines
kotlin
override fun onFinish(event: FinishEvent) {
    if (event is TaskFinishEvent) {
        val r = event.result
        durations[event.descriptor.taskPath] = r.endTime - r.startTime  // ms already measured
    }
}

go deeper

for a junior

Know that you can observe task durations via a listener; details of thread-safety are beyond junior scope.

for a middle

Describe reading start/end times from TaskFinishEvent and aggregating per task path.

for a senior

Cover the BuildService wiring, configuration-cache compatibility, and thread-safe aggregation under --parallel.

for a principal

Provide profiling as a reusable cache-safe convention plugin, and weigh build-scan/Gradle Enterprise instrumentation versus a homegrown listener for org-wide observability.

## Goal Produce "task X took Y ms" for the whole build and a sorted summary at the end — the classic build-profiling feature. ## Preferred implementation: BuildService + OperationCompletionListener The Tooling-API event already contains timing, so you don't have to bracket each task yourself: ```kotlin abstract class TaskTimers : BuildService<BuildServiceParameters.None>, OperationCompletionListener, AutoCloseable { private val durations = java.util.concurrent.ConcurrentHashMap<String, Long>() override fun onFinish(event: FinishEvent) { if (event is TaskFinishEvent) { val r = event.result durations[event.descriptor.taskPath] = r.endTime - r.startTime } } override fun close() { durations.entries.sortedByDescending { it.value } .forEach { println("${it.value} ms ${it.key}") } } } abstract class TimingPlugin : Plugin<Project> { @get:Inject abstract val registry: BuildEventsListenerRegistry override fun apply(project: Project) { val svc = project.gradle.sharedServices .registerIfAbsent("taskTimers", TaskTimers::class.java) {} registry.onTaskCompletion(svc) } } ``` ## Why a BuildService A shared `BuildService` is instantiated **once** for the build and managed by Gradle. Registering it for completion events keeps the whole thing configuration-cache compatible, and its single instance is the right place to hold the cross-task aggregate. Its `close()` runs at build end — a natural spot to print the report. ## The concurrency pitfall With `org.gradle.parallel=true`, tasks from independent projects run on multiple threads, and the worker API adds more parallelism. `onFinish` can therefore be invoked **concurrently from different threads**. If you aggregate into a plain `HashMap` or a non-atomic counter, you get lost updates or `ConcurrentModificationException`. Use `ConcurrentHashMap`, `AtomicLong`, or synchronize. This is the number-one mistake in homegrown profilers. ## The legacy approach and its problems ```kotlin val starts = mutableMapOf<String, Long>() gradle.taskGraph.beforeTask { starts[it.path] = System.nanoTime() } gradle.taskGraph.afterTask { val ms = (System.nanoTime() - starts[it.path]!!) / 1_000_000 } ``` This works only without the configuration cache, captures a mutable map across the build, and you must measure time yourself. It also shares the same concurrency hazard. Acceptable for a quick local hack, wrong for shared build logic. ## Distinguishing real work For a meaningful report, you may want to mark cache hits and up-to-date tasks: from a `TaskSuccessResult` check `isFromCache()` / `isUpToDate()` so a 0 ms cached task isn't mistaken for a fast real run. ## Summary Use `OperationCompletionListener` in a `BuildService`, read timings straight off `TaskFinishEvent`, aggregate in a thread-safe container, and print in `close()`. Treat parallel callback delivery as the key correctness concern.

  • Why don't you need beforeExecute to capture start times with OperationCompletionListener?
    TaskFinishEvent's result already carries startTime and endTime measured by Gradle, so the duration is endTime - startTime — no manual bracketing needed.
  • What concurrency hazard appears when --parallel is enabled?
    onFinish callbacks can be delivered from multiple threads simultaneously, so a non-thread-safe aggregate (plain HashMap, non-atomic counter) risks lost updates or ConcurrentModificationException; use ConcurrentHashMap/atomics.
  • How do you avoid counting a cached task as a real fast execution?
    Check the TaskSuccessResult's isFromCache()/isUpToDate() and label or exclude those entries in the report.

saying these in an interview costs you the question

  • Aggregating timings in a plain HashMap shared across parallel callbacks.
  • Re-implementing timing with nanoTime in beforeExecute/afterExecute when the event already provides it.
  • Assuming a single-threaded callback model — parallel builds break that assumption.

context