skip to content

Why does Gradle compute the entire task graph before executing any task, and what capabilities does this 'plan-then-run' model enable?

level: seniorimportance: should knowfreq 38%

answer

  1. global view of all work + ordering
  2. parallel: run independent tasks together
  3. inputs/outputs => up-to-date + build cache
  4. fail-fast on cycles before side effects
  5. configuration cache serializes the graph

basics

~20 s

Knowing the whole graph up front lets Gradle schedule independent tasks in parallel, decide up-to-date/cached status correctly, fail fast on cycles, and report progress — because it sees all the work and its ordering before running anything.

solid answer

~40 s

Gradle deliberately resolves the **complete** DAG at the end of configuration so it can reason about the *entire* plan, not discover work as it goes. This enables several capabilities: **parallel execution** (`--parallel`/worker API) can run tasks with no edge between them concurrently because the scheduler knows which tasks are independent; **up-to-date checking and the build cache** use each task's declared inputs/outputs — visible from the graph — to skip or restore work; **fail-fast on cycles** happens before any action runs; and **accurate progress/reporting** (and `--dry-run`) is possible because the full task list and order are known. It also underpins the **configuration cache**, which can serialize the computed graph and skip configuration entirely on later runs. In short, the global view is what makes Gradle's incrementality, caching, and parallelism sound.

code

bash · 3 lines
bash
gradle build --parallel --build-cache --configuration-cache
# parallel scheduling, cache reuse, and skipping configuration on reruns
# are all enabled by computing the full task graph before execution

go deeper

for a junior

State that knowing the whole graph lets Gradle skip up-to-date tasks and run independent ones in parallel.

for a middle

Connect declared inputs/outputs and the DAG to up-to-date checks, build cache, and parallel scheduling.

for a senior

Explain configuration cache, fail-fast, and the configuration-time trade-off of plan-then-run.

for a principal

Weigh configuration cost against caching/parallel gains and set build-wide conventions to keep the up-front planning cheap at scale.

## Why plan everything first? Gradle could, in principle, run tasks as it discovers them. Instead it builds the **whole** task graph after configuration. The payoff is a **global view** of all work and its ordering, which unlocks optimizations that are impossible with incremental discovery. ## Capabilities enabled ### 1. Parallel execution With `--parallel` (and the **Worker API**), Gradle runs tasks that have **no dependency edge** between them at the same time. Only a complete graph reveals which tasks are mutually independent and which must be serialized. ### 2. Up-to-date checks & build cache Each task declares **inputs/outputs**. Knowing them across the whole graph lets Gradle: - mark a task **UP-TO-DATE** and skip it when inputs/outputs are unchanged; - pull outputs from the **build cache** (`@CacheableTask`) keyed by input hashes — even across machines. The graph tells Gradle the producer→consumer relationships so cached results stay consistent. ### 3. Fail-fast on cycles Because the DAG is validated up front, a circular dependency aborts **before** any side-effecting action runs — no half-built state. ### 4. Reporting, dry-run, and continue `--dry-run` prints the plan; progress bars and `--continue` (keep going after a failure to maximize work done) all rely on knowing the remaining graph. ### 5. Configuration cache The **configuration cache** serializes the computed task graph and task state. On the next compatible run, Gradle **skips initialization and configuration**, deserializing the graph and going straight to execution — a major speedup for large builds. ```bash gradle build --parallel --build-cache --configuration-cache # parallel scheduling + cache reuse + skip configuration on reruns, # all enabled by having the full graph computed up front ``` ## The trade-off The cost is that **all** scripts must be configured before anything runs (configuration time). This is exactly why configuration avoidance (`tasks.register` over `create`) and the configuration cache exist — to keep that up-front cost small while preserving the global-plan benefits.

  • How does the graph enable parallel execution to be safe?
    The scheduler runs only tasks with no dependency edge between them concurrently. Because the full DAG is known, Gradle can prove which tasks are independent and must never reorder ones connected by an edge.
  • What does the configuration cache do with the computed graph?
    It serializes the configured task graph and state to disk. On a later compatible invocation Gradle deserializes it and skips initialization and configuration entirely, jumping straight to execution.
  • What is the cost of the plan-then-run model?
    Every project's scripts must be configured before any task runs, so configuration time scales with build size — which is why configuration avoidance and the configuration cache exist.

A flight scheduler that sees every flight and connection at once can route planes in parallel and reuse aircraft efficiently — far better than dispatching each plane only when the previous one lands.

saying these in an interview costs you the question

  • Saying Gradle discovers and runs tasks incrementally as it goes (it plans the full graph first).
  • Claiming parallelism works without knowing the full graph.
  • Confusing the build cache (task outputs) with the configuration cache (serialized graph).

context