Given the inter-project task DAG, how does Gradle exploit parallelism and incrementality, and what limits it?
answer
- DAG enables parallel + incremental
- org.gradle.parallel + max-workers
- up-to-date via input/output hashing
- build cache keyed on input hashes
- coarse deps / leaky api / cycles limit gains
basics
~20 sThe DAG lets Gradle run independent tasks in parallel (--parallel) and skip up-to-date ones via input/output checks. Producer→consumer edges serialize what must be ordered; only the affected slice rebuilds. Cycles or coarse dependencies kill parallelism.
solid answer
~50 sThe induced DAG is what makes parallel, incremental builds correct. With `org.gradle.parallel=true`, Gradle executes tasks with no dependency relationship concurrently across projects, while still honoring producer→consumer edges (so `:lib:jar` finishes before `:app:compileJava`). Incrementality comes from each task's declared inputs/outputs and up-to-date checks: change :app only and :lib stays up-to-date, so its tasks are skipped. The build cache extends this across machines/branches by keying task outputs on input hashes. Limits: a *false* or overly coarse dependency (e.g. depending on a whole project when you need one configuration) serializes work that could run in parallel and widens the rebuild blast radius. Cyclic project dependencies are illegal. Tasks with undeclared inputs/outputs break up-to-date checks. Configuration-time work that touches other projects (cross-project configuration) undermines the configuration cache. The goal: a lean, accurate DAG so the scheduler can maximize concurrency and minimize rebuilt nodes.
code
toml · 7 lines# gradle.properties
org.gradle.parallel=true
org.gradle.caching=true
org.gradle.configuration-cache=true
org.gradle.workers.max=8
# Now independent :lib / :other-lib tasks run concurrently;
# unchanged projects stay UP-TO-DATE; outputs are cache-keyed on inputs.go deeper
Know --parallel runs independent tasks at once and unchanged projects are skipped as up-to-date.
Explain how the DAG permits parallelism while preserving ordering, and how input/output declarations drive up-to-date skipping.
Diagnose limited speedup via over-serialized/leaky edges; relate build cache, configuration cache, and narrow consumption to blast radius.
Set org-wide conventions (lean api surface, narrow cross-project consumption, declared I/O) and measure DAG width/critical path to maximize CI parallelism.
## The DAG is the scheduling substrate Every downstream optimization rides on an accurate task DAG: - **Topological correctness**: producers run before consumers. - **Parallelism**: nodes with no path between them can run simultaneously. - **Incrementality**: nodes whose inputs are unchanged are skipped. ## Parallel execution ```properties # gradle.properties org.gradle.parallel=true org.gradle.workers.max=8 ``` With parallel mode, Gradle runs tasks from *different* projects concurrently when no edge connects them. `:lib:jar` and `:other-lib:jar` (siblings :app both depends on) can build at once; `:app:compileJava` waits for both. The `--max-workers` count bounds concurrency. Intra-task parallelism (e.g. the Worker API) is separate but composes. ## Incremental builds & up-to-date Each task declares **inputs** (`@InputFiles`, `@Input`) and **outputs** (`@OutputDirectory`). Before running, Gradle hashes inputs; if inputs and outputs match the prior run, the task is **UP-TO-DATE** and skipped. Change only :app's sources and :lib's `compileJava`/`jar` are up-to-date → skipped → :app rebuilds against the unchanged :lib jar. ## Build cache across the graph ```properties org.gradle.caching=true ``` For `@CacheableTask`s, outputs are stored keyed by input hashes. A teammate (or CI) that compiles the same :lib inputs fetches outputs from the cache instead of recompiling — incrementality that survives clean checkouts and branch switches. ## What limits these gains 1. **Coarse/false dependencies** — `implementation(project(':lib'))` when you only need one artifact pulls :lib's whole api closure onto your compile cp, enlarging the rebuild blast radius and serializing more. Prefer narrow consumption (specific configuration / `api` discipline). 2. **Leaky api** — over-using `api` in libraries widens consumers' compile classpaths, so changes recompile more modules. 3. **Cycles** — a cyclic project dependency is an error; no valid topological order. 4. **Undeclared inputs/outputs** — breaks up-to-date and caching (task always reruns or, worse, produces stale results). 5. **Cross-project configuration** at configuration time (reaching into another project's tasks/extensions) hurts the configuration cache and parallel configuration; prefer isolated, lazy wiring (`Provider`, `tasks.named`). ## Takeaway The build engineer's job is to keep the DAG **lean and accurate**: minimal edges, well-declared task I/O, narrow cross-project consumption. The scheduler then gives you maximum parallelism and minimum rebuilt nodes for free.
- You enabled --parallel but see little speedup. What graph issue would explain it?Over-serialized dependencies: too many projects funnel through a single producer, or coarse/leaky api edges chain everything together, leaving few independent nodes for the scheduler to run concurrently.
- Why does a cyclic project dependency fail the whole model?The task graph must be a DAG; a cycle has no valid topological order, so Gradle can't determine which producer runs first and reports the cycle as an error.
- How does the build cache differ from up-to-date checks?Up-to-date checks skip a task when its own prior outputs match current inputs in this build dir. The build cache reuses outputs across builds/machines by storing them keyed on input hashes, surviving clean checkouts.
saying these in an interview costs you the question
- Claiming --parallel reorders tasks and can violate producer→consumer ordering.
- Thinking enabling caching fixes builds that have undeclared task inputs/outputs.
- Ignoring that leaky api and coarse project deps enlarge the rebuild blast radius.