skip to content

Mechanically, how does a retained VFS speed up Gradle's up-to-date checks compared with a cold snapshot?

level: seniorimportance: should knowfreq 28%

answer

  1. fingerprint = metadata + content hash
  2. cost = stat directories + hash bytes
  3. retained reuses cached hashes
  4. invalidate only OS-flagged paths
  5. feeds cache key + InputChanges

basics

~20 s

Up-to-date checks compare current input/output hashes to the last build's. A cold VFS re-stats and re-hashes every input; a retained VFS reuses cached hashes for unchanged files and only re-snapshots paths the OS reported changed.

solid answer

~50 s

An up-to-date check needs a **fingerprint** of each task's declared inputs and outputs — built from file metadata plus **content hashes** — and compares it to the fingerprint stored from the previous run. The cost is dominated by **stat-ing directories and hashing file content**. With a **cold VFS** (no daemon, or watching off), Gradle must walk and re-hash all declared inputs to rebuild those fingerprints. With a **retained VFS**, the daemon already holds metadata and hashes from last time; the OS watch backend has flagged exactly which paths changed since. Gradle **invalidates only those flagged paths**, re-snapshots them, and **reuses cached hashes** for everything else. The up-to-date comparison itself is unchanged — same hashes, same decision — but the *gathering* of input state reads far fewer files, which is why warm builds reach the task graph quickly. The build cache and incremental tasks then consume the very same snapshots, so VFS retention compounds with them.

code

kotlin · 7 lines
kotlin
// Declaring inputs/outputs is what makes a task snapshottable;
// the VFS just makes recomputing these fingerprints cheap on warm builds.
tasks.register<MyGenTask>("gen") {
    inputs.dir(layout.projectDirectory.dir("schemas"))
        .withPropertyName("schemas")
    outputs.dir(layout.buildDirectory.dir("generated"))
}

go deeper

for a junior

Know it reuses cached file hashes so fewer files are read.

for a middle

Explain stat+hash cost and that only changed paths are re-snapshotted.

for a senior

Articulate that the comparison is unchanged, how it feeds cache keys and InputChanges, and the correctness boundary on missed events.

for a principal

Reason about the compounding effect across cache + incremental builds and where it materially moves end-to-end build latency.

## What an up-to-date check actually computes For every task, Gradle has declared **inputs** (files, properties) and **outputs**. To decide `UP-TO-DATE`, it builds an **input/output fingerprint**: per input file it combines path/structure with a **content hash**; per task it folds those into a single key. It compares this fingerprint to the one persisted from the **last execution**. If they match (and outputs are intact), the task is skipped. The same fingerprints feed: - the **build cache key** (a cacheable task's hash of inputs), and - **incremental task** change sets (`@InputChanges` / `InputChanges`), which need to know *which* input files changed. So snapshotting input files is on the hot path for all three. ## Cost: stat + hash Producing a fingerprint requires: 1. **Walking** declared input directory trees (directory reads / `stat`). 2. **Hashing** file content (read bytes, compute the digest), unless an unchanged metadata signature lets Gradle skip re-hashing. On a repo with many inputs this filesystem I/O dominates the time between 'configuration done' and 'executing tasks'. ## Cold VFS vs retained VFS **Cold (no retention):** every build rebuilds the VFS from scratch — walk all input trees, re-hash content. Even a no-op build pays the full scan. **Retained + watching:** the daemon kept the VFS in memory and the OS told it precisely which paths changed: 1. For paths the OS **did not** flag, Gradle trusts the retained metadata/hash — **no disk read**. 2. For **flagged** paths, Gradle invalidates the VFS entry and re-snapshots just those. 3. Fingerprints are recomputed from the (mostly cached) hashes. The decision logic is identical; only the **data-gathering** got cheaper. ```text Build N: walk + hash all inputs -> VFS populated, fingerprints stored (idle) OS pushes change events -> daemon invalidates changed paths in VFS Build N+1: re-snapshot only changed -> reuse cached hashes for the rest ``` ## Why it compounds Because the retained snapshots are exactly what the **build cache** and **incremental task** machinery consume, a faster VFS makes those features start faster too. It does **not** change cache hits or task outputs — it changes how quickly Gradle can compute the inputs that decide them. ## Correctness boundary If the OS misses an event (rare) or a path isn't watched (out-of-root, symlink target, network mount), Gradle re-snapshots that path normally, so a missed event can never produce a stale up-to-date result for watched-and-flagged content; un-watched content is simply always re-read.

  • Does VFS watching change which tasks are considered up-to-date?
    No. The up-to-date decision is the same hash comparison either way. Watching only changes how cheaply Gradle gathers the input/output state used in that comparison.
  • How does retained VFS relate to the build cache and incremental tasks?
    All three consume the same input snapshots/hashes. A retained VFS makes producing those snapshots cheaper, so the cache-key computation and InputChanges change-set start faster, without altering hits or outputs.
  • Could a missed OS event cause a stale up-to-date result?
    For un-watched paths Gradle always re-snapshots, so they're safe. The risk is only for watched paths if the OS drops an event; Gradle mitigates by re-validating on registration changes and falling back to full snapshots when watching is unreliable.

saying these in an interview costs you the question

  • Saying VFS retention changes cache hit rates or task outputs.
  • Confusing the fingerprint comparison (always done) with the snapshot gathering (what VFS accelerates).
  • Claiming it replaces declaring task inputs/outputs.

context