skip to content

Beyond declaring inputs, what makes a task's outputs safe to cache, and how would you make a task that currently produces non-reproducible outputs cacheable?

level: principalimportance: should knowfreq 28%

answer

  1. key stable AND outputs reproducible
  2. no timestamps/abs paths/random
  3. preserveFileTimestamps=false
  4. reproducibleFileOrder=true
  5. validatePlugins gate + remote rollout

basics

~20 s

Outputs must be reproducible: the same inputs must always yield the same bytes, with no timestamps, absolute paths, or randomness baked in. Make archives reproducible (fixed timestamps, sorted entries) and remove machine-specific data before marking the task cacheable.

solid answer

~40 s

Correctly-declared inputs get you a stable *key*; cache *safety* also needs **reproducible outputs** — identical inputs producing identical output bytes. If a task embeds wall-clock timestamps, absolute paths, hostnames, or non-deterministic ordering, then a cache hit replays stale or misleading content and you can't trust shared/remote reuse. To make such a task cacheable: strip non-determinism — use Gradle's reproducible-archive settings (`isPreserveFileTimestamps = false`, `isReproducibleFileOrder = true` on `AbstractArchiveTask`), pin timestamps/versions instead of reading the clock, sort generated output, and avoid absolute paths in generated files. Only then add `@CacheableTask` with proper input annotations and path sensitivity. At org scale this becomes a policy: cacheability is a *correctness* claim, so it should be gated by `validatePlugins`, reproducibility checks, and review, not added casually for a quick speed-up.

code

kotlin · 15 lines
kotlin
// project-wide reproducible archives (a precondition for safely caching them)
tasks.withType<AbstractArchiveTask>().configureEach {
    isPreserveFileTimestamps = false
    isReproducibleFileOrder = true
}

// then a codegen task is safe to mark cacheable
@CacheableTask
abstract class GenSources : DefaultTask() {
    @get:Input abstract val version: Property<String>
    @get:InputFiles
    @get:PathSensitive(PathSensitivity.RELATIVE)
    abstract val schema: ConfigurableFileCollection
    @get:OutputDirectory abstract val outDir: DirectoryProperty
}

go deeper

for a junior

Know outputs must be the same for the same inputs (no timestamps/randomness).

for a middle

Name concrete non-reproducibility sources and the archive settings that fix them.

for a senior

Lay out the two-halves model (correct inputs + reproducible outputs) and how to remediate a task before caching it.

for a principal

Operationalize cacheability as a correctness claim: validatePlugins gates, reproducible-build defaults, staged remote-cache rollout, and hit-rate monitoring across teams.

## Two halves of cache safety Caching reuses a task's outputs whenever its key matches. That's only sound if **both** halves hold: 1. **Complete, correct inputs** — every value that affects the output is a declared input (so the key changes when it should). Covered by `@Input`/`@InputFiles`/path sensitivity. 2. **Reproducible outputs** — the same inputs always produce byte-identical outputs. Without this, a cache hit can silently replay content that differs from what a fresh run would produce, and remote reuse across machines becomes unreliable. `@CacheableTask` is your assertion that *both* hold. Gradle can check inputs are annotated; it cannot verify determinism — that's on the author. ## Common sources of non-reproducibility - **Timestamps** in archives (`Jar`/`Zip` store entry mtimes) or in generated files (`Built-On: <date>`). - **Absolute paths** written into generated code/manifests. - **Unstable ordering** — iterating a `HashSet`, filesystem order, or parallel writers. - **Hostnames / user names / random UUIDs** captured into output. ## Making outputs reproducible ```kotlin // reproducible archives tasks.withType<AbstractArchiveTask>().configureEach { isPreserveFileTimestamps = false isReproducibleFileOrder = true } // a generated-source task: pin time, sort, avoid abs paths @CacheableTask abstract class GenSources : DefaultTask() { @get:Input abstract val version: Property<String> @get:InputFiles @get:PathSensitive(PathSensitivity.RELATIVE) abstract val schema: ConfigurableFileCollection @get:OutputDirectory abstract val outDir: DirectoryProperty @TaskAction fun run() { // write with a fixed header, deterministic ordering, relative paths only } } ``` ## Org-level governance Because cacheability is a correctness claim, treat it as one: - Gate plugins with **`validatePlugins`** in CI so missing annotations/path sensitivity fail the build. - Adopt **reproducible builds** defaults project-wide (the archive settings above). - When enabling a **remote** cache, start with high-value deterministic tasks (compile, test, codegen) and require evidence of reproducibility before broadening. - Watch cache **hit rates** and investigate divergence with build scans — a sudden drop usually means a new non-reproducible input crept in. ## Why principal-level The trap is treating caching as a free speed knob. A non-reproducible task marked cacheable produces *correctness* incidents (stale outputs reused across CI) that are painful to debug. The senior/principal judgment is knowing that the annotation is a guarantee you must engineer for, and operationalizing that guarantee across many teams.

  • Gradle can validate input annotations — why can't it just verify outputs are reproducible too?
    Reproducibility depends on arbitrary task logic (clocks, ordering, network). Gradle can't statically prove determinism, so it makes the author assert it via @CacheableTask and relies on conventions/reviews to enforce it.
  • How would you detect that a non-reproducible input slipped into a previously-good cacheable task?
    Monitor remote cache hit rate; a drop signals divergence. Compare two build scans' task cache keys/input fingerprints to find the newly-varying input, or run the task twice and diff outputs byte-for-byte.
  • What's a safe rollout order when introducing a remote cache org-wide?
    Start with expensive, deterministic tasks (compile, test, codegen) that you've verified reproducible; gate with validatePlugins; watch hit rates; then expand. Avoid caching cheap I/O tasks where overhead outweighs gains.

Caching is like reusing a photocopy: only safe if the original is always printed identically. If each print silently adds today's date, sharing copies spreads wrong versions.

saying these in an interview costs you the question

  • Believing correctly-declared inputs alone make a task safe to cache (ignores output determinism).
  • Marking archive tasks cacheable without reproducible-archive settings.
  • Treating @CacheableTask as a pure performance flag rather than a correctness guarantee.

context