Beyond declaring inputs, what makes a task's outputs safe to cache, and how would you make a task that currently produces non-reproducible outputs cacheable?
answer
- key stable AND outputs reproducible
- no timestamps/abs paths/random
- preserveFileTimestamps=false
- reproducibleFileOrder=true
- validatePlugins gate + remote rollout
basics
~20 sOutputs must be reproducible: the same inputs must always yield the same bytes, with no timestamps, absolute paths, or randomness baked in. Make archives reproducible (fixed timestamps, sorted entries) and remove machine-specific data before marking the task cacheable.
solid answer
~40 sCorrectly-declared inputs get you a stable *key*; cache *safety* also needs **reproducible outputs** — identical inputs producing identical output bytes. If a task embeds wall-clock timestamps, absolute paths, hostnames, or non-deterministic ordering, then a cache hit replays stale or misleading content and you can't trust shared/remote reuse. To make such a task cacheable: strip non-determinism — use Gradle's reproducible-archive settings (`isPreserveFileTimestamps = false`, `isReproducibleFileOrder = true` on `AbstractArchiveTask`), pin timestamps/versions instead of reading the clock, sort generated output, and avoid absolute paths in generated files. Only then add `@CacheableTask` with proper input annotations and path sensitivity. At org scale this becomes a policy: cacheability is a *correctness* claim, so it should be gated by `validatePlugins`, reproducibility checks, and review, not added casually for a quick speed-up.
code
kotlin · 15 lines// project-wide reproducible archives (a precondition for safely caching them)
tasks.withType<AbstractArchiveTask>().configureEach {
isPreserveFileTimestamps = false
isReproducibleFileOrder = true
}
// then a codegen task is safe to mark cacheable
@CacheableTask
abstract class GenSources : DefaultTask() {
@get:Input abstract val version: Property<String>
@get:InputFiles
@get:PathSensitive(PathSensitivity.RELATIVE)
abstract val schema: ConfigurableFileCollection
@get:OutputDirectory abstract val outDir: DirectoryProperty
}go deeper
Know outputs must be the same for the same inputs (no timestamps/randomness).
Name concrete non-reproducibility sources and the archive settings that fix them.
Lay out the two-halves model (correct inputs + reproducible outputs) and how to remediate a task before caching it.
Operationalize cacheability as a correctness claim: validatePlugins gates, reproducible-build defaults, staged remote-cache rollout, and hit-rate monitoring across teams.
## Two halves of cache safety Caching reuses a task's outputs whenever its key matches. That's only sound if **both** halves hold: 1. **Complete, correct inputs** — every value that affects the output is a declared input (so the key changes when it should). Covered by `@Input`/`@InputFiles`/path sensitivity. 2. **Reproducible outputs** — the same inputs always produce byte-identical outputs. Without this, a cache hit can silently replay content that differs from what a fresh run would produce, and remote reuse across machines becomes unreliable. `@CacheableTask` is your assertion that *both* hold. Gradle can check inputs are annotated; it cannot verify determinism — that's on the author. ## Common sources of non-reproducibility - **Timestamps** in archives (`Jar`/`Zip` store entry mtimes) or in generated files (`Built-On: <date>`). - **Absolute paths** written into generated code/manifests. - **Unstable ordering** — iterating a `HashSet`, filesystem order, or parallel writers. - **Hostnames / user names / random UUIDs** captured into output. ## Making outputs reproducible ```kotlin // reproducible archives tasks.withType<AbstractArchiveTask>().configureEach { isPreserveFileTimestamps = false isReproducibleFileOrder = true } // a generated-source task: pin time, sort, avoid abs paths @CacheableTask abstract class GenSources : DefaultTask() { @get:Input abstract val version: Property<String> @get:InputFiles @get:PathSensitive(PathSensitivity.RELATIVE) abstract val schema: ConfigurableFileCollection @get:OutputDirectory abstract val outDir: DirectoryProperty @TaskAction fun run() { // write with a fixed header, deterministic ordering, relative paths only } } ``` ## Org-level governance Because cacheability is a correctness claim, treat it as one: - Gate plugins with **`validatePlugins`** in CI so missing annotations/path sensitivity fail the build. - Adopt **reproducible builds** defaults project-wide (the archive settings above). - When enabling a **remote** cache, start with high-value deterministic tasks (compile, test, codegen) and require evidence of reproducibility before broadening. - Watch cache **hit rates** and investigate divergence with build scans — a sudden drop usually means a new non-reproducible input crept in. ## Why principal-level The trap is treating caching as a free speed knob. A non-reproducible task marked cacheable produces *correctness* incidents (stale outputs reused across CI) that are painful to debug. The senior/principal judgment is knowing that the annotation is a guarantee you must engineer for, and operationalizing that guarantee across many teams.
- Gradle can validate input annotations — why can't it just verify outputs are reproducible too?Reproducibility depends on arbitrary task logic (clocks, ordering, network). Gradle can't statically prove determinism, so it makes the author assert it via @CacheableTask and relies on conventions/reviews to enforce it.
- How would you detect that a non-reproducible input slipped into a previously-good cacheable task?Monitor remote cache hit rate; a drop signals divergence. Compare two build scans' task cache keys/input fingerprints to find the newly-varying input, or run the task twice and diff outputs byte-for-byte.
- What's a safe rollout order when introducing a remote cache org-wide?Start with expensive, deterministic tasks (compile, test, codegen) that you've verified reproducible; gate with validatePlugins; watch hit rates; then expand. Avoid caching cheap I/O tasks where overhead outweighs gains.
Caching is like reusing a photocopy: only safe if the original is always printed identically. If each print silently adds today's date, sharing copies spreads wrong versions.
saying these in an interview costs you the question
- Believing correctly-declared inputs alone make a task safe to cache (ignores output determinism).
- Marking archive tasks cacheable without reproducible-archive settings.
- Treating @CacheableTask as a pure performance flag rather than a correctness guarantee.