skip to content

Why would you register an artifact transform instead of using a normal task, and what caching benefit does triggering it via attribute matching give?

level: seniorimportance: should knowfreq 26%

answer

  1. per-artifact, on demand vs whole-set task
  2. wired into variant matching + chaining
  3. cache key = input content + parameters
  4. dedup across consumers/projects/builds
  5. transform = stateless per-artifact; task = aggregation

basics

~20 s

A transform runs per-artifact, on demand, only when its to attributes are requested, and Gradle caches each result keyed by input + parameters. A task would run once for the whole set and isn't tied into dependency-graph variant matching.

solid answer

~50 s

Use a transform when you need to convert *dependency artifacts* on the fly as part of resolution — e.g. unzip every jar to classes, or minify each jar. Because you trigger it by requesting the `to` attributes (via `artifactView`), Gradle slots the conversion into variant matching: it runs **per artifact**, **lazily**, and **only** for artifacts that are actually consumed in that shape. Each transformed result is cached keyed by the input artifact's content plus the transform's parameters, so the same jar is unzipped once and reused across builds and projects. A normal task isn't aware of variant matching — you'd have to manually enumerate inputs, and it processes everything eagerly as one unit rather than per-artifact on demand. Transforms also de-duplicate work across consumers that request the same shape. The trade-off: transforms are best for stateless per-artifact conversions; cross-artifact aggregation still belongs in a task.

code

kotlin · 15 lines
kotlin
// Parameters fingerprinted into the cache key via normalization annotations
interface MinifyParams : TransformParameters {
    @get:Input
    val keepClasses: MapProperty<String, Set<String>>
}

dependencies {
    registerTransform(Minify::class) {
        from.attribute(minified, false)
        to.attribute(minified, true)
        parameters {
            keepClasses.put("guava", setOf("com.google.common.collect.ImmutableMap"))
        }
    }
}

go deeper

for a junior

Know a transform converts dependency artifacts on demand and results are cached.

for a middle

Explain per-artifact, lazy execution triggered by attribute requests, and basic input+parameter caching.

for a senior

Contrast with tasks, explain de-duplication across consumers and the input-content + parameters cache key, and when aggregation still needs a task.

for a principal

Weigh transforms vs tasks for build performance at scale, and govern parameter normalization so cross-machine build-cache hits are reliable.

## The conceptual difference A **task** produces outputs from declared inputs as a single unit of work in the task graph. An **artifact transform** is a per-artifact conversion rule wired into **dependency resolution**: it activates when a resolution requests the `to` attributes and Gradle inserts it to bridge variant attributes. ## Why a transform, not a task - **Per-artifact, on demand.** Requesting `artifactType=classes` via `artifactView` makes Gradle apply the transform to each jar that is actually consumed in that view — not to artifacts nobody asks for. - **Integrated with matching.** The conversion participates in variant selection and can be chained automatically. A task sits outside this; you'd hand-roll the wiring and lose chaining. - **De-duplication across consumers.** If `:app` and `:tests` both request `classes`, the same dependency jar is transformed once and the result shared. ## Caching keyed by input + parameters This is the headline benefit. Gradle caches a transform's output keyed by: 1. The **content** of the input artifact (its fingerprint), and 2. The transform's **parameters** (the values bound at registration, fingerprinted via normalization annotations on the Parameters interface). So a given (input, parameters) pair is computed **once** and reused — across tasks, across projects, and (with the build cache) across builds and machines. The same upstream jar unzipped in ten modules costs one transform execution. This is why the registration/trigger model matters: because the trigger is *requesting an attribute*, Gradle owns the lifecycle and can cache aggressively, which a manual task pipeline cannot match without significant effort. ## Practical guidance - Reach for a transform when the operation is a **stateless, per-artifact** conversion of dependency files (unzip, minify, instrument, extract metadata). - Keep parameters minimal and properly annotated (e.g. `@Input`, `@InputFiles` + `@PathSensitive`) so the cache key is correct. - Don't use a transform for **cross-artifact aggregation** (e.g. merging all artifacts into one report) — that's a task, because a transform sees one input artifact at a time. ## Net effect of the register+trigger design Registration declares the rule; the `artifactView` request triggers it; attribute matching decides where it applies; and the input+parameters cache key makes it run the minimum number of times. Together these give correct, deduplicated, cacheable artifact processing that a plain task can't easily replicate.

  • What forms the cache key for a transform's output?
    The fingerprint of the input artifact's content plus the transform's parameters (fingerprinted via normalization annotations on the Parameters interface). The same pair is computed once and reused.
  • When should you still use a task instead of a transform?
    For cross-artifact aggregation — merging or summarizing many artifacts into one output — because a transform processes one input artifact at a time and can't see the whole set.
  • Does the transform run for every dependency in the configuration?
    No. Only for artifacts actually consumed through the requesting view in the target shape; unrequested artifacts are never transformed.

A task is a batch job you schedule; a transform is a vending machine — it produces an item only when someone presses the button for that shape, and it remembers (caches) the last item it made for that exact request.

saying these in an interview costs you the question

  • Saying transforms run eagerly for all dependencies — they run per consumed artifact, on demand.
  • Using a transform for aggregation across artifacts (it sees one input at a time).
  • Forgetting that parameter normalization annotations affect the cache key's correctness.

context