skip to content

Which inputs does a builder compare to decide whether a build step can reuse the layer from the previous run?

level: middleimportance: should knowfreq 48%

answer

  1. inputs are computed before running
  2. parent, definition, copied file content
  3. outputs are never compared
  4. argument values are part of the definition
  5. a hit promises inputs, not freshness

basics

~20 s

Three: the identity of the layer the step starts from, the step's own definition including any argument values in it, and, for a step that copies files in, the content of those files. What the step produced is not an input.

solid answer

~40 s

The builder computes a step's inputs *before* running it and looks for a stored layer recorded against exactly those inputs. The inputs are the **parent layer's identity**, the **step's definition** with any argument or environment values substituted in, and - only for steps that bring files into the image - the **content of the copied files**. Builders differ on that last point: some compare content digests, some file metadata, which is why an otherwise identical checkout can miss on one setup and hit on another. What is deliberately absent is everything on the output side: the layer's size, how long the step took, and above all whatever the step's command fetched from the network. The builder is matching inputs, not judging results.

code

pseudocode · 14 lines
pseudocode
parentLayerId = identityOf(baseImage)

for each step in build:
    inputs = {
        parent:   parentLayerId,
        text:     definitionOf(step) with argument values substituted,
        files:    contentOf(files the step copies in)   # copy steps only
    }
    if storedLayerFor(inputs) exists:
        layer = storedLayerFor(inputs)          # reuse: step is not executed
    else:
        layer = execute(step)                   # miss: real work happens
        store(inputs -> layer)
    parentLayerId = identityOf(layer)           # new on a miss -> next lookup misses

go deeper

for a junior

Know that the builder decides before it runs a step, and that a step's starting layer counts as one of the things it compares.

for a middle

Name all three inputs and say what is excluded - outputs, timing and anything fetched from outside the build.

for a senior

Use the three inputs as a diagnostic order when someone reports a miss they cannot explain, and start at the first step that missed rather than the one they noticed.

for a principal

Recognise that input-matching is the only check that is both cheap and correct, and that making intent visible as an input is the engineer's job, not the builder's.

## The decision is made before the step runs It is tempting to picture a builder executing a step, noticing the result is the same as last time, and discarding the work. That is not what happens and could not be - the point of reuse is to *avoid* the work. Instead the builder computes a description of the step's inputs, looks that description up in its record of previous runs, and either takes the stored layer or executes. That single fact explains almost every surprising cache behaviour, in both directions: a step that should obviously be re-run is not, and a step where nothing seemed to change is. ## The inputs - **The parent layer's identity.** Every step is defined as *this change, applied to that starting point*. A stored layer is only meaningful on the exact parent it was computed against, so the parent is always part of the match. This is why one miss cascades to the end of the build. - **The step's own definition.** The command or declaration as written, with any argument values or environment values substituted in. Change the text, change the inputs. Bind a per-build value into a step and that step can never be reused. - **The content of files the step copies in.** Only steps that bring files from outside contribute this. Here designs genuinely differ - some builders compare a digest of the file content, others rely on file metadata such as size and modification time - which is why the same project can hit on one machine's builder and miss on another's without anything meaningful having changed. ## What is deliberately not an input | candidate | is it compared? | why | |---|---|---| | the layer the step will produce | no | it does not exist yet; that is the point | | how long the step took last time | no | timing is not part of the step's meaning | | the size of the resulting layer | no | an output, available only after executing | | what the step's command downloaded | no | outside the build, invisible to the matcher | | the age of the stored layer | no | reuse does not expire on a clock | The last two rows are where real production surprises come from. A step whose command reaches out to the network is matched on its *text*, so it is reused no matter what the network would return today. And a stored layer does not go stale with time; only a changed input retires it. ## Reading a surprising miss When a step re-runs and an engineer insists nothing changed, work the three inputs in order: 1. **Did something before it change?** A new parent explains a miss on its own, even when the step is untouched. Check the first step that missed, not the one you noticed. 2. **Did the step's definition change?** Including anything interpolated into it - an argument value, an environment value, a version string someone bumped. 3. **Did the copied file set change?** For a copy step, any file it brings in counts, including ones nobody thinks of as source. ## Reading a surprising hit The inverse question is the more dangerous one: a step that should have re-run and did not. Since the match is over inputs, a step whose *intent* is time-dependent - fetch the newest of something - has no time-dependent input, so it is reused indefinitely. The general rule is worth memorising: **a reused layer is a promise about the step's inputs, never about its results**. ## Why this design rather than a smarter one Comparing outputs would require running the step, which defeats reuse. Verifying freshness would require the builder to understand every command a step could run. Matching declared inputs is the only check that is both cheap and correct: if the inputs are identical, the step is deterministic, and the stored layer is right. Where a step is *not* deterministic in its inputs, the model does exactly what it promised and the engineer is the one who has to make the intent visible - by turning the thing that should trigger a re-run into an actual input.

  • Why is the parent layer part of the match at all?
    Because a layer is a set of changes applied to a specific starting point, not a standalone artifact. Reusing a layer computed on a different parent would produce an image that no build ever actually produced. Including the parent is what keeps reuse correct.
  • Two engineers build the same commit and one of them misses on a copy step. What is the likely cause?
    Something in that step's inputs differs even though the tracked content does not - most often file metadata on a fresh checkout, since builders differ in whether they compare content digests or metadata. The second candidate is a parent that differs, such as a base image that was refetched.

saying these in an interview costs you the question

  • Says the builder compares the step's output to decide reuse
  • Thinks a stored layer expires after some period of time
  • Forgets that argument values are part of a step's definition
  • Believes the builder checks whether a fetched result is current
  • Ignores the parent layer and matches on the step text alone