skip to content

A dependency-fetch step reused its cached layer for weeks and shipped stale package versions - why did the builder not re-run it?

level: seniorimportance: should knowfreq 40%

answer

  1. a hit promises inputs, not freshness
  2. the step's text never moved
  3. latest is a time-dependent intent
  4. no clock retires a stored layer
  5. make the intent a copied-in input

basics

~20 s

Because nothing the builder compares had changed: the step's definition was fixed, the layer beneath it was the same, and it copied no files. Reuse is a promise about a step's inputs, never about the freshness of what it fetched.

solid answer

~50 s

A step that says *refresh the package manager's index and install the latest updates* has a definition that never changes, sits on a parent that rarely changes, and copies nothing in. So its inputs are constant, it matches on every build, and the layer that gets reused is whatever the outside world looked like the day it last executed. The builder is not wrong - it never promised freshness, only that identical inputs give an identical layer. The failure is that the step's **intent** is time-dependent while its **inputs** are not. The fixes all work by making the intent an input: pin the versions you want in a file the step copies in, so editing that file invalidates the step; or deliberately rebuild without reuse on a cadence you choose, so the step executes whether or not anything matched.

go deeper

for a junior

Remember the one-liner: a reused step is matched on its inputs, so a step that fetches the newest of something is not re-run just because the newest changed.

for a middle

Explain which three inputs stayed constant for that step, and why a fixed definition plus a stable parent means a permanent hit.

for a senior

Recognise the class - intent depends on the outside world, inputs do not - and give the fix as making the intent an input, with a bounded drift window as the fallback.

for a principal

Own the trade explicitly: always-current and reproducible-from-unchanged-inputs pull against each other, and a team should decide which failure it prefers rather than discover it.

## Why the step was never re-run The builder decides reuse from a step's inputs: the parent layer's identity, the step's own definition, and the content of any files it copies in. A fetch-the-latest step has all three constant. Its text is fixed - that is the whole idea of *latest*. Its parent is an early, stable layer. It copies nothing. So every build after the first computes the same inputs, finds the stored layer, and skips the work. The layer it takes instead is a frozen photograph: whatever the index said and whatever versions resolved on the day the step last actually ran. Weeks later the image still ships those, and nothing in the build output looks unusual, because from the builder's point of view this is a perfectly ordinary cache hit. ## The general defect This is one instance of a pattern worth naming: > A **reused layer is a promise about the step's inputs, not about its results.** Any step whose correctness depends on something *outside* the build - a network fetch, a moving upstream reference, a clock - has a time-dependent intent and a time-independent set of inputs. The mismatch is invisible because the mechanism is behaving exactly as specified. Other steps with the same shape: - fetching an archive from a mutable location rather than one identified by content; - resolving a dependency range rather than an exact version; - cloning a source reference that moves, such as a branch tip; - copying in a generated file that the build itself regenerates elsewhere. ## Making the intent an input The fixes are all the same move - turn the thing that should trigger a re-run into something the builder compares: 1. **Pin what you actually want, in a file the step copies in.** If the versions live in a file that the step brings into the image, then updating them changes that step's inputs and the step re-runs. The update becomes a reviewable change rather than an accident of timing. 2. **Rebuild without reuse on a cadence you choose.** A periodic full build - reuse disabled - forces every such step to execute, so the drift window is bounded by a number you picked rather than by whenever the inputs happened to move. 3. **Re-point the step at content-identified inputs.** A fetch of something named by its content digest is either the same bytes or a different identity; there is no silent third case. ## The trade-off nobody gets to avoid | approach | what you get | what it costs | |---|---|---| | leave it unpinned and reused | fastest builds, no maintenance | silent drift; the image ships old versions indefinitely | | leave it unpinned, rebuild fully on a cadence | bounded drift window | a slow build every cadence, and versions can change under you without a code change | | pin versions as a copied-in file | changes are explicit and reviewable | someone must do the updating, or the pin rots in the other direction | There is no option where the image is both always current and always reproducible from unchanged inputs - those are opposite requirements, and the choice is a judgment about which failure the team would rather have. ## What this is *not* It is not a network problem, and it is not the stored layer having gone bad. The layer is exactly what the step produced; the record is accurate. It is also not fixed by removing one layer from a built image or by rebuilding on a different machine - a different builder with no record simply executes everything once and then reproduces the same drift from its own start date. ## How to answer this in an interview Name the mechanism first, then the class of defect, then the fix, in that order: the step's inputs never changed, so it was reused; a hit guarantees input equality and says nothing about freshness; make the intent an input by pinning it into a copied file, and bound the residual drift with a deliberate no-reuse rebuild. If asked what you would *not* do, say: relying on the cache to expire by itself - it does not have a clock.

  • Would splitting the refresh and the install into two steps fix it?
    No. Both halves still have fixed definitions, constant parents and no copied files, so both are reused. Splitting changes the granularity of reuse, not whether either step's inputs ever move.
  • Is pinning versions strictly better than fetching the latest ones?
    It is a trade, not an upgrade. Pinning makes builds reproducible and changes reviewable, but a pin that nobody updates silently ages in the other direction. Whichever you choose, the update has to be somebody's job - the builder will not do it for you.
  • The team says the cache must have been corrupt. How do you answer?
    The stored layer is exactly what the step produced and the record of its inputs is accurate; nothing is corrupt. The step simply had no input that moved. Reframe it as a step whose intent depends on the outside world while its inputs do not.

It is like a standing order at a newsagent that is answered from a filed copy rather than re-run: the order still reads bring me today's paper, and you keep receiving the edition from the day it was first filled.

saying these in an interview costs you the question

  • Thinks the builder checks whether a fetched result is still current
  • Believes a fetch-the-latest step keeps the image up to date
  • Says a stored layer expires on its own after a while
  • Treats the drift as a network or registry fault
  • Proposes deleting one layer from the built image instead of rebuilding
  • Assumes building on another machine proves the layer was fine