skip to content

How does the container image builder decide whether an individual Dockerfile instruction can reuse a cached layer? Contrast the rule used for RUN with the rule used for COPY and ADD.

level: middleimportance: must knowfreq 72%

answer

  1. RUN = command-string match, never result inspection
  2. COPY/ADD = content + mode checksum, mtime-insensitive in BuildKit
  3. Parent layer identity is in every key → cascade
  4. Referenced ARG/ENV expands into the RUN key
  5. ADD <url> keys on the URL unless you pin --checksum

basics

~20 s

For RUN the key is the literal command string plus the parent layer and environment — the builder never inspects what the command would produce. For COPY and ADD from the build context the key is a checksum of the copied files' content and metadata. Any mismatch rebuilds that step and everything below it.

solid answer

~60 s

Every instruction gets a cache key combining the parent layer's identity, the instruction text, and for file-copying instructions the content of the inputs. - **RUN**: the key is the *command string as written*, plus the base image and relevant build environment (expanded `ARG`s, `ENV`, `WORKDIR`, `USER`). Docker does not execute the command to see whether the result would differ. `RUN apt-get update` hits the cache forever because the string never changes, even though upstream package indexes move daily. Conversely, a cosmetic edit — reordering flags, changing whitespace — is a miss. - **COPY / ADD from the context**: the key is a checksum over the copied files' content plus path and mode/ownership. Editing a byte misses; touching a file's mtime alone does not, under BuildKit. - **ADD from a URL**: keyed on the URL (and `--checksum` if you pin one), so a changed file behind a stable URL can be silently reused. - Metadata instructions (`ENV`, `WORKDIR`, `LABEL`) key on their text and invalidate everything below when changed. And invalidation cascades: no instruction after a miss can hit.

code

dockerfile · 8 lines
dockerfile
# stale-prone: update hits cache forever
RUN apt-get update
RUN apt-get install -y curl

# correct: one string, one cache key
RUN apt-get update \
 && apt-get install -y --no-install-recommends curl \
 && rm -rf /var/lib/apt/lists/*

go deeper

for a junior

Know the two headline rules — RUN compares the command text, COPY compares file contents — and that a miss cascades downward.

for a middle

Add the environment inputs (base image, referenced ARG/ENV, WORKDIR/USER) and the classic apt-get update staleness consequence, plus why volatile metadata belongs at the bottom.

for a senior

Be precise about what does and does not enter the key (mtime, --chown, mounts, secrets), and use that precision to explain real cache-miss incidents.

for a principal

Talk about cache keys as a contract: build inputs must be declared, deterministic and pinned, so a cache hit is provably the same build — a supply-chain and reproducibility property, not just a speed one.

## The shape of a cache key When the builder processes an instruction it produces a cache key from three things: (1) the identity of the layer or state the instruction is applied to, (2) the instruction itself as written, and (3) for instructions that read the build context, a digest of the data they read. If a previously built result exists for that key — in the local build cache, or in an imported remote cache — the builder reuses it and reports `CACHED`. The first component is what produces the **cascade**: because the parent's identity is part of the key, one miss changes the parent of everything below, so nothing below can hit. Cache is sequential and positional, never per-instruction in isolation. ## RUN: string matching, not result matching For `RUN`, the key is essentially the command string plus the build environment that would affect it. Crucially, **the builder does not run the command to find out whether the output changed**. It compares text. Two consequences follow, and interviewers probe both: **Staleness.** `RUN apt-get update` will hit the cache weeks later even though upstream package indexes have moved on. The command string is identical, so the builder happily reuses a layer containing an index snapshot from a month ago. That is why `update` and `install` must live in the *same* `RUN`: they then form a single string that either hits or misses together. **Over-invalidation.** Any textual change is a miss — adding a flag, reordering packages, even reflowing a line continuation. The command is semantically identical; the cache does not care. What else feeds the key: the base image (a different `FROM` digest changes the state everything is applied to), the current `WORKDIR`/`USER`/`SHELL`, and environment values. An `ARG` that is *referenced* by the command is expanded into the string, so changing that build argument busts the step; an `ARG` that is declared but never used does not affect it. A changed `ENV` invalidates because the resulting configuration and the command environment differ. BuildKit adds `RUN --mount` to the picture. `--mount=type=bind` and `--mount=type=secret` bring inputs whose identity participates in the key (secrets are deliberately not baked into layers). `--mount=type=cache` gives the command a persistent scratch directory — a package-manager download cache — that is *not* part of the layer or the key: it makes a cache **miss** cheaper, it does not create a hit. Confusing those two is a common error. ## COPY and ADD: content checksums For `COPY` (and `ADD` reading local files), the key includes a digest computed over the files actually being copied — their content, their paths, and metadata such as mode and ownership — not their modification times under BuildKit. So a `git checkout` that rewrites timestamps does not by itself bust the cache; changing a byte does. Note also that only the *matched* files count: `COPY src/ ./src/` is unaffected by edits under `docs/`, which is why narrow copies cache better than `COPY . .`. A `--chown` or `--chmod` change alters metadata and therefore the key. And what is in the build context matters: `.dockerignore` excludes paths from the context entirely, so they can neither be copied nor perturb a checksum. `ADD` with a remote URL behaves differently and is worth calling out: the key is based on the URL (plus the `--checksum=sha256:…` you pin, when you pin one), so if the artifact behind that URL is replaced, the build can keep reusing a layer holding the old bytes. Pinning a checksum makes the dependency explicit and the failure loud. ## Metadata instructions `ENV`, `WORKDIR`, `LABEL`, `EXPOSE`, `USER`, `ENTRYPOINT`, `CMD` change image configuration. They key on their own text, and because they change the state below them, editing one invalidates the rest of the stage. That is why volatile metadata — a build timestamp, a commit SHA label — belongs at the *bottom* of the Dockerfile, not the top. Putting `ARG GIT_SHA` and its `LABEL` near the top guarantees a full rebuild every commit. ## Multi-stage nuance Each stage is cached on its own chain of keys. `COPY --from=builder /app/dist ./dist` is keyed on the *content produced by that stage*, so if the builder stage recompiles to byte-identical output, the copy can still hit. This is one reason reproducible compilers help caching downstream. ## Putting it together The interview-ready summary: **RUN matches on text, COPY matches on content, everything matches on its parent — and one miss propagates to the end of the stage.** From that you can derive nearly every practical rule: keep `apt-get update` fused with its install, copy manifests narrowly and early, keep volatile labels last, and never assume the builder is smart enough to notice that a command's output would have been the same.

  • Does `RUN --mount=type=cache` make a step hit the build cache?
    No. It gives the command a persistent directory — for example a package-manager download cache — that survives across builds but is not part of the layer or the cache key. It makes a cache *miss* much cheaper because downloads are reused, but the step still executes. It also lives on the builder, so an ephemeral CI runner loses it unless the builder itself is persistent.
  • Why can putting `ARG GIT_SHA` near the top of a Dockerfile destroy build times?
    Because the argument changes on every commit, and any instruction that references it — or any instruction after the `ARG`/`LABEL` that consumes it — gets a new cache key. The miss then cascades to the end of the stage, so the whole build reruns. Declare and consume volatile build arguments as late as possible, ideally after the expensive install and compile steps.

saying these in an interview costs you the question

  • Saying Docker re-runs a RUN and compares the produced filesystem to decide on the cache
  • Claiming a file's modification time busts a COPY under BuildKit
  • Assuming `ADD https://…/tool.tgz` re-downloads and notices the artifact changed
  • Thinking cache mounts and cache hits are the same mechanism
  • Believing an unused `ARG` invalidates the steps below it

context