skip to content

A team reports that their image build 'never hits the cache' even though they changed almost nothing between runs. How would you diagnose which instruction invalidates first, and what are the usual causes?

level: seniorimportance: should knowfreq 48%

answer

  1. Find the FIRST non-CACHED step; the rest is cascade
  2. --progress=plain shows CACHED markers
  3. .dockerignore: .git, node_modules, dist, .env
  4. Volatile ARG/LABEL near the top = full rebuild
  5. No cache in CI = wrong scope, mode=min, eviction, fresh builder

basics

~20 s

Build with plain progress output and find the first step not marked CACHED — everything below it is collateral. Then ask what feeds that step's key: a volatile build argument, a wide COPY with no .dockerignore, a moved base-image tag, a different builder or scope per job, or cache eviction.

solid answer

~50 s

**Localise first.** Run `docker buildx build --progress=plain` and read down for the first step that is not `CACHED`. Only that instruction is the real failure; every miss below it is the cascade. **Then interrogate that step's inputs:** - *A `COPY` misses* → the copied content changed. Check `.dockerignore`: `.git`, `node_modules`, `dist/`, local env files and editor state routinely enter the context. Check for generated files (version stamps, compiled assets) written before the build. Check `--chown`/`--chmod` churn. - *A `RUN` misses* → its expanded string changed: a referenced `ARG` carrying a commit SHA or timestamp, a changed `ENV`, a different `WORKDIR`/`USER`. - *The first steps miss* → the base image moved (a mutable tag plus `--pull`), or the build targets a different platform than the cached one. - *Everything misses in CI only* → the cache is not there at all: fresh builder per job, wrong import reference, `mode=min` for multi-stage, eviction or a prune policy, or a leftover `--no-cache`. Fix the first miss, re-measure, repeat.

code

bash · 7 lines
bash
docker buildx build --progress=plain -t app:dev . 2>&1 | grep -nE 'CACHED|^#[0-9]+ ' | head -40

# what is the builder actually holding?
docker buildx du --verbose | head -20

# cold baseline for comparison
time docker buildx build --no-cache -t app:cold .

go deeper

for a junior

Know to look for the first step that is not marked CACHED and to check .dockerignore and instruction ordering.

for a middle

Map the miss to the specific input that feeds that step's key — copied content and mode for COPY, expanded string and environment for RUN.

for a senior

Work the environment layer too: import/export scope, mode=max, platform, eviction, stray flags — and verify with a controlled two-build experiment before and after each fix.

for a principal

Make it systemic: cache hit rate and build-duration metrics as tracked signals, a repo template with .dockerignore and ordering baked in, and pinned base digests so cold builds are scheduled events rather than surprises.

## Step 1: find the first miss, ignore the rest Cache invalidation cascades: once one instruction misses, everything after it must rebuild. So a build log full of misses usually contains exactly **one** interesting event — the first one. Debugging any later step is wasted effort. `docker buildx build --progress=plain .` prints each step with `CACHED` where the cache was used. Read top-down and stop at the first step without it. Compare two consecutive builds' logs; the boundary moves as you fix causes. Useful supporting probes: `docker buildx du` shows what the builder is actually holding; `docker history <image>` shows the layer chain and sizes; a `--no-cache` run gives you a cold-build baseline so you know what the cache is worth. If the whole build is uncached in CI but cached locally, the problem is the environment, not the Dockerfile. ## Step 2: interrogate the inputs of that instruction ### If it is a COPY or ADD The key is a content checksum over the matched files (content, path, mode/ownership under BuildKit). Something in that set genuinely changed. Common real causes: - **No or weak `.dockerignore`.** `COPY . .` with `.git` included means any fetch, any branch switch, any new commit object changes the context. `node_modules`, `target/`, `build/`, `.venv`, `coverage/` and stray OS metadata files do the same. Excluding them is usually the single biggest fix — it also shrinks context upload time. - **Generated files inside the context.** A pre-build step that writes a version file, a timestamp, a `.env`, or compiled assets guarantees a new checksum every run. - **Permission churn.** A CI checkout that yields different file modes than a developer's machine changes the metadata portion of the key, so cache exported by one environment never matches the other. - **Line-ending or formatting rewrites** (CRLF conversion, an autoformatter running before the build) change bytes. ### If it is a RUN The key is the expanded command string plus the environment. Look for: - **A volatile referenced `ARG`.** `ARG GIT_SHA` or `ARG BUILD_TIME` consumed near the top of the file changes every build and takes the entire stage with it. Move it below the expensive steps. - **A changed `ENV`, `WORKDIR`, `USER` or `SHELL`** above the step. - **Deliberate cache-busting left behind** — someone's temporary `--build-arg CACHEBUST=$(date +%s)` that stayed in the pipeline. ### If the miss is at the very top - **The base image moved.** `FROM node:22` is a mutable tag. With `--pull`, or on a runner that resolves it fresh, you can be building on a different digest than the cached layers assumed. Pinning by digest makes this deterministic and turns 'mystery cold build' into an explicit, reviewed bump. - **Platform mismatch.** Cache is per platform. A cache produced for `linux/amd64` gives nothing to a `linux/arm64` build, and multi-platform builds need cache for each. ## Step 3: environment causes — when nothing in the Dockerfile is wrong If the file is fine and it caches locally but never in CI: - **A fresh builder per job** with no `--cache-from`, so there is simply no cache to hit. - **Import/export reference mismatch** — the job writes `app:cache-$BRANCH` and reads `app:cache-main`, or reads a tag no job ever wrote. - **`mode=min` on a multi-stage build**, so builder-stage layers were never exported. - **Eviction / GC.** Provider caches evict under size pressure; a shared builder with a keep-storage limit or a scheduled prune drops old entries. Long gaps between builds then look like random cold builds. - **Parallel jobs racing** to export to the same reference, each overwriting the other's entry. - **A stray `--no-cache`** in the pipeline definition, often added during an incident and never removed. ## Step 4: verify with a controlled experiment Build twice back to back with no changes. If the second build is fully cached, the Dockerfile is fine and the problem is per-run input or environment. If the second build still misses, something inside the build is non-deterministic in a way that feeds a key — usually a generated file or an argument. ## Step 5: fix in order of leverage Add or tighten `.dockerignore`; move volatile arguments and labels to the bottom; split wide copies into narrow ones (manifest first); pin base images by digest; export cache with `mode=max` and a scope that actually matches on import; give a persistent builder a sane GC policy. Then re-measure hit rate and wall-clock time — the point is faster builds, not a prettier log.

  • The build caches perfectly on a developer laptop but never in CI. Where do you look?
    At the environment, not the Dockerfile. Check that the job imports cache at all and that the import reference matches what a previous job exported, that a multi-stage build exports with `mode=max`, and that the target platform matches the cached one. Then check eviction and any `--no-cache` or volatile build argument injected only by the pipeline.
  • How would you prove that a `.dockerignore` change actually helped?
    Take a cold-build baseline and a warm-build time before the change, apply it, then run two consecutive no-op builds and compare where the first non-CACHED step lands and the total wall time. Tracking cache hit rate and p50/p95 build duration in CI over a week turns it from an anecdote into evidence.

saying these in an interview costs you the question

  • Debugging the last failing step instead of the first non-CACHED one
  • Reaching for `--no-cache` as the fix, which hides the problem and makes every build worst-case
  • Assuming a source file's timestamp busted a COPY under BuildKit (content and mode do, mtime does not)
  • Ignoring that mutable base-image tags and platform differences produce legitimate cold builds
  • Never checking whether CI imports cache at all before blaming the Dockerfile

context