skip to content

Layers & Build Cache

Each instruction adds one read-only layer, every layer is named by its sha256 digest, and the build cache is keyed per instruction so one changed line invalidates all the rest. Ordering a Dockerfile around that is the biggest build-speed win available.

part ofDockeroverview, primer and where to startread it →
on this pageshow

questions

22

In a Dockerfile, why do teams copy the dependency manifest (for example package.json or pom.xml) and run the dependency install before copying the rest of the application source? What goes wrong if the whole source tree is copied first?

level: juniorimportance: must knowfreq 80%

answer

  1. Layers cached in order — one miss, all misses below
  2. COPY key = content checksum
  3. Manifest + lockfile first, install, then COPY . .
  4. Least-changing at top, most-changing at bottom
  5. .dockerignore keeps the context (and checksum) stable

basics

~20 s

Each build instruction becomes a cached layer, and once one instruction misses the cache every later one is rebuilt. Copying source first means any code edit invalidates the dependency install. Copy the manifest, install, then copy source, so installs stay cached until dependencies change.

solid answer

~50 s

An image is a stack of layers, one per build instruction. The builder reuses a cached layer only if that instruction's inputs are unchanged **and** every preceding instruction also hit the cache — invalidation cascades forward and never recovers. A `COPY` is keyed on a checksum of the copied content, so `COPY . .` misses on any source edit. If the dependency install (`npm ci`, `mvn dependency:go-offline`, `pip install -r requirements.txt`) sits after that copy, it re-runs on every commit: the slowest, most network-heavy step in the build, invalidated by a one-character change. The fix is to order instructions from least- to most-frequently changing: copy only the manifest and lockfile, run the install, then copy the source and build. The install layer is now reused until the lockfile itself changes. Pair it with a `.dockerignore` so `.git`, `node_modules` and build output never enter the build context and never churn the checksum.

code

dockerfile · 6 lines
dockerfile
FROM node:22-alpine
WORKDIR /app
COPY . .
RUN npm ci
RUN npm run build
CMD ["node", "dist/server.js"]

go deeper

for a junior

Recall the rule: instructions become layers, a miss cascades downward, so copy the manifest and install dependencies before copying source. Be able to write the corrected Dockerfile.

for a middle

Explain why COPY misses — a content checksum — and connect it to .dockerignore and lockfile determinism. Mention that the same ordering applies inside each stage of a multi-stage build.

for a senior

Frame it as ordering by rate of change, quantify the saving (network-bound install per commit versus per lockfile change), and note the correctness angle: deterministic installs so a cached layer matches the lockfile you tested.

for a principal

Treat it as a policy question — a repo template or lint rule so every service gets the layout by default, plus measurement of cache hit rate and build time to show the ordering actually pays off across the fleet.

## What the build cache is A container image is a stack of read-only filesystem layers. Most Dockerfile instructions produce one layer (`RUN`, `COPY`, `ADD`); others only change image configuration (`ENV`, `WORKDIR`, `LABEL`, `EXPOSE`). When you build, the builder walks the instructions top to bottom and, for each one, computes a cache key from the instruction and the state it is applied to. If a layer built earlier from the same key is available, the builder reuses it and prints `CACHED` instead of executing anything. ## The cascade rule The cache key of an instruction includes the identity of the layer it is applied to. That single fact drives everything else: if instruction #4 misses, then instruction #5 is being applied to a *different* parent, so its key differs too, so it must miss as well — and so on to the end of the stage. There is no such thing as a cache hit after a miss. The practical consequence is that the position of an instruction matters as much as its content: an instruction that changes often poisons every step below it. ## Why `COPY . .` is the expensive mistake For `COPY` and `ADD` from the build context, the cache key is derived from the *contents* of the files being copied (BuildKit hashes file content plus path and mode/ownership metadata), not from a timestamp. So `COPY . .` produces a new key whenever any file in the context changes — which is every commit, by definition. Consider the naive ordering: ``` FROM node:22 WORKDIR /app COPY . . RUN npm ci RUN npm run build ``` Changing one line in a component invalidates `COPY . .`, which invalidates `npm ci`, which re-downloads the entire dependency tree from the registry. A build that could take 8 seconds takes 3 minutes, on every push, for every developer and every CI job. ## Ordering by change frequency The rule of thumb: **put the things that rarely change high in the file and the things that change every commit low**. Dependencies change on the cadence of a lockfile edit — perhaps weekly. Source changes hourly. So split the copy: ``` FROM node:22 WORKDIR /app COPY package.json package-lock.json ./ RUN npm ci COPY . . RUN npm run build ``` Now `npm ci` sits above the volatile copy. Editing source invalidates only `COPY . .` and the build step; the install layer is reused. Editing `package-lock.json` correctly invalidates the install — which is exactly what you want, because the dependency set genuinely changed. The cache is not being cheated; it is being told the truth about what each step depends on. The same shape applies in every ecosystem: `pom.xml` before `src/`, `requirements.txt` before the app package, `go.mod`/`go.sum` before the Go sources, `Cargo.toml` before `src/`. ## Supporting habits **`.dockerignore`.** The build context is everything sent to the builder. If `.git`, `node_modules`, `target/`, `dist/` and local `.env` files are included, the `COPY . .` checksum changes for reasons unrelated to your code — a fresh `git fetch` alone can do it. A `.dockerignore` keeps the context small (faster to transfer) and the checksum stable. **Lockfiles.** Cache correctness depends on the manifest fully describing the dependency set. `npm ci` against a committed lockfile is deterministic; `npm install` without one can resolve differently on a cache hit versus a cold build, producing images that differ from what you tested. **Multi-stage builds.** The same ordering discipline applies inside each stage. A builder stage compiles with the deps-first layout; the final stage copies only the produced artifact, so the runtime image is small and its layers change only when the artifact does. ## How to talk about it The crisp framing for an interview is: *layers are cached in order, a miss cascades, and `COPY` misses on any content change — so put slow, stable work above fast, volatile work.* That one sentence explains the manifest-first pattern, why `.dockerignore` matters, and why installing build tools belongs near the top of the file.

  • You reordered the Dockerfile but the install still re-runs on every build. What would you check first?
    Check the build context and `.dockerignore`. If `.git`, `node_modules` or generated output are being copied, or if the manifest glob accidentally pulls in volatile files, the copy above the install keeps changing. Also confirm the copy above the install lists only the manifest and lockfile — a stray `COPY . ./` or `COPY src ./` before the install defeats the ordering.
  • Does this ordering change anything about the final image's size or contents?
    No. The layer set and total content are the same; only build time differs, because layers are reused instead of recomputed. Splitting the copy adds one extra tiny layer for the manifest files, which is negligible. Size is governed by what you install and which stage you ship, not by cache ordering.

Like prepping a kitchen: you stock the pantry once and reuse it all week. Putting the pantry restock after 'chop today's vegetables' means re-stocking for every single meal.

saying these in an interview costs you the question

  • Believing Docker re-checks whether a RUN would produce the same output and skips it — it does not inspect results, only cache keys
  • Thinking a later instruction can hit the cache after an earlier one missed
  • Claiming `.dockerignore` only affects image size, not cache behavior
  • Using `npm install` instead of `npm ci` and assuming the cached layer still matches the committed lockfile

context

open as a page

In a container image reference, what is the difference between the tag form nginx:1.27 and the digest form nginx@sha256:1a2b3c..., and what guarantee does each one give you?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A tag is a mutable label the registry can repoint at new content at any time. A digest is the sha256 hash of the image manifest, so one digest always resolves to byte-identical content. Tags name an image; digests identify it.

open as a page

A running container writes a 2 GB file into a path inside itself. Where does that data physically live on the host, what happens to it when the container is removed, and why does the image on disk stay unchanged?

level: juniorimportance: must knowfreq 55%

basics

~20 s

It lands in the container's own thin writable layer on the host, on top of the read-only image layers. Removing the container deletes that layer and the data. Image layers are never written to, which is why many containers can share one image.

open as a page

In a Dockerfile, someone installs packages in one RUN instruction and deletes the package cache in a separate, later RUN instruction, but the built image is no smaller. Explain why, and how you would fix it.

level: juniorimportance: must knowfreq 70%

basics

~20 s

Every Dockerfile instruction commits its own read-only layer. A later layer can only mark files as deleted; the bytes still sit in the earlier layer and still ship. Delete in the same RUN that created the files.

open as a page

How does the container image builder decide whether an individual Dockerfile instruction can reuse a cached layer? Contrast the rule used for RUN with the rule used for COPY and ADD.

level: middleimportance: must knowfreq 72%

basics

~20 s

For RUN the key is the literal command string plus the parent layer and environment — the builder never inspects what the command would produce. For COPY and ADD from the build context the key is a checksum of the copied files' content and metadata. Any mismatch rebuilds that step and everything below it.

open as a page

How can one reference such as alpine:3.20 work on both linux/amd64 and linux/arm64 machines — what does the registry return, and how is the correct variant selected?

level: middleimportance: must knowfreq 50%

basics

~20 s

The tag points at a manifest list (OCI image index): a JSON document listing one manifest digest per platform, each annotated with os/architecture/variant. The client picks the entry matching its own platform and pulls only that image's config and layers.

open as a page

Docker's default storage driver on Linux is overlay2, which builds a container's root filesystem with an overlay mount. Explain what the lowerdir, upperdir, workdir and merged directories are and how they combine.

level: middleimportance: must knowfreq 48%

basics

~20 s

lowerdir is the stack of read-only image layers, upperdir is the container's writable layer, merged is the unified view the container uses as /, and workdir is a private scratch directory the kernel needs for atomic operations. Upper entries shadow lower ones.

open as a page

Inside a running container you delete a directory such as /usr/share/doc that came from a read-only image layer. How does the union filesystem record that deletion, and why does the deleted data still exist on the host?

level: middleimportance: must knowfreq 44%

basics

~20 s

Lower layers cannot be modified, so the deletion is recorded as a whiteout in the writable layer that hides the name. The original bytes stay in the read-only image layer, still on disk and still extractable, just invisible through the merged view.

open as a page

Why does rebuilding a container image from the same Git commit produce a different digest?

level: middleimportance: must knowfreq 52%

basics

~20 s

An image records file bytes and metadata, not your commit. A fresh checkout gives every file a new mtime, the image config gets a new created timestamp, unpinned installs pull newer packages, and changed build args alter the config.

open as a page

CI builds run on fresh, ephemeral runners, so every image build starts with an empty local build cache. How do you get cache hits anyway, and what are the tradeoffs of the approaches?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Export the cache to a shared backend and import it next run: BuildKit's --cache-to/--cache-from with a registry, a CI-provider cache, or a local directory. Use mode=max to keep intermediate stage layers. Alternatively keep a persistent remote builder so its cache never disappears.

open as a page

You need a container image build to genuinely re-fetch things rather than reuse cached layers — for example to pick up new upstream OS package versions. What options does the Docker CLI give you, and how do they differ from each other?

level: middleimportance: should knowfreq 55%

basics

~20 s

--no-cache ignores all cached layers for the build; --pull refreshes the base image referenced by FROM; --no-cache-filter <stage> busts only chosen stages. Changing a referenced build argument invalidates from that instruction down. Each has a different blast radius and cost.

open as a page

`docker images` shows an IMAGE ID starting with sha256, and `docker inspect` also reports a RepoDigests entry like repo@sha256:.... Why are those two sha256 values different, and which one identifies what you pulled from the registry?

level: middleimportance: should knowfreq 44%

basics

~20 s

The image ID is the sha256 of the image's local config JSON blob. The repo digest is the sha256 of the manifest the registry served. The manifest points at the config plus layers, so the two hash different documents. Registry references use the repo digest.

open as a page

Your Dockerfile installs OS packages unpinned and curls a tarball. What does that cost you?

level: middleimportance: should knowfreq 44%

basics

~20 s

Every fetch resolves at build time against something you do not control, so two builds of the same commit produce different filesystems. The image you promote is not the image you tested, and the build breaks when upstream moves.

open as a page

Walk through the package-manager flags and cleanup steps you use to keep Debian/Ubuntu and Alpine container images small during a build, and explain what each one saves.

level: middleimportance: should knowfreq 55%

basics

~20 s

On Debian/Ubuntu: apt-get install with --no-install-recommends, then rm -rf /var/lib/apt/lists/* in the same RUN. On Alpine: apk add --no-cache, and apk del a --virtual group for build-only tools. For language tools use pip --no-cache-dir and npm cache clean --force.

open as a page

A team reports that their image build 'never hits the cache' even though they changed almost nothing between runs. How would you diagnose which instruction invalidates first, and what are the usual causes?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Build with plain progress output and find the first step not marked CACHED — everything below it is collateral. Then ask what feeds that step's key: a volatile build argument, a wide COPY with no .dockerignore, a moved base-image tag, a different builder or scope per job, or cache eviction.

open as a page

Walk through what a registry stores for a single-platform container image: which fields in the manifest identify the layers, and why does the config blob's rootfs.diff_ids list different sha256 values than the manifest's layer descriptors?

level: seniorimportance: should knowfreq 34%

basics

~20 s

The manifest lists a config descriptor and layer descriptors — digests of the compressed blobs as transferred. The config blob lists rootfs.diff_ids — digests of the same layers uncompressed. Different bytes hashed, so different values; diff_ids identify layers on disk, layer digests identify them on the wire.

open as a page

A process inside a container opens a 4 GB file that came from a read-only image layer and writes one byte near the start. What does the union filesystem do, and what does that imply for write-heavy workloads?

level: seniorimportance: should knowfreq 40%

basics

~20 s

overlayfs copies the whole file from the image layer into the container's writable layer before applying the write — a copy-up. One byte costs 4 GB of I/O and 4 GB of disk. Write-heavy or large-file workloads belong on volumes, which skip this.

open as a page

A team's service image has grown to 1.8 GB and nobody knows why. Describe how you would find which build steps and which files account for the size, and what you would do with the findings.

level: seniorimportance: should knowfreq 50%

basics

~20 s

Use docker history to attribute bytes to Dockerfile instructions, then dive to see per-layer files and bytes wasted by files deleted or overwritten later. Fix the guilty steps: .dockerignore, multi-stage build, smaller base, same-layer cleanup. Re-measure after.

open as a page

Your deployment manifests reference every container image by tag. A colleague proposes pinning all of them to sha256 digests instead. How would you decide, and what machinery does digest pinning require to be sustainable?

level: principalimportance: should knowfreq 38%

basics

~20 s

Pinning buys reproducible, auditable, exactly rollback-able deployments and closes tag-mutation risk. It costs you automatic patching, so it only works with automated digest bumps, registry retention that never garbage-collects pinned manifests, and a clear rule about which references may still be tags.

open as a page

How far should a team chase bit-identical container image rebuilds versus an auditable rebuild?

level: principalimportance: should knowfreq 28%

basics

~20 s

Bit-identity is a verification property, worth buying only where an outside party must check you. For internal services the goal is knowing what went into an image and being able to rebuild an equivalent one. Take the cheap wins everywhere.

open as a page

Your organisation is standardising container base images and someone proposes moving every service to Alpine or to a distroless image. How would you evaluate that, and what would you actually recommend?

level: principalimportance: should knowfreq 48%

basics

~20 s

Smaller bases cut bytes and vulnerability surface but cost compatibility and debuggability: Alpine uses musl instead of glibc, distroless has no shell or package manager. Decide per runtime, keep a debug path, and prefer slim glibc images where native dependencies matter.

open as a page

What does setting SOURCE_DATE_EPOCH change about a BuildKit container image build?

level: seniorimportance: nice to knowfreq 22%

basics

~10 s

SOURCE_DATE_EPOCH is a Unix timestamp BuildKit uses instead of the wall clock: it writes that value into the image config's created and history fields and normalises file timestamps in the layers the build produces.

open as a page