How does splitting a Dockerfile into several build stages change build parallelism and layer-cache invalidation compared with one long linear Dockerfile, and how should you order instructions inside each stage to benefit?
answer
- Stages = DAG, not a chain
- Independent stages build concurrently
- Unreferenced stages never execute
- COPY --from keys on artifact content
- Order by volatility: base → deps → source
basics
~20 sStages form a dependency graph, not a list. Independent stages build concurrently, unreferenced stages are skipped, and a cache miss invalidates only the rest of that stage's chain — not sibling stages. Inside each stage, put rarely-changing steps first.
solid answer
~60 sA single-stage Dockerfile is a chain: one cache miss invalidates every instruction after it. Multi-stage builds turn that chain into a **DAG**. `FROM x AS y` and `COPY --from=y` are the edges; BuildKit resolves the whole graph before executing anything. Three practical consequences: - **Parallelism.** Stages with no dependency between them run concurrently — a frontend-asset stage and a backend-compile stage overlap instead of queueing. - **Skipping.** Anything outside the requested target's closure is never executed. - **Isolated invalidation.** Editing the runtime stage does not rebuild the builder; changing a source file busts the compile step but not the dependency-resolution stage that precedes it. Inside each stage the usual rule applies: order instructions from least to most volatile — pinned base, OS packages, dependency manifests + install, then application source. The consuming `COPY --from` also compares content: if the builder reruns but emits a byte-identical artifact, downstream layers can still hit cache. On ephemeral CI runners remember that a registry cache export keeps only final-stage layers unless you export in max mode.
code
dockerfile · 18 linesFROM node:22 AS web
WORKDIR /w
COPY web/package*.json ./
RUN npm ci
COPY web/ .
RUN npm run build
FROM golang:1.22 AS api
WORKDIR /a
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN CGO_ENABLED=0 go build -o /out/app ./cmd/api
FROM gcr.io/distroless/static:nonroot
COPY --from=api /out/app /app
COPY --from=web /w/dist /public
ENTRYPOINT ["/app"]go deeper
Know that each stage caches separately and that dependency manifests should be copied before the source so installs stay cached.
Describe the DAG, the two edge types (FROM and COPY --from), and that unreferenced stages are skipped; explain per-instruction cache keys.
Diagnose real builds: which stage is being invalidated and why, artifact-checksum reuse on the consuming COPY, mode=max cache export on ephemeral runners, per-platform caches.
Frame it as build-time economics — where cache lives across a fleet, reproducible artifacts to preserve downstream reuse, and whether emulated multi-arch builds justify native runners.
## From chain to graph In a single-stage Dockerfile, the layer cache is a prefix match: instruction N reuses cache only if every instruction before it hit cache and its own cache key matches. One edit near the top rebuilds everything below it. Multi-stage builds change the topology. Each `FROM` starts an independent chain, and the only links between chains are `FROM <stage>` (inheritance) and `COPY --from=<stage>` (artifact transfer). BuildKit parses the file into a directed acyclic graph of these dependencies and then schedules it, so **file order stops dictating execution order**; dependency order does. ## Parallelism Two stages that do not depend on each other are built at the same time, subject to available parallelism on the builder. The classic win is a service whose image needs both compiled frontend assets and a compiled backend binary: ``` FROM node:22 AS web # npm ci && npm run build FROM golang:1.22 AS api # go build FROM distroless/static COPY --from=api /out/app /app COPY --from=web /src/dist /public ``` `web` and `api` have no edge between them, so they run concurrently and total wall time is roughly the slower of the two rather than their sum. Writing the same work as one linear stage forces serialisation for no reason. This scheduling is a BuildKit property; the legacy pre-BuildKit builder executed stages strictly in file order. The corollary is that *creating* independence is a design lever: if a stage needs only `go.mod` and `go.sum`, give it only those, so it does not become artificially dependent on a stage that touches the whole source tree. ## Skipping BuildKit builds only the dependency closure of the requested target (the last stage by default, or `--target`). Stages outside it produce no work at all — no pulls, no RUN steps, no time. This is why lint/test/docs stages are free to leave in the file, and simultaneously why they do not gate a production build unless CI targets them explicitly or the production stage copies something they produced. ## Isolated invalidation Cache keys are computed per instruction within a stage: for `RUN`, from the command string plus the parent layer's identity; for `COPY`/`ADD`, from the metadata and content checksum of the copied files plus the parent. Because chains are separate, invalidation is contained: - Changing `ENTRYPOINT` or adding a `LABEL` in the runtime stage rebuilds only that stage's tail — the builder stage is untouched. - Editing application source busts the `COPY . .` and `RUN build` steps of the builder, but a `deps` stage that only copied manifests still hits cache, so dependency resolution is not repeated. - Two stages that both inherit from a shared ancestor share that ancestor's cached layers instead of duplicating the work — provided the shared work really is in the ancestor and not copy-pasted into both. There is a second, subtler effect on the consuming side. `COPY --from=builder /out/app /app` keys on the *content* of `/out/app`. If the builder stage reruns (say a comment changed in the source) but the compiler emits a byte-identical binary, the checksum is unchanged and the downstream layers can still hit cache — the runtime image is not rebuilt or re-pushed. Reproducible builds therefore pay off twice. Conversely, embedding a build timestamp or a git SHA into the artifact guarantees a new checksum on every commit and forfeits that reuse. ## Ordering inside a stage The DAG does not rescue a badly ordered stage. Within each stage, sort by volatility: 1. Pinned `FROM` and OS package installs (change monthly). 2. Dependency manifests only — `go.mod`/`go.sum`, `package*.json`, `pom.xml`, `requirements.txt` — followed by the install command (change weekly). 3. Application source and the build command (change per commit). Copying the whole tree before installing dependencies is the single most common cache-destroying mistake, and it is worse in a multi-stage build because the expensive step being re-run is dependency resolution in the builder. ## Cross-run caching in CI Ephemeral CI runners start with an empty local cache, so the DAG's benefits only materialise if cache is imported from somewhere. Two things bite here. First, exporting an image to a registry stores the final stage's layers only; importing that as cache gives you nothing for the builder stage. Keeping intermediate-stage cache requires an explicit max-mode cache export to a registry or remote cache backend. Second, cache is per-platform: a multi-architecture build maintains separate caches per target platform, and emulated (QEMU) builds are slow enough that cache import matters far more than on native runners. ## What to say in an interview Name the DAG, give the parallel frontend/backend example, state that invalidation is per-chain, and mention that stage skipping is why test stages don't gate. Then land the practical rule: order by volatility inside stages, and export cache in max mode if your runners are ephemeral.
- Your CI builds are always cold even though you pass --cache-from pointing at the last pushed image. Why doesn't the builder stage ever hit cache?Pushing an image exports only the final stage's layers, so importing it as cache can never match instructions that ran in earlier stages. You need a cache export that includes intermediate stages — a registry or remote cache written with mode=max — and to import from that reference. Also check that the cache is for the same target platform, since caches are per-platform.
- Does changing the ENTRYPOINT in the final stage force the builder stage to rebuild?No. The builder stage is a separate chain in the graph and its cache keys are unaffected by instructions in the runtime stage. Only that final stage's tail is rebuilt, which is one of the concrete reasons multi-stage layouts iterate faster than one long linear file.
A linear Dockerfile is a single-file queue; multi-stage is a kitchen with several stations. Independent dishes cook at once, and burning the sauce doesn't mean re-baking the bread.
saying these in an interview costs you the question
- Thinking stages always execute strictly in the order written in the file
- Believing a change anywhere in the Dockerfile invalidates every stage
- Expecting --cache-from on a pushed image to restore builder-stage cache
- Copying the whole source tree before installing dependencies, then blaming the cache
- Assuming parallelism removes the need to order instructions by volatility