Under BuildKit, why is there no intermediate image to run after a failed build step, and how do you get a shell at that point?
answer
- The classic builder committed after each step
- BuildKit exports only what you ask for
- Stage results are the exportable unit
- Cut a boundary above the failing instruction
- --target the stage, then run it with sh
basics
~20 sBuildKit commits no image per instruction, so a failed build leaves no id to run — the classic-builder trick is gone. Instead cut a stage boundary just before the failing instruction, build with --target that stage, and run it interactively.
solid answer
~50 sThe classic builder printed `---> <id>` after every instruction because it committed a container to an image at each step, so after a failure you could `docker run -it <last id> sh`. BuildKit executes the Dockerfile as a graph of steps and keeps each step's result as a snapshot in the **builder's own cache**, not in the image store; nothing is exported unless you ask for it, so `docker images` shows nothing new and there is no id to run. What replaces it is `--target`: name a stage ending immediately before the failing instruction, run `docker build --target <stage> -t probe .`, then `docker run --rm -it probe sh`. You are standing in exactly the filesystem, working directory and environment the failing `RUN` saw, and can paste the command in. If no boundary sits where you need one, add a temporary `FROM <previous stage> AS <name>` line to cut one.
code
dockerfile · 14 linesFROM golang:1.24-alpine AS deps
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
FROM deps AS prebuild
COPY . .
FROM prebuild AS build
RUN CGO_ENABLED=0 go build -trimpath -o /out/gateway ./cmd/gateway
FROM scratch
COPY --from=build /out/gateway /gateway
ENTRYPOINT ["/gateway"]go deeper
Know that a failed build leaves you no image to run and that --target <stage> builds only up to a named stage. Being able to build a probe image and open a shell in it is the expectation here.
Explain the mechanism: the classic builder committed an image per instruction, BuildKit keeps step results in the builder's cache and exports only the stage you request. Then show the --target probe and why the boundary goes above the failing instruction.
Demonstrate that you reproduce the failing step's exact environment rather than approximating it in the base image, and that you know the probe stage costs almost nothing because earlier steps come from cache.
Own the tradeoff BuildKit made — concurrency, cache sharing and secret handling in exchange for per-instruction images — and decide whether your Dockerfiles carry permanent, named debug stages so on-call engineers have a supported entry point instead of editing files under pressure.
## Why the old trick stopped working The classic builder built an image by running one container per instruction and committing it. Its output looked like this: ``` Step 12/18 : RUN go build -o /out/gateway ./cmd/gateway ---> Running in 7c1f2a ---> 9b4e11 ``` Each `---> <id>` was a real image, and after a failure the id printed by the last *successful* step was still sitting in `docker images` (as a dangling image). The universal debugging move was `docker run --rm -it 9b4e11 sh`, and from there you re-ran the failing command by hand. BuildKit — the default builder in Docker Engine 23.0 and later — does not work that way. It resolves the Dockerfile into a graph of steps, executes independent steps concurrently, and stores each completed step's filesystem as a content-addressed snapshot in the **builder's cache**, which is a separate store from the engine's images. An image is produced only by an *export* at the end of the build, for the stage you asked for. When the build fails, no export happens, so nothing lands in `docker images`, there are no dangling intermediates to list, and there is no id to run. The snapshot of the step before the failure exists, but it is not addressable as an image. This is a real ergonomics loss that BuildKit trades for concurrency, better caching and secret handling, and it is why interviewers ask the question: a candidate who still answers "run the intermediate image" has not built an image since 2022. ## The replacement: --target a probe stage `--target <stage>` tells BuildKit to build only up to the named stage and export *that* as the image, skipping every stage that stage does not depend on. Stage results **are** exportable images, so a stage boundary is the supported way to get a runnable artefact from the middle of a build. So the technique is: cut a stage boundary immediately before the failing instruction. ```dockerfile FROM golang:1.24-alpine AS deps WORKDIR /src COPY go.mod go.sum ./ RUN go mod download FROM deps AS prebuild COPY . . FROM prebuild AS build RUN CGO_ENABLED=0 go build -trimpath -o /out/gateway ./cmd/gateway FROM scratch COPY --from=build /out/gateway /gateway ENTRYPOINT ["/gateway"] ``` `docker build --target prebuild -t gateway-probe .` gives you an image holding everything the failing `go build` could see. `docker run --rm -it gateway-probe sh` puts you in it. Now paste the command: ``` /src # CGO_ENABLED=0 go build -trimpath -o /out/gateway ./cmd/gateway ``` and you get the error interactively, with a shell to explore: `ls`, `env`, `cat go.mod`, run it again with more verbosity, edit a file and retry. That loop is worth far more than another full build. If the Dockerfile has no boundary where you need one, add one temporarily — a `FROM <previous stage> AS probe` line costs nothing and is deleted afterwards — or simply comment out the failing instruction and everything below it and build without `--target` at all. Both produce the same thing. ## Details that matter in the room **The probe must be *before* the failure, not the enclosing stage.** `--target build` in the file above still runs the failing `go build`; you want the stage that ends just above it. **Watch the shape of the final stage.** A Go binary shipped `FROM scratch` has no shell at all, so "just run the image" was never available here even on a success; the thing you debug is the builder stage, which is a full distribution image with a shell. **Running the base image alone is a weaker version of the same idea.** `docker run --rm -it golang:1.24-alpine sh` and typing the commands by hand does reproduce some failures, but you are missing every `COPY`, `ENV`, `ARG`, `WORKDIR` and `USER` the build had applied. Differences between your hand-run and the build's environment are exactly where these bugs hide, so prefer a `--target` probe, which reproduces them for free. **Non-interactive context.** Even inside the probe, remember the build ran with no TTY and as whatever `USER` was in effect; a command that behaves differently under a terminal (a package manager prompting, a tool colourising output) can pass by hand and fail in the build. Re-run with the same flags the Dockerfile used before concluding the environment is identical. **There is an experimental shortcut.** buildx ships a debug monitor that can drop you into the container of the step that failed, gated behind the `BUILDX_EXPERIMENTAL=1` environment variable. It is genuinely useful, but its command surface has moved between buildx versions, so know it exists and do not build a team runbook on it. The `--target` probe works on every BuildKit version and in CI.
- The Dockerfile has one long stage and no boundary where you need one. What do you do?Add one temporarily: `FROM <previous stage> AS probe` immediately above the failing instruction, build with `--target probe`, and delete the line afterwards. The equivalent without editing stage structure is to comment out the failing instruction and everything below it and build the file as it stands. Both give you an exportable image at the point you care about.
- Why is running the base image and typing the commands by hand a weaker technique than a --target probe?The base image has none of the build's accumulated state: no COPY'd files, no ENV or ARG values, no WORKDIR, no USER, and none of the packages earlier steps installed. Those differences are frequently the bug. A `--target` probe reproduces the exact filesystem and environment the failing instruction ran in, so anything that still fails there is the real defect.
- After a BuildKit build fails, is the completed work lost?No. Each finished step's result stays in the builder's cache, so re-running the build replays those steps as CACHED in seconds and stops at the same place. What is missing is only an *image* for them — the cache is not the image store, and nothing is exported until a build completes the stage you targeted.
The classic builder was a photographer taking a print at every step of the journey, so you could always pick up the last print. BuildKit keeps negatives in its own darkroom and prints only the frame you name — so you name a frame just before the crash.
saying these in an interview costs you the question
- Says to docker run the last intermediate image id
- Expects docker images to list per-instruction layers
- Thinks a failed build leaves dangling images to run
- Targets the stage that contains the failing instruction
- Claims BuildKit throws away all completed step results
- Debugs in the base image and ignores COPY, ENV and WORKDIR