skip to content

Why does `docker build .` never run the tests in your Dockerfile's test stage?

level: middleimportance: should knowfreq 55%

answer

  1. The build does not read the file top to bottom
  2. Something has to select the stage
  3. The default target is the last stage
  4. Unreachable stages are pruned before anything runs

basics

~20 s

BuildKit builds only the target stage and the stages that target depends on. The default target is the last stage in the file, so a test stage nothing depends on is pruned and never executed. Run it explicitly with docker build --target test.

solid answer

~50 s

BuildKit resolves the Dockerfile into a graph and builds only what the target needs. With no `--target`, the target is the **last** stage in the file — typically the runtime stage — so a test stage that no later stage pulls anything from is simply not part of the graph and never runs. Two fixes. Either invoke it explicitly, `docker build --target test .`, as its own step whose exit status gates the pipeline, and then build the runtime image in a second invocation that reuses the shared cache. Or make the final stage depend on something the test stage produced, so the graph cannot be pruned — at the price of coupling every image build to the suite. This is also a builder difference: the classic pre-BuildKit builder executed every stage in file order up to the target, so old pipelines that relied on the tests running by accident break on upgrade.

code

bash · 3 lines
bash
# gate on the suite, then build what ships
docker build --target test -t chatfanout-test .
docker build -t chatfanout:1.4.2 .

go deeper

for a junior

Know that --target names a stage and that a plain docker build targets the last stage in the file. If the pipeline never passes --target, the test stage is not being built.

for a middle

Explain the graph: BuildKit builds the target's transitive dependencies and prunes the rest, so a test stage nothing consumes never runs. Be able to name both fixes and say which invocation gates the pipeline.

for a senior

Show you would diagnose this from the build log rather than assume, and know the classic-builder difference that makes it appear after an engine upgrade with the build still green.

for a principal

Own the choice between two invocations and a Dockerfile-enforced dependency: one keeps builds fast and the policy visible in the pipeline, the other guarantees the gate but couples every image build to the suite.

### The default target is the last stage `docker build .` does not mean "execute this file top to bottom". BuildKit parses the Dockerfile into a dependency graph, picks a target stage, and then builds the transitive closure of that stage's dependencies — nothing else. When you do not pass `--target`, the target is the last stage in the file. Everything reachable from it gets built; everything else is pruned before a single container starts. A test stage is usually a leaf in that graph. It depends on a base or dependency stage, but the runtime stage takes nothing from it, so no path leads from the target back to the tests. The stage is dead code as far as the build is concerned, and the build finishes green having never run an example. ### Building it on purpose `docker build --target test .` makes the test stage the target. BuildKit then builds it and its dependencies, the suite executes as a `RUN` step, and a red suite exits the build non-zero — which is exactly the gate you wanted. The command is a normal build, so it can be tagged (`-t chatfanout-test`) if you want an image of the test stage to start a container from afterwards. Running both — the test target, then the default runtime build — costs far less than twice the time. The two builds share the base and dependency stages, so the second invocation hits the cache for everything they have in common, and it does not re-run the test step because the test stage is not in the runtime target's graph at all. ### Making the tests impossible to skip If you want a single `docker build` that cannot ship an untested image, the graph itself has to force it: the final stage must consume something the test stage produced — a marker file, the coverage summary, anything — so the test stage becomes a dependency of the target. It works, and it is the only way one invocation can guarantee the suite ran. The price is real. Every image build now waits for the suite, including builds you only wanted for a quick local check. Cache invalidation gets coarser, because anything that busts the test stage also busts the final image. And you have wired a policy decision ("images must be tested") into a Dockerfile, where the next person to add `--target` or a second final stage can quietly undo it. Most teams keep the two invocations and enforce the ordering in the pipeline instead. ### The builder difference that surprises people BuildKit prunes unreachable stages. The classic builder that predates it did not: it walked the stages in file order and executed each one up to the target, whether or not anything depended on it. A Dockerfile written against the old behaviour, with a test stage sitting in the middle and no `--target` anywhere in the pipeline, ran its suite on every build. Move that same file to a modern engine, where BuildKit is the default builder for `docker build` on Linux from Docker Engine 23.0 onward, and the tests silently stop running. The build stays green, only faster — which is how the regression goes unnoticed. Setting `DOCKER_BUILDKIT=0` restores the old behaviour on engines that still ship the classic builder, but that is a diagnosis aid, not a fix. ### A concrete shape For a Ruby chat-message fan-out worker: ```dockerfile FROM ruby:3.3-bookworm AS deps WORKDIR /app COPY Gemfile Gemfile.lock ./ RUN bundle install FROM deps AS test COPY . . RUN bundle exec rspec FROM ruby:3.3-slim-bookworm AS runtime WORKDIR /app COPY . . CMD ["bundle", "exec", "ruby", "worker.rb"] ``` `docker build .` targets `runtime`, whose graph contains only itself — the 1,847 examples never execute. `docker build --target test .` runs them. Reading the build output tells you which happened: a build that never printed the suite's output did not run it, and `--progress=plain` makes that easy to confirm in a CI log. ### What to check when someone reports "the tests stopped running" Ask which target the pipeline builds, look for a `--target` flag, and check whether the final stage consumes anything from the test stage. Then read the build log for the test step's output. Almost every instance of this bug is one of three things: no `--target` in the pipeline, a stage renamed so `--target` now names something else, or a Dockerfile that used to rely on the classic builder's sequential behaviour.

  • How would you make the tests impossible to skip in a single build?
    Make the final stage consume something the test stage produced — a marker file or the coverage summary — so the test stage is a dependency of the target and cannot be pruned. It guarantees the suite ran, but every image build now waits for it, cache invalidation becomes coarser, and the policy is buried in the Dockerfile where a later `--target` can undo it.
  • Does running --target test and then a normal build do the work twice?
    No. The two builds share their base and dependency stages, so the second invocation hits the cache for all of them. The test step itself is not repeated either, because the test stage is not in the runtime target's graph. You pay for the suite once and for a little extra graph resolution.
  • What does docker build --target test produce when it succeeds?
    An ordinary image of that stage, tagged if you passed `-t`. That is useful beyond the gate: you can start a container from it to re-run a single failing spec interactively, or run it with a bind mount so the suite writes its reports somewhere you can read them.

saying these in an interview costs you the question

  • Assumes every stage in a Dockerfile is built on every build
  • Thinks --target only changes which image gets tagged
  • Believes stage order in the file decides what executes
  • Says the tests run because the stage appears before runtime
  • Expects two invocations to repeat all the shared work

context