How can a build run its test suite without the test tooling reaching the shipped image?
answer
- stages form a graph, not a list
- a leaf stage ships nothing
- final stage copies nothing from tests
- stop the build at a chosen stage
- unreferenced stage may never run
basics
~20 sGive the tests their own stage built on the toolchain stage. The final stage copies nothing from it, so no test runner, fixture or report reaches the shipped layers, and the build can be asked to stop at that stage when you want the tests or their output.
solid answer
~50 sAdd a third stage. The toolchain stage produces the artifact; a `test` stage starts from the toolchain stage and runs the suite, with all the fixtures, runners and coverage tooling it needs; the `final` stage copies the artifact from the toolchain stage and nothing at all from the test stage. Because the test stage is not an ancestor of the final image, none of that tooling ships. Two practical levers follow. You can ask the builder to stop at a chosen stage as the build's endpoint, which is how you run the tests or extract a report without producing a shipped image at all. And if you want the final image to be unbuildable when tests fail, make the dependency real: have the test stage write a result file only on success and have the final stage copy that file.
code
pseudocode · 17 linesstage toolchain:
base = image carrying compiler and build tooling
compile source -> /out/app
stage tests:
from stage toolchain # already has artifact, source and tools
install test runner and fixtures
run suite
on pass: write /out/passed.txt
stage final:
base = small run-time base
take from stage tests: /out/passed.txt -> /passed.txt # real dependency
take from stage toolchain: /out/app -> /app
entry = /app
# suite fails -> /out/passed.txt absent -> final stage cannot be builtgo deeper
The takeaway is that a build can contain a stage whose output nobody ships. Tests can run inside the build and still leave no trace in the image, because the final stage simply does not copy anything from that stage.
Explain stages as a graph with copy edges, and why a stage that nothing references contributes nothing to the shipped image. Then name the two ways to actually guarantee the suite ran: make it the build's endpoint, or make it a real dependency.
The judgment is knowing which tests belong in a build stage at all — fast, hermetic ones — and which need real dependencies started around them, where the pipeline is the right owner and a build step is the wrong one.
Decide where verification is expressed across the organisation. Gates hidden inside build files are invisible to anyone reading the pipeline; the tradeoff is between a guarantee that travels with the build and a policy that can be audited in one place.
## Stages are a graph, not a sequence It is tempting to read a staged build as a list of steps that run top to bottom. It is better read as a **graph**: each stage names a base and may take files from other stages, and those references are the edges. The shipped image is one node in that graph — the final stage — together with everything it transitively depends on. A stage that no other stage takes files from is a leaf hanging off the side, and nothing it contains is in the image. That is what makes a test stage possible. Testing needs the opposite of what shipping needs: a test runner, fixtures, sample data, mock services, coverage instrumentation, sometimes a second copy of the source. All of it is legitimate during the build and none of it should be inside the boundary when the workload runs. ## The shape 1. **Toolchain stage** — compiler and build tooling, produces the artifact at a known path. 2. **Test stage** — starts from the toolchain stage, so it already has the artifact, the source and the tools; installs whatever else the suite needs; runs the suite. 3. **Final stage** — starts from a small base and copies the artifact from the *toolchain* stage. It does not copy from the test stage, so nothing from the test stage is in the shipped image. Two things you can then do that a single-stage build cannot: - **Stop the build at a chosen stage.** Naming an intermediate stage as the build's endpoint produces that stage's filesystem instead of the final image. That is how you run just the tests, or how you pull a coverage report or a generated inventory file out onto the local filesystem, without building or pushing anything. - **Export a result without shipping it.** A small stage that starts from an empty base and contains only the report is a clean thing to extract, because there is nothing else in it to sift through. ## The trap: a test stage nothing depends on If the final stage takes nothing from the test stage, the test stage is exactly the leaf described above — and **builders differ in whether they execute a stage that nothing references**. Some build the whole file; others build only what the requested output transitively needs, and quietly skip the rest. So a build that produces the final image is not, on its own, evidence that the tests ran. Make the gating explicit, one of two ways: - **Run the tests as their own build**, by naming the test stage as the endpoint. The build fails if the suite fails, and it produces no image. - **Make the dependency real.** Have the test stage write a result file only when the suite passes, and have the final stage copy that file in. Now the final image cannot be produced unless the test stage ran and succeeded, because the file it needs would not exist. The cost is one tiny file in the shipped image, which is usually a fair trade, and it is honest: the edge in the graph is the guarantee. ## What this does and does not buy - It keeps the test tooling out of the shipped layers, which is the size and attack-surface point this whole leaf is about. - It keeps the tests close to the artifact under test: the same filesystem, the same dependency set, the same build. - It does **not** decide *when* the suite runs, on which trigger, how its results are reported, or which stages run in parallel on which machine — that is the delivery pipeline's job, and a build file is a poor place to encode a release policy. - It does not replace tests that need real dependencies around them. A suite that needs a database, a message broker or the network is usually run by the pipeline against started services, not inside a build step. ## The sentence to say out loud *The test stage builds on the toolchain stage and the final stage takes nothing from it, so the tooling never reaches the shipped image — and because a stage nothing references may not be executed at all, the tests are either the build's endpoint or a real dependency of the final stage.* That covers both the mechanism and the failure mode people hit with it.
- How do you make the final image impossible to build when the suite fails?Give the final stage something to copy that only exists on success. If the test stage writes its result file only when the suite passes, and the final stage copies that file, the copy fails when the tests failed. The edge in the graph is what enforces it — merely defining a test stage does not, because a stage nothing references may not be executed.
- Where does this stop being the build's job?At policy. The build can express that a stage exists, what it needs and what depends on it. Which tests run on which trigger, how long the whole verification takes, where results are published and who may proceed to a deployment are decisions for the delivery pipeline, and encoding them in a build file hides them from everyone who reads the pipeline instead.
saying these in an interview costs you the question
- Assumes defining a test stage means the tests always ran
- Copies the test report into the final image and ships the fixtures too
- Believes every stage in a build file must end up in the image
- Thinks the build file is the right place to express release policy
- Runs tests in the final stage and then removes the test tooling