skip to content

In a multi-stage Docker build, what does a green test stage not prove about the shipped image?

level: seniorimportance: should knowfreq 38%

answer

  1. The suite and the shipment are two images
  2. Everything the slim base left out is untested
  3. Production flags were never applied during the suite
  4. Start the artefact you built, with its real limits

basics

~20 s

The test stage is a different image: fuller base, dev dependencies, usually root, no runtime limits. A green suite proves the code passes, not that the shipped image starts — missing libraries, USER permissions, a wrong ENTRYPOINT and memory limits all survive it.

solid answer

~50 s

A test stage and a runtime stage are two different images that happen to share a Dockerfile. The test stage usually sits on a fuller base, carries dev and test dependencies, runs as root, has the source tree present, and runs with none of the flags production uses. So a green suite says the code is correct; it says nothing about whether the shipped image can start. What it misses is a specific list: a shared library present in the build base but absent from the slim runtime base, files the app writes to a directory its `USER` cannot write, a wrong `ENTRYPOINT`/`CMD` or exec-vs-shell form, missing CA certificates or timezone data, and resource limits — a worker that passes 1,847 examples on a laptop gets OOM-killed with exit 137 under `--memory=128m`. Close the gap with a cheap smoke run of the built runtime image under its production flags, not by moving the whole suite into it.

code

bash · 3 lines
bash
docker build -t chatfanout:1.4.2 .
docker run --rm --memory=128m --user 10001 --read-only --tmpfs /tmp \
  chatfanout:1.4.2 --self-check

go deeper

for a junior

Know that the stage the tests ran in and the stage that ships are different images with different contents, and that starting the built image once is a cheap check worth doing.

for a middle

Explain concrete gaps: a library present in the build base but not the slim runtime base, permissions under a non-root USER, CMD in shell form breaking signal delivery, and limits that were never applied during the suite.

for a senior

Show the pipeline design: keep the deep suite in the test stage and add a short smoke run of the built image under production flags, and be able to read exit 137 and OOMKilled as a limit failure rather than a code failure.

for a principal

Own the fidelity-versus-cost tradeoff — how close to production the pipeline should model, what you deliberately leave unmodelled, and how you stop teams from restoring fidelity by dragging test dependencies into the shipped image.

### Two images, one file A multi-stage Dockerfile makes it easy to forget that the stage the suite ran in and the stage you push are separate images with separate filesystems. Typically the test stage is `FROM ruby:3.3-bookworm` — full toolchain, headers, dev and test gems, root user, the whole source tree, an environment that says `test` — and the runtime stage is `FROM ruby:3.3-slim-bookworm` with a non-root `USER`, only production dependencies, no source for the specs, and a `CMD` nobody has ever executed. The suite validated the first image. You ship the second. ### What that gap actually hides The failures that survive a green suite are boring and repetitive, which is what makes them worth listing: - **Missing shared libraries.** A native extension compiled in the build stage links against a library that the slim base does not ship. The suite loads it fine; the runtime image dies at require time. - **Missing OS data.** CA certificates, timezone data, locales: absent from a minimal base, so the first outbound TLS call or timestamp formatting fails, and only in production. - **User and permissions.** The suite ran as root. Under `USER 10001` the app cannot write its pid file, its log directory, or a cache path, and a read-only root filesystem turns every temp write into an error. - **Entrypoint and signals.** `CMD` in shell form makes the shell PID 1, so the process never receives `SIGTERM` on `docker stop` and every deploy takes the full timeout. No unit test can see this. - **Configuration shape.** Environment variables the test stage set inline are supplied differently at runtime; a missing one that defaulted harmlessly in tests is fatal in the shipped image. - **Resource limits.** The suite ran with the whole host's memory. A chat-message fan-out worker that buffers per-recipient batches sits comfortably under test and is killed at `--memory=128m`, showing exit code 137 and `OOMKilled: true` in `docker inspect`. ### Narrowing the gap without pretending to close it The instinct — "then run the suite inside the runtime image" — is usually wrong. To do it you must add the test framework, the specs, and their dependencies to the image, at which point it is no longer the image you ship, and you have re-created the same gap one layer down while also inflating what you push. The cheaper and more honest move is a **smoke check against the artefact you actually built**, with the flags production uses: ```bash docker build -t chatfanout:1.4.2 . docker run --rm --memory=128m --user 10001 --read-only \ --tmpfs /tmp -e QUEUE_URL=... chatfanout:1.4.2 --self-check ``` That single run exercises the runtime base's libraries, the `USER`, the read-only filesystem, the entrypoint, the config wiring and the memory limit — the entire list above — in a couple of seconds, and it fails loudly with a non-zero exit or a 137 rather than silently in production. If the app has no self-check mode, start it, wait for its `HEALTHCHECK` to report healthy or poll its readiness endpoint, then stop it and assert it exited on `SIGTERM` rather than being killed after the timeout. ### Where the line sits Think of it as two questions with two owners. "Is the code correct?" belongs to the suite in the test stage, where the toolchain makes it fast and debuggable. "Is this artefact runnable as configured?" belongs to a short check against the built image under production flags. Teams that conflate them either ship untested artefacts on a green suite, or drag the whole suite into the runtime image and lose the minimality that made it worth building. And be explicit about what the smoke check still does not prove: it runs on one host, with one architecture unless you check each platform of a multi-arch build, without the network topology, secrets, or neighbours of the real environment. It moves the class of failures from "discovered in production" to "discovered in the pipeline", which is the win — not a guarantee of correctness in the environment you have not modelled.

  • How would you smoke-test a runtime image that has no test framework in it?
    Black-box it. Start the built image with its production flags and either invoke a self-check subcommand, wait for its HEALTHCHECK to report healthy, or poll a readiness endpoint from another container on the same user-defined network. Then `docker stop` it and check it exited promptly on SIGTERM instead of after the timeout. No test framework needs to be in the image for any of that.
  • How do you catch the memory failure that the test stage cannot show you?
    Run the built image under the same `--memory` limit production uses and give it representative work. If the kernel kills it, the container exits 137 and `docker inspect` reports `"OOMKilled": true` — an unambiguous signal that the limit, not the code, is what failed. The test stage runs with the host's whole memory, so it can never surface this.
  • Is it worth running the entire suite inside the runtime image?
    Rarely. You would have to add the framework, the specs and their dependencies to the image, so the thing you tested is no longer the thing you ship and the image you push grows for no runtime benefit. Keep the deep suite in the test stage and spend the fidelity budget on a short check against the real artefact.

A dress rehearsal in the rehearsal room proves the cast knows the lines. It does not prove the set fits on the actual stage, that the doors open, or that the lighting rig can carry the load.

saying these in an interview costs you the question

  • Treats the test stage and the runtime stage as the same image
  • Assumes a passing suite means the container will start
  • Adds test dependencies to the shipped image to run the suite
  • Ignores USER, read-only rootfs and memory limits when testing
  • Never starts the built image before pushing it
  • Blames the application when a container exits 137

context