A docker build that takes 11 minutes dies at step 14 in CI but passes on your laptop — how do you triage it?
answer
- Evidence before hypotheses
- The 11-minute loop is the real enemy
- Stop the build above the failure
- Re-run one stage, not the whole build
- Then enumerate CI-vs-laptop differences
basics
~20 sCapture the real failure first with plain progress in CI, then shorten the loop before theorising: --target to stop just above the failing step, --no-cache-filter to re-run only that stage, and a shell there to run the command by hand.
solid answer
~50 sWork in two phases: get evidence, then shorten the loop. Evidence means `BUILDKIT_PROGRESS=plain` in the pipeline so the failing step's real output and the `failed to solve: process ... exit code: N` line are in the log, not a collapsed tail. Then stop paying 11 minutes per attempt: `--target <stage-just-above-the-failure>` builds only what that stage needs, and `--no-cache-filter=<stage>` re-runs one named stage against fresh state while everything below it stays cached — `--no-cache` would cost the full build every time. With a probe image in hand, `docker run --rm -it` into it and run the failing command by hand. Finally, enumerate what differs between the two environments rather than guessing: build args and secrets that exist locally but not in CI, `--platform` (an arm64 laptop against an amd64 runner), builder network or proxy access, an unpinned dependency that moved, and a stale or shared remote cache. Change one at a time.
code
bash · 3 linesexport BUILDKIT_PROGRESS=plain
docker build --target prebuild -t gw-probe .
docker run --rm -it gw-probe shgo deeper
Learn the two commands that make this tractable: plain progress to see the real error, and --target to stop the build just above the failing step so you can open a shell there.
Explain why --no-cache-filter=<stage> beats --no-cache on a long build, and what a probe stage reproduces that running the base image by hand does not.
Show a disciplined ladder — evidence, then loop time, then one variable per attempt — and a concrete list of laptop-versus-runner differences: build args and secrets, platform, builder egress, unpinned inputs, runner disk, imported cache.
Own the systemic side: how long a failing build takes to diagnose is a platform metric. Decide the pipeline's progress and cache-scoping conventions, whether Dockerfiles carry permanent debug stages, and what gets pinned so a green build stays green.
## The shape of the problem Take a concrete case: a Go telemetry ingest gateway shipped as a `scratch` image, 22 steps, an 11-minute cold build, failing at step 14 on the CI runner and green on a developer laptop. The temptation is to start pushing commits with speculative fixes. At 11 minutes an attempt, six speculative pushes is over an hour, and you still have no evidence. Triage is a ladder, taken in order. ## Rung 1 — make the failure legible Set `BUILDKIT_PROGRESS=plain` in the pipeline environment (or add `--progress=plain`). CI logs rendered with the TTY progress renderer are collapsed and full of control characters; plain output gives you every line of step 14 and, at the bottom, `ERROR: failed to solve: process "/bin/sh -c <command>" did not complete successfully: exit code: 2`. That line names the literal command and its exit status. Read it before forming any hypothesis. Half of these incidents end here: the message says `no space left on device`, `permission denied`, `dial tcp: i/o timeout` or a compiler error, and the hypothesis space collapses to one. ## Rung 2 — shorten the loop before you experiment An 11-minute iteration is the real enemy. Two flags fix it. `--target <stage>` builds only up to a named stage and skips every stage that stage does not need. Cut a boundary immediately above the failing instruction, and your attempt stops at roughly the time the build took to reach step 14 — and after the first run, most of it is cache. `--no-cache-filter=<stage>` re-runs one *named stage* from scratch while leaving the rest of the cache intact. Note the unit carefully: it takes stage names, not instruction numbers, and the stage must be named with `AS`. This is the flag for "I need step 14 to actually execute again, not replay as CACHED, but I am not paying to re-download the module cache." Reaching for `--no-cache` instead is the common error — it is correct but it costs the full 11 minutes on every attempt. ```bash docker build --target prebuild -t gw-probe . docker run --rm -it gw-probe sh docker build --no-cache-filter=build --progress=plain . ``` ## Rung 3 — run the failing command by hand With the probe image, open a shell and paste the failing command. You are now in the exact filesystem, working directory, environment and user the step had, with a terminal, and you can iterate in seconds: add verbosity flags, `env | sort`, check whether a file the step expects actually arrived, run the command with strace-free simple bisection. Note that the *final* image here is `scratch` and has no shell at all, so the builder stage is the only thing you can ever get a prompt in — another reason the probe is the technique, not an afterthought. ## Rung 4 — enumerate what differs, one variable at a time When the command succeeds by hand and fails in CI, the answer is an environment difference. The productive list: - **Build arguments and build secrets.** A value you have exported locally may simply be absent on the runner, and a missing build secret usually surfaces as an authentication or empty-file error inside the step rather than as a clear message. - **Architecture.** An arm64 laptop and an amd64 runner are different builds. `--platform` on the build, cross-compilation flags, and any dependency shipping per-architecture binaries all become suspects. Reproduce locally with the runner's platform before concluding anything. - **Network egress from the builder.** The step that fetches modules, packages or a private repository may face a proxy, an egress allowlist or a DNS resolver on the runner that you do not have locally. Timeouts and TLS errors point straight here. - **Unpinned inputs.** A floating base-image tag, an unpinned package or a dependency resolved at build time can move between your last local build and the CI run. If the CI build worked yesterday and nothing in the repository changed, this is the leading hypothesis. - **Disk and memory on the runner.** `no space left on device` and a step killed with no message are runner-capacity symptoms, not Dockerfile bugs. - **A stale or shared remote cache.** If the pipeline imports a cache produced by another branch or another architecture, a step can be served a result that does not match the source in front of it. Re-running the stage with `--no-cache-filter` is exactly the experiment that rules this in or out: if it passes with the stage forced to execute, the cache was the problem. Change one variable per attempt and write down what each attempt eliminated. Two changes at once on an 11-minute build buys confusion at 22 minutes a round. ## Rung 5 — leave the build easier to triage next time When it is fixed, spend ten minutes on the next incident: keep the probe stage in the Dockerfile with a real name so nobody has to invent one under pressure, set plain progress permanently in CI, pin whatever moved, and make the step that fails most often print enough context to be diagnosed from the log alone. The measure of a good triage is not only that the build is green — it is that the next failure is diagnosed from the first CI log instead of the sixth.
- What exactly does --no-cache-filter take as its value, and how is it different from --no-cache?It takes one or more build *stage names* — the names given with `FROM ... AS <name>` — and forces only those stages to execute rather than being served from cache. `--no-cache` invalidates the entire build. On a long build the difference is the whole point: you re-run the one stage you are investigating and keep the expensive dependency-fetch stages cached.
- The step passes when you run it by hand in the probe image but keeps failing in CI. What is your next move?Stop debugging the command and start comparing environments. Diff build args and secrets present locally versus on the runner, the build platform, the builder's network egress and proxy settings, and the runner's disk and memory headroom. Re-run the build on the runner with plain progress and one variable changed at a time until the failure moves.
- How would you tell whether a stale imported build cache is causing the failure?Force the suspect stage to execute with `--no-cache-filter=<stage>` while leaving everything else alone. If the build then passes, the step was being served a cached result that did not match the current inputs, and the fix is in how the pipeline imports or scopes its cache rather than in the Dockerfile.
- The failing build's final stage is a scratch image. Does that change your approach?It removes the fallback of building the image and poking around inside it — a scratch image has no shell, so there is nothing to exec into even on success. The debuggable artefact is the builder stage, which is a full distribution image. That makes the `--target` probe the primary technique here rather than a convenience.
saying these in an interview costs you the question
- Pushes speculative fixes and waits 11 minutes each
- Adds --no-cache to every attempt out of habit
- Thinks --no-cache-filter takes an instruction number
- Reads only the collapsed CI tail, never plain output
- Changes several variables in one attempt
- Assumes the laptop and the runner build identically