Why is the Docker build cache empty on every job when CI runners are ephemeral, and where can it live instead?
answer
- Ask what survives the machine
- Flat build times are the symptom
- Builder state lives in a volume
- The cache must go somewhere external
- Importing over the network is not free
basics
~20 sBuildKit keeps its cache in the builder's own state — a volume on the runner — and an ephemeral machine is destroyed after the job. To reuse anything, export cache to an external backend such as a registry and import it next build.
solid answer
~50 sA build cache is only useful if the next build can reach it. BuildKit keeps its cache in the builder's state — for a `docker-container` builder, a Docker volume on the runner — and an ephemeral runner throws that machine away when the job ends. So every job re-executes every instruction, and the tell-tale symptom is a build time that never varies no matter how small the change. To break that, the cache has to outlive the machine: BuildKit can export cache to an external backend (commonly a registry repository, sometimes a CI-provided cache service or a directory on shared storage) with `--cache-to`, and the next build imports it with `--cache-from`. The tradeoff is that importing is a network operation, so you must measure: a cache import that costs 38 s only pays for itself if it saves more than 38 s of rebuilding.
code
bash · 5 linesdocker buildx create --name ci --driver docker-container --use --bootstrap
docker buildx build \
--cache-from type=registry,ref=registry.example.com/geo/tile-api:buildcache \
--cache-to type=registry,ref=registry.example.com/geo/tile-api:buildcache \
--push -t registry.example.com/geo/tile-api:build-4712 .go deeper
Recall that the build cache is local to the machine that built, so a fresh runner has none, and that CI has to fetch cache from somewhere external for a build to be faster than the first one.
Explain where BuildKit keeps cache — the builder's own state volume — and how export and import move it to and from an external backend, plus why that requires a builder that is not the engine's embedded one.
Demonstrate measurement: import and export cost against time saved, recognising flat build durations as the cold-cache signature, and knowing that an image pull is not a warm build cache.
Own the fleet economics: which pipelines get a cache, where it lives, who pays for its storage and retention, and when the right answer is warm builders instead of shipping a large cache over the network for every job.
### Why the cache is gone, mechanically BuildKit records the result of each build step in a content-addressed cache held in **the builder's own state**. For a builder created with the `docker-container` driver, that state is a Docker volume attached to the BuildKit container on whichever machine ran the job. It is not in the image store, not in the repository, and not in the build context — it is local storage belonging to a container on one host. An ephemeral runner is a machine created for one job and destroyed afterwards. Both the BuildKit container and its state volume die with it. The next job gets a clean machine, creates a fresh builder, and BuildKit correctly concludes it has never seen any of these steps before. Nothing is broken; the cache simply does not exist. The symptom is characteristic and worth recognising in a log: build duration that is flat. A geospatial tile server built as a Node.js API on an alpine base takes 6 m 41 s to build from scratch, and on a 47-node fleet of throwaway runners it takes about 6 m 41 s for a one-line change to a route handler, for a README typo, and for a dependency bump alike. Nothing gets faster, ever. Compare that with a developer laptop, where the second build of the same commit finishes in seconds, and the difference is entirely the presence of local state. ### Making the cache outlive the machine BuildKit can move cache in and out of an external backend: `--cache-to` writes it at the end of a build, `--cache-from` reads it at the beginning. The backends differ in where the bytes live — a registry repository, a CI provider's own cache service, a directory on shared storage — but the shape is the same, and exporting a standalone cache is one of the capabilities that requires a builder on the `docker-container` (or `remote`) driver rather than the engine's embedded one. A registry backend is the common choice in CI because the runner already has network access and credentials to a registry, and because the cache then follows the same access rules as the images. Import is lazy: the builder fetches the cache manifest, works out which steps it can reuse, and pulls only the blobs it actually needs rather than the whole cache. ### The economics — measure, do not assume An external cache converts CPU time into network time, and that trade is not always favourable: - **Import cost is real.** In the tile-server example the cache image is 1.34 GB and importing it costs about 38 s on a runner with ordinary network. If reusing it turns a 6 m 41 s build into 1 m 12 s, the trade is excellent. If a build only takes 50 s cold, an import that costs 38 s buys almost nothing. - **Export cost is real too.** Writing cache at the end of every build adds upload time and registry storage, on every job, including the ones that will never be reused. - **Hit rate decides everything.** Cache that no later build can match is pure overhead paid twice, so the practical question is which builds should read from and write to which cache location — for example whether feature-branch jobs read a cache produced by the mainline and refrain from writing their own. - **Storage grows.** A cache backend accumulates; on a 47-node fleet producing many builds a day it needs a retention policy, or it becomes the largest thing in the registry. ### What is not a substitute Several things get mistaken for a warm cache and are not: - **Pulling the application image.** Having the previous image on the runner does not populate BuildKit's step cache; layers in an image and cache entries for build steps are different records. - **A shared base image.** Starting from a base that is already local saves the pull, not the build steps that follow it. - **A longer runner lifetime by accident.** If some runners happen to be reused, builds get randomly fast and slow, which is worse than uniformly cold: it hides the real cost and makes timing regressions invisible. ### Diagnosing it honestly Run the same commit twice and compare. A warm build should show reused steps in the log and finish far faster; a cold one re-executes everything. Watching the difference between an import that happens and one that finds nothing to match is how you tell a *missing* cache from an *ineffective* one — and it is the number worth putting into a job's log, because on a large fleet a few seconds per build multiplies into hours of runner time a day.
- How would you prove that adding an external cache actually made CI faster?Measure the same commit both ways on the same class of runner: cold with no cache import, then warm. Record build duration, import duration and export duration separately, because a build that got 90 s faster while paying 38 s to import and 25 s to export gained less than the headline suggests. Then look at the distribution across many jobs, not one run, since hit rate is what dominates on a fleet.
- Why does having the previous application image on the runner not warm the build?Image layers and build-step cache records are different things. An image tells you what the filesystem looked like; the build cache records which step produced which result so BuildKit can decide a step need not run again. Pulling the image gives you the former, so the build re-executes its steps regardless.
- What operational cost does a registry-backed cache add on a large fleet?Storage and lifecycle. Every job that exports cache writes into the registry, and unattended that grows without bound, sometimes outweighing the images themselves. It needs retention just like any other artefact, plus attention to which jobs are allowed to write cache at all — otherwise short-lived branch builds fill it with entries nothing will ever reuse.
It is a hotel room, not your desk: every night you get a different one and it is always empty. Anything you want tomorrow has to go into a locker somewhere else — and fetching it from the locker each morning is a cost you can measure.
saying these in an interview costs you the question
- Blames a slow registry rather than an absent cache
- Thinks the daemon shares cache between machines
- Assumes importing a cache is free
- Claims pulling the previous image warms the build cache
- Never measures cold versus warm build time
- Adds cache export everywhere without checking hit rate