Your CI runs Docker builds on ephemeral runners, so every job starts with an empty builder and rebuilds everything. How would you design the build cache strategy, and what are the trade-offs of the options BuildKit gives you?
answer
- ephemeral runner = cold builder every job
- inline = free but final stage only
- registry mode=max = all stages, costs push + storage
- cache MOUNTS never export: needs a persistent builder
- scope per branch with main fallback, and per platform
basics
~20 sExport the build cache to somewhere shared. BuildKit supports --cache-to/--cache-from with an inline cache (embedded in the pushed image, final stage only), a registry cache (separate tag, mode=max keeps intermediate stages), or a local or cloud backend. Weigh cache hit rate against push/pull cost and storage, and scope caches per branch and per platform.
solid answer
~60 sEphemeral runners make the builder-local cache worthless, so the cache must live somewhere shared, and the choice is about hit rate versus transfer cost. - **Inline cache** (`--cache-to type=inline`, then `--cache-from <image>`): metadata is embedded in the pushed image, so there is no extra artefact. Cheap, but only the final stage's layers are cached, which is weak for multi-stage builds. - **Registry cache** (`--cache-to type=registry,ref=repo/app:buildcache,mode=max`): a separate tag holding all stages including intermediates. Much better hit rate; costs registry storage and a push every build, and needs a retention policy. - **Local / GHA / S3 backends** for CI-native storage, with the same trade-offs. Note cache **mounts** are not exported by any of these, so package-manager caches still start cold unless you keep a long-lived builder. At scale I usually prefer a persistent shared builder (remote or Kubernetes driver) over exporting on every job: hit rates are higher and there is no upload cost. Either way, scope the cache key by branch with a fallback to the main branch, and per platform.
code
bash · 5 linesdocker buildx build \
--cache-from type=registry,ref=repo/app:cache-$BRANCH \
--cache-from type=registry,ref=repo/app:cache-main \
--cache-to type=registry,ref=repo/app:cache-$BRANCH,mode=max \
-t repo/app:$SHA --push .go deeper
Know that the build cache lives on the builder and that CI can share it by exporting to a registry with --cache-to and --cache-from.
Distinguish inline from registry cache and know that mode=max includes intermediate stages, plus that cache mounts are not exported.
Design it: branch-scoped refs with a main fallback, per-platform separation, retention and pruning, and choosing between export/import and a persistent builder.
Frame the decision economically and operationally: measured build-time breakdown, storage and egress cost, shared-builder isolation and failure domain, and when release builds should deliberately run cold for provenance.
## The problem BuildKit's cache lives on the builder. Ephemeral CI runners create a new builder per job, so the layer cache is empty, cache mounts are empty, and every build is cold. On a large image this can dominate pipeline time and cost. There are two families of solution: move the cache to shared storage, or stop throwing the builder away. ## Exporting cache to shared storage BuildKit separates cache **export** (--cache-to) from cache **import** (--cache-from), and supports several backends. **Inline** (--cache-to type=inline). Cache metadata is embedded in the image config that you push anyway. Import with --cache-from type=registry,ref=<your image tag>. Zero extra artefacts and no extra storage. The limitation is decisive for modern Dockerfiles: only layers present in the final image are described, so in a multi-stage build the entire compile stage is not cached. It suits single-stage or thin-final-stage images poorly and is mostly a legacy convenience. **Registry** (--cache-to type=registry,ref=repo/app:buildcache,mode=max). Writes a dedicated cache artefact to a registry tag. mode=min stores only the final stage's layers; **mode=max stores every stage including intermediates**, which is what makes it useful. Hit rates are good, any runner anywhere can import it, and it survives the runner. Costs: an extra push per build (sometimes large), registry storage that grows without a retention policy, and pull time on import. Some registries do not support the artefact types used by cache manifests, which forces image-manifest mode. **CI-native backends**: type=gha for GitHub Actions cache, type=s3 or type=azblob for object storage, type=local for a directory the CI system caches. These trade registry coupling for CI coupling, and each carries its own size limits and eviction rules; the GitHub Actions cache in particular evicts by age and total repository quota, which silently degrades hit rate. ## Keeping the builder instead The alternative is to stop destroying the cache. A long-lived buildx builder (the docker-container driver on a persistent host, the remote driver, or the kubernetes driver running BuildKit pods with persistent volumes) keeps both the layer cache and the cache mounts warm. Jobs connect to it rather than building locally. This usually beats export-import on hit rate, because nothing is serialised, uploaded and re-downloaded, and because cache mounts (npm, Maven, Go, cargo) work at all, which no export backend provides. The costs are operational: you now run infrastructure, you must size disks and prune (docker builder prune with keep-storage), you must think about multi-tenancy and isolation between untrusted builds sharing a builder, and a busy shared builder can become a queueing bottleneck. Isolation is a genuine security consideration: a shared builder means one repository's build can, through cache poisoning or resource exhaustion, affect another's. ## Scoping and correctness Whatever the backend, cache scope matters. Cache keys are content-derived, so importing a cache from an unrelated branch is safe but often useless. A common pattern is to export per branch and import with an ordered fallback: the current branch's cache first, then the default branch's. That gives feature branches a warm start without them polluting each other. Per-platform scoping is mandatory for multi-architecture builds: an amd64 build's cache does not accelerate arm64, and mixing them into one ref wastes storage. For release builds, some teams deliberately build cold, or with --no-cache, to avoid the possibility that a poisoned or stale cache contributes to a shipped artefact, and to keep the provenance story simple. That is a real trade: reproducibility and supply-chain confidence against speed. ## Deciding Measure before engineering. Establish how much of pipeline time is build, and which stages dominate. If dependency resolution dominates, cache mounts on a persistent builder are worth far more than layer cache export. If the base and system-package layers dominate, registry cache with mode=max is easy and effective. If images are small and builds are already a minute, the simplest option that works is correct, and elaborate cache infrastructure is cost without return. Also weigh money and blast radius: registry storage and egress are not free at fleet scale, and a shared builder is a shared failure domain. The principal-level answer states the decision criteria and the exit conditions, not a single blessed configuration.
- Why is inline cache often disappointing for a multi-stage build?Inline cache embeds cache metadata in the pushed image, so it can only describe layers that are actually part of that image. In a multi-stage build the final image contains just the runtime layers, so the expensive compile and dependency stages are not represented and will be rebuilt from scratch. Registry cache with mode=max exports intermediate stages as well, which is where the real time is spent.
- Even with registry cache configured, dependency downloads still happen on every CI build. Why?Cache mounts are builder-local state and are excluded from every cache export backend, so npm, Maven, Go and cargo caches always start empty on a fresh builder. The only ways to keep them warm are a persistent or remote builder that owns that storage, or snapshotting the directory through the CI system's own cache, which reintroduces upload and download cost.
- What are the risks of a single shared BuildKit builder serving many repositories?It is a shared failure domain and a shared trust boundary: disk exhaustion or a hot build queues everyone, and builds from different repositories share cache state, so a poisoned or malicious cache entry can influence another team's artefact. Mitigations are per-tenant builders or namespaces, storage quotas with aggressive pruning, and cold or --no-cache builds for release artefacts where provenance matters most.
saying these in an interview costs you the question
- Assuming --cache-to also exports RUN --mount=type=cache contents
- Using inline cache with a multi-stage build and expecting the build stage to be cached
- Exporting cache with mode=max and never setting a retention policy, then blaming registry cost
- Sharing one cache ref across architectures or across all branches indiscriminately
- Adding cache infrastructure without measuring which stage actually dominates build time