Would you give a 47-node CI fleet one shared BuildKit builder or a fresh builder per job?
answer
- Two ends: disposable or long-lived
- Warm cache versus isolation and ownership
- A shared builder is stateful infrastructure
- Cache mounts are shared between tenants
- Pools bound the blast radius
basics
~20 sIt trades warm cache against isolation. A long-lived shared builder keeps cache local and skips import cost, but is stateful infrastructure with shared cache mounts and fleet-wide blast radius; a per-job builder is disposable but always cold.
solid answer
~40 sPer-job builders created with `docker buildx create --driver docker-container` are the default: disposable, isolated, and needing no operator. The price is a cold start on every job and the cost of importing cache over the network. A long-lived builder shared through the `remote` driver keeps cache on local disk, so hits are effectively free — but it becomes a stateful service someone must size, garbage-collect and answer for when it dies and the whole fleet has nowhere to build. It also merges trust boundaries, since its cache and build-time cache mounts are shared by every tenant. I decide on build shape, the trust boundary, measured import cost against build time, and whether anyone can own the service. The usual answer is a middle one: pooled builders for the few hot repositories, per-job builders elsewhere.
code
bash · 3 linesdocker buildx create --name shared --driver remote tcp://buildkit.ci.internal:1234 --use
docker buildx du --builder shared
docker buildx prune --builder shared --filter until=168hgo deeper
Recall the basic difference: a builder created inside a job disappears with it, while a long-lived builder elsewhere keeps its cache between jobs. Knowing that both exist is enough at this level.
Explain why a warm builder is faster — local cache instead of an imported one — and what state a builder holds, so it is clear why a persistent one needs pruning and a disposable one does not.
Show the operational picture: sizing, garbage collection, health and the failure mode when a shared builder dies mid-pipeline, plus the fallback path that keeps jobs building.
Own the tradeoff end to end — measured benefit against isolation, blast radius, cross-tenant cache exposure and who is on call for a stateful build service — and be willing to say the disposable default stays because nobody can own the alternative.
### The two ends of the spectrum **Per-job builder.** Each runner creates a `docker-container` builder, builds, and dies with it. Properties: no shared state, no operator, no coordination, uniform behaviour across the fleet, and a cold cache every time. Any reuse comes from importing an external cache over the network, which costs real seconds on every job. **Long-lived shared builder.** One or a few BuildKit instances run on dedicated hosts; every job points at them through the `remote` driver. Properties: cache is local to the builder, so a hit costs a disk read rather than a transfer; builds start immediately with no builder bootstrap; and the builder is now a production service. ### What the shared model actually buys The honest gain is import cost plus hit quality. On a 47-node fleet building a geospatial tile server as a Node.js API on an alpine base, the cold build is 6 m 41 s; with an imported registry cache it is 1 m 12 s plus about 38 s to fetch a 1.34 GB cache. A warm shared builder skips that 38 s entirely and, because its local cache holds far more history than any exported snapshot, it also *hits more often* — on steps an exported cache never carried. Multiply a minute per build by hundreds of builds a day and the number is large enough to justify the conversation. ### What the shared model costs 1. **It is stateful.** Cache accumulates until something prunes it. BuildKit's own garbage-collection policy, configured in its config file, plus visibility from `docker buildx du` and deliberate `docker buildx prune`, are not optional extras — they are the difference between a builder that works and one that fills a disk at 03:00 and takes every pipeline with it. 2. **The blast radius is the whole fleet.** A per-job builder failing loses one job. A shared builder failing loses all of them at once, so it needs redundancy, health checks and someone on call. Capacity planning becomes real: concurrent builds compete for CPU, memory and disk bandwidth on a machine you now have to size. 3. **Trust boundaries merge.** Everything that builds on a shared builder shares its cache and its build-time cache mounts. A step that writes into a cache mount is writing into storage the next tenant's build can read; a poisoned cache entry can be reused by an unrelated pipeline. If contributions from outside the team can trigger builds, a shared builder is the wrong answer, and the isolation of a disposable one is worth the cold start. 4. **Noisy neighbours.** One repository's enormous build can starve everyone else's, and there is no natural fairness between tenants sharing one BuildKit. ### How I would actually decide - **Measure first.** What fraction of fleet time is builder setup plus cache import? If it is a few per cent, the shared model is buying very little and costing an operational commitment. - **Look at the build shape.** A handful of large, slow, frequently-built repositories concentrates the benefit; a long tail of small services spreads it thin and each still pays the shared builder's operational overhead. - **Draw the trust boundary before the performance line.** Untrusted or externally-triggered builds get isolated disposable builders, full stop. This is not negotiable for a performance gain. - **Prefer pools to a singleton.** Builders pinned per team or per big repository keep most of the cache locality while bounding both blast radius and cross-tenant exposure — and give you a natural unit to size and retire. - **Keep the escape hatch.** Whatever the default, jobs should be able to fall back to a per-job builder when the shared one is unavailable, so a builder outage degrades pipelines rather than stopping them. That fallback is only real if it is exercised, so let some jobs use it routinely. ### The organisational half The technical comparison usually understates the deciding factor: a shared builder needs an owner. If the platform team can own a stateful build service — monitoring, disk, upgrades, capacity, an on-call rotation — the warm-cache gain is available. If nobody can, choosing it anyway means the fleet's throughput now depends on an unowned machine, and the first disk-full incident will cost more than a year of the seconds it saved. Per-job builders are the honest default precisely because their operational cost is near zero, and a team should have to justify leaving it, with measurements, rather than adopt the shared model because it sounds faster.
- What would make you refuse a shared builder outright, whatever the speed gain?Builds triggered by parties I do not trust. A shared BuildKit means a shared cache and shared build-time cache mounts, so one build can leave content another build reuses. Where contributions from outside the team can start a build, the isolation of a disposable per-job builder is the requirement and the cold start is simply the price.
- How do you keep a shared builder from becoming the fleet's single point of failure?Run more than one and pool jobs across them, pin teams or large repositories to specific builders so a failure is partial, health-check them and drain rather than kill, and keep a per-job fallback path that jobs genuinely exercise. Also treat disk as a first-class alarm: a garbage-collection policy plus disk-usage monitoring, because full storage is the most common way one of these dies.
- What single measurement would most change your mind here?The share of total fleet wall-clock spent on builder bootstrap plus cache import. If it is a few per cent, no operational commitment is justified and per-job builders stay. If it is a fifth of the fleet's time on a handful of hot repositories, that is a concrete, bounded case for pooled warm builders serving exactly those repositories.
saying these in an interview costs you the question
- Assumes a shared builder is always faster with no downside
- Ignores disk growth and garbage collection on a long-lived builder
- Treats a build host as stateless when it holds all the cache
- Lets untrusted builds run on a shared builder for speed
- Picks a model without measuring setup and import cost
- Plans no fallback when the shared builder is unavailable