skip to content

How would you design a Docker disk-reclamation policy for a shared build fleet?

level: principalimportance: should knowfreq 42%

answer

  1. Two host classes, two different policies
  2. Continuous timers beat emergency cleanups
  3. Age windows plus label exclusions
  4. Budget the build cache, never wipe it
  5. A nearby mirror makes images cheap to lose

basics

~10 s

Give build runners and service hosts separate rules, reclaim continuously with age and label filters instead of during incidents, budget the build cache rather than wiping it, and keep --volumes out of automation.

solid answer

~50 s

Start by deciding what each host class is for, because build runners and service hosts have opposite economics: on a runner the build cache is the asset and images are disposable, on a service host the images are the asset and there is no cache. Then make reclamation continuous and boring — a timer running age-filtered `docker container prune`, `docker image prune`, and `docker builder prune`, plus label-based exclusions so pinned base images survive. Give the build cache a space budget rather than a scheduled wipe, so builds stay warm until the disk actually needs the room. Make images cheap to lose by treating the registry as the source of truth with immutable tags and a pull-through mirror, so an over-eager sweep costs a pull, not an outage. Finally, monitor free space and alert on the trend, and never put `--volumes` in an automated job on shared hardware.

code

bash · 8 lines
bash
#!/usr/bin/env bash
# scheduled reclamation for a build runner: safest first, always filtered
set -euo pipefail
docker container prune -f --filter "until=48h"
docker image prune -f
docker image prune -a -f --filter "until=336h" --filter "label!=retention"
docker builder prune -f --filter "until=168h"
docker system df

go deeper

for a junior

Know that reclamation should be scheduled and filtered rather than typed by hand during an outage, and that the same commands are not appropriate on a host running production services.

for a middle

Be able to write the scheduled job: individual prunes rather than a blanket one, until= windows, and label exclusions so pinned images survive. Explain why each step comes in that order.

for a senior

Show that you separate build and service hosts, budget the build cache instead of wiping it, and make images cheap to lose by pushing to a registry with a nearby mirror before pruning aggressively.

for a principal

Own the tradeoff explicitly: retention windows derived from rollback horizon and cache reuse, disk spend weighed against CI minutes, guardrails on destructive flags with a named owner, and alerting on trend so the policy is never exercised under pressure.

## Decide what the host is for before you decide what to delete The first architectural move is refusing to write one policy. Two host classes behave in opposite ways: - **Build runners.** The build cache is the valuable object — it is what keeps a large Python inference image with CUDA wheels building in three minutes instead of half an hour. Images produced locally are disposable the moment they are pushed. Containers are short-lived. Volumes are usually irrelevant. - **Service hosts.** There is no build cache. Images are the valuable object, because deleting the wrong one means an outage-length re-pull at the worst moment, and volumes hold state that must never be swept. Running both classes on the same machine is the root cause of most disk incidents I would design away: a build filling the disk should never be able to threaten an order-checkout API on a 340 ms p99 budget. If they must share, they at least should not share a filesystem. ## Make reclamation continuous, not heroic A policy that only runs when someone is paged is not a policy. The shape that works is a timer, several times a day, running the individually-scoped prunes with age filters, in increasing order of risk, non-interactively: ``` docker container prune -f --filter "until=48h" docker image prune -f docker image prune -a -f --filter "until=336h" --filter "label!=retention" docker builder prune -f --filter "until=168h" ``` Three design choices are embedded there. First, individual commands rather than `docker system prune`, so each step's blast radius is visible and one of them can be dropped per host class. Second, **age filters everywhere** — `until=` is what converts a destructive command into a retention rule, and the numbers are policy: two days of stopped containers, two weeks of images, a week of unused cache. Third, **label-based exclusion**: builds stamp `LABEL retention=keep` on the images that must never be re-pulled at 3 a.m. — the heavy base images and the current release — and `--filter "label!=retention"` spares them. Labels are how you express intent that an age filter cannot. ## Budget the build cache instead of wiping it Deleting build cache is free of correctness risk and expensive in CI minutes, which makes it a capacity decision rather than a cleanup. The better instrument than a scheduled wipe is a size budget: give the cache an allowance and let the daemon's own garbage collection evict least-recently-used records to stay inside it. `docker builder prune` also accepts a space-keeping flag for the same purpose when you drive it yourself. Either way, size the allowance from what you are buying: measure build duration with a warm cache versus cold, cost that in runner-minutes, and set the budget where the marginal gigabyte stops paying for itself. A team that cannot state that number has picked its retention by superstition. ## Make images cheap to lose The reason aggressive image pruning is frightening is usually that the local daemon has become the only copy of something. Fix the cause rather than the symptom: immutable, digest-addressable tags pushed to a registry for anything that could ever need a rollback, and a pull-through mirror close to the fleet so that re-pulling a 6.2 GB base costs seconds of LAN traffic instead of minutes of internet. Once that is true, `image prune -a` degrades from a hazard into a slow build, and the policy can be far more aggressive on space. ## Guardrails, ownership and blast radius A few rules I would make non-negotiable across the fleet: - **No `--volumes` in any automated job**, on any host that is shared or holds state. Volume reclamation is a human decision with a named owner, because the registry cannot give data back. - **No hand-deletion under the data root.** Every reclaim goes through the API, or the daemon's metadata and the filesystem disagree. - **The same policy is configuration, not tribal knowledge** — it ships with the host image, so a new runner is born with it rather than acquiring it after its first incident. ## Close the loop with signals Finally, export the buckets — the four rows of `docker system df` — as metrics per host alongside filesystem free space, and alert on the trajectory rather than a fixed threshold: a runner that will hit 90% in four hours is actionable, a runner that is at 78% is not. The success criterion for the whole policy is that nobody ever runs a prune during an incident, because the timer already did, and that the failure mode when the policy is too aggressive is a slower build rather than a service that cannot start.

  • How do you choose the retention window for `--filter until=` rather than guessing at it?
    Derive it from what the window protects. For images it is the rollback horizon: how far back a release could realistically be reverted, plus a margin, so two weeks is a defensible starting point. For build cache it is the reuse distribution — how old a cache record typically is when a build still hits it. Both are measurable from deploy history and build logs, and both should be revisited when either changes.
  • Would you ever run a fleet-wide `docker system prune -a --volumes` from configuration management?
    Not on shared or stateful hosts. It removes unused local volumes, and unused only means no container is attached right now, which is true of a stopped database's data. On disposable, single-purpose build runners that hold nothing but caches it is defensible, because the host is rebuildable by definition — and that is the real distinction: the command is acceptable exactly where the host itself is expendable.
  • What signal tells you the policy is too aggressive rather than too lax?
    Build duration and pull volume, not disk. If median build time climbs because cache records are being evicted before they are reused, or registry pull bandwidth spikes on every deploy because base images keep being re-fetched, the retention windows are too short. A policy that is too lax shows up instead as free space trending toward a cliff. Tracking both means the two failure modes are distinguishable before either becomes an incident.

saying these in an interview costs you the question

  • One prune policy applied to build and service hosts alike
  • Schedules `docker system prune -a --volumes` fleet-wide
  • Treats build cache deletion as free because nothing breaks
  • Relies on engineers pruning manually when paged
  • Sets retention windows with no rollback or reuse data
  • Alerts on a fixed disk threshold rather than the trend

context