skip to content

A Docker host is out of disk mid-build. How do you reclaim space without disrupting running containers?

level: seniorimportance: must knowfreq 63%

answer

  1. Measure before you delete anything
  2. Sweep in increasing order of risk
  3. Running containers keep their images safe
  4. --volumes destroys data on shared hosts
  5. Unrotated logs escape df accounting

basics

~20 s

Measure first with docker system df -v and du on the data root, then sweep in rising order of risk: stopped containers, dangling images, aged build cache, finally image prune -a with an until= filter.

solid answer

~50 s

Measure before deleting: `docker system df -v` tells you which bucket is large, and `sudo du -sh /var/lib/docker/*` catches what the report misses, usually container log files. Then sweep from safest outward — `docker container prune --filter "until=48h"`, `docker image prune`, `docker builder prune --filter "until=72h"` — checking free space between steps. Only if that is not enough do you run `docker image prune -a --filter "until=168h"`, having confirmed the images you may need for a rollback are still in the registry. Nothing in that sequence can touch a running container: prune never deletes an image a container references, so the checkout API keeps serving at its 340 ms p99. What you must not do is reach for `docker system prune -a --volumes`, which deletes data in unused volumes and every image not currently in use, on a host you share with someone else's stack.

code

bash · 11 lines
bash
# 1. measure
docker system df -v
sudo du -sh /var/lib/docker/*

# 2. sweep, safest first, checking free space between steps
docker container prune -f --filter "until=48h"; df -h /var/lib/docker
docker image prune -f;                          df -h /var/lib/docker
docker builder prune -f --filter "until=72h";   df -h /var/lib/docker

# 3. only if still short, and only with an age filter
docker image prune -a -f --filter "until=168h"

go deeper

for a junior

Know the individual commands and that the safe first moves are docker container prune and docker image prune. Be able to say that prune never removes an image a running container is using.

for a middle

Explain what each prune targets and how the flags widen it, especially that -a covers images with no container and --volumes covers data. Know the ordering effect: removing old containers first frees their images for a later sweep.

for a senior

Show the diagnostic habit — measure with docker system df -v and du, reconcile the gap, then sweep in increasing risk order — and name the traps: held-open deleted files, unrotated logs, and a growing writable layer that no prune can reclaim.

for a principal

Frame the incident as a systems failure, not a cleanup: decide who is allowed to run destructive prunes on shared hosts, whether build and run workloads should share a disk at all, and what alert would have fired before the host wedged.

## Step 1 — find out what is actually full Deleting before measuring is how a cleanup turns into an incident. Two commands, always in this order: ``` docker system df -v sudo du -sh /var/lib/docker/* ``` The first attributes space to images, containers, local volumes and build cache, and itemises each. The second catches everything the daemon does not account for. If `du` says 213 GB and the report only explains 148 GB, the gap is real and it is very often container log files under the data root — a service logging to stdout with no rotation configured on its log driver. That case is fixed by configuring rotation, not by pruning anything, and a candidate who prunes for twenty minutes without noticing is missing the actual bug. Worth knowing at this point: freeing a file that a process still holds open does not return the space. If someone already deleted a huge log out from under the daemon, `df` stays full until the holder closes the descriptor; `lsof +L1` finds those. ## Step 2 — sweep in increasing order of risk Run these one at a time and re-check free space in between, because the first cheap sweep is often enough: 1. **`docker container prune --filter "until=48h"`** — removes stopped containers older than two days along with their writable layers. Running containers are untouchable here by definition. This also *unlocks* their images for a later image sweep, which is why the ordering matters. 2. **`docker image prune`** — dangling images only. Cannot delete anything tagged, cannot delete anything a container references. 3. **`docker builder prune --filter "until=72h"`** — BuildKit cache records not used in three days. On a build host this is frequently the largest single bucket; on a host that only runs containers it is empty. Note it prunes the cache of the current builder; a buildx builder using a container driver holds its cache inside that builder, so you prune it with `docker buildx prune --builder <name>`. 4. **`docker image prune -a --filter "until=168h"`** — the aggressive step, deliberately last and deliberately filtered. It removes every image older than a week that no container references. Before running it, confirm the images you would need for a rollback are pushed and re-pullable, because this is what deletes them. At no point in that ladder can a running container be affected. An image referenced by a running container is never eligible for prune, so an order-checkout API container serving at a 340 ms p99 keeps its image, its writable layer and its volumes throughout. That property is what makes the ladder safe to run during an incident. ## Step 3 — the things that bite **`docker system prune -a --volumes` is not a cleanup, it is a reset.** Plain `docker system prune` already removes stopped containers, networks with nothing attached, dangling images and unused build cache. `-a` extends the image sweep to every image no container uses, and `--volumes` extends it to unused local volumes — that is *data*. On a shared host, the volume of a database container that someone stopped an hour ago is unused by that definition, and it is not recoverable. The rule is simple: on a host you do not exclusively own, do not pass `--volumes`, and prefer the individual prunes so you can see what each one takes. **A running container's writable layer only grows.** If `docker ps -s` shows a container at 47 GB, no prune will help: that space is freed when the container is removed. The fix is to find what it is writing — an application logging to a file path inside the container, a cache directory that was never mounted — and either mount that path as a volume or stop writing there. **Pruning the build cache has a cost you should name out loud.** Deleting BuildKit's cache is completely safe for correctness and directly expensive for build time: the next build of a large Python inference image with CUDA wheels runs cold, re-downloading and re-compiling work that a cache hit would have skipped. During an out-of-disk incident that is the right trade, but on a schedule it is a decision about CI minutes, not just disk. **Deleting from the host filesystem is not a shortcut.** Removing directories under the data root by hand, rather than through the daemon, leaves the daemon's metadata pointing at layers that no longer exist and can corrupt the installation. Every reclaim goes through the API. ## Step 4 — leave the host better than you found it The incident answer ends with a prevention answer: configure log rotation on the log driver so the invisible bucket stops growing, put an age-filtered prune on a timer instead of waiting for the next page, and monitor free space with an alert threshold that fires while there is still room to act.

  • What exactly does a plain `docker system prune` remove, before you add any flags?
    Stopped containers, networks with no container attached, dangling images, and unused build cache. It does not touch volumes, and it does not remove tagged images that are simply not running. Adding `-a` widens the image sweep to every image no container references; adding `--volumes` brings unused local volumes — real data — into scope. Those two flags are the whole difference between a tidy-up and a data-loss event.
  • You pruned everything you safely could and `df` barely moved. Where else can the space be?
    Outside the daemon's accounting: container log files under the data root with no rotation configured, files a process still holds open after deletion, bind-mounted host directories that containers are filling, or a fat writable layer in a container that is still running and therefore cannot be pruned. `du` on the data root and `lsof +L1` distinguish these; each has a different fix, and none of them is another prune.
  • Why is `docker builder prune` the step you often reach for on a CI runner but never on a service host?
    Because build cache only accumulates where builds happen. A host that just runs containers has an essentially empty Build Cache row in `docker system df`, so pruning it frees nothing. On a runner it is regularly the biggest bucket. The cost is symmetric: on the runner you pay for it with slower cold builds afterwards, which is why an `until=` filter beats deleting the cache wholesale.

saying these in an interview costs you the question

  • Runs `docker system prune -a --volumes` as the first move
  • Deletes directories under the data root by hand
  • Prunes before measuring which bucket is actually large
  • Ignores container log files because df does not show them
  • Thinks pruning can kill a running container's image
  • Expects prune to shrink a running container's writable layer

context