A Dockerfile contains the instruction `VOLUME /var/lib/data`. What does the Docker daemon do when someone runs a container from that image, and why do hosts running such images accumulate dozens of unnamed volumes?
answer
- VOLUME = metadata, applied at container create
- random 64-hex name = anonymous volume
- one per `docker run`, never reused
- cleanup: --rm, rm -v, compose down, prune
- writes after VOLUME are lost at build time
basics
~20 sEvery container started from that image gets a fresh anonymous volume (a random 64-hex-character name) mounted at /var/lib/data unless the run command supplies its own mount. Each docker run creates another one, and they are only cleaned up by --rm, docker rm -v or a prune — hence the pile-up.
solid answer
~60 s`VOLUME` is a declaration in image metadata, not a mount at build time. At container **creation**, for every declared path that the run command has not already covered with an explicit mount, the daemon creates an anonymous volume — a real volume with a random 64-hex-char ID and no human name — and mounts it there. Run the image ten times and you get ten volumes. They are cleaned up only by `docker run --rm`, `docker rm -v`, `docker compose down`, or `docker volume prune` (which on modern Engine versions targets exactly these anonymous volumes by default). Plain `docker rm` leaves them behind, which is how a CI host ends up with hundreds of dangling volumes and a full disk. The second, nastier effect is build-time: once `VOLUME /var/lib/data` is declared, any later `RUN` that writes under that path writes into a temporary volume that is discarded when the layer finishes — so the files silently vanish from the image. That is why most modern images drop `VOLUME` and let operators choose the mount.
code
dockerfile · 5 linesFROM debian:12
VOLUME /var/lib/data
# BROKEN: this writes into a throwaway volume, not into the layer
RUN mkdir -p /var/lib/data && dd if=/dev/zero of=/var/lib/data/seed.bin bs=1M count=50
# resulting image has an empty /var/lib/datago deeper
Recall that VOLUME makes Docker attach an unnamed volume at container start, and that these accumulate until pruned.
Explain the mechanics: metadata applied at create time, skipped when the run command mounts that path, one volume per run, plus the four cleanup commands.
Bring the operational angle — CI hosts filling with hex-named volumes, docker system df as the detector, --rm as the standing habit, and the build-time data-loss trap when a RUN follows the declaration.
Argue the policy: image authors should not decide persistence for operators; declare nothing, document the state paths, and require named, labelled volumes so every byte of state has an owner and a backup story.
## What `VOLUME` really is `VOLUME ["/var/lib/data"]` in a Dockerfile writes an entry into the image's config metadata (`Config.Volumes`). It does not create storage while building. Its effect is entirely at container-creation time: the daemon reads that list, and for each path **not already covered by an explicit `-v`/`--mount` from the run command**, it creates a fresh volume and mounts it there. ## Anonymous volumes The volume Docker creates in that case has no name you chose. It gets a random 64-character hexadecimal ID and appears in `docker volume ls` as exactly that: ``` local 9f2c1a...c7e4 ``` Internally it is an ordinary volume with the `local` driver — same mountpoint layout, same inspect output. What distinguishes it is a piece of bookkeeping: because you never gave it a name, Docker marks it as anonymous, and that flag is what cleanup commands key on. The original intent was helpful: an image author signalling "this path holds state, do not leave it on the container's throwaway writable layer." The container gets working persistence even if the operator forgets to mount something. In practice the cost outweighs the benefit, which is why images like the official Postgres still carry it but most application images do not. ## Why hosts fill up Each `docker run` creates a *new* anonymous volume — there is no reuse across containers, because nothing identifies one anonymous volume as "the same state" as another. So the multiplication is linear in container starts: - CI that runs an integration test image 200 times a day creates 200 volumes a day. - A crash-looping container recreated by an orchestrator creates one per recreation (restarts of the *same* container reuse its volume; recreation does not). And plain `docker rm mycontainer` does **not** delete them. The cleanup paths are: - `docker run --rm ...` — removes the container and its anonymous volumes on exit. The habit to teach for CI. - `docker rm -v mycontainer` — the `-v` flag removes anonymous volumes attached to that container (never named ones). - `docker compose down` — removes anonymous volumes belonging to the project's containers; `-v` additionally removes the named volumes the file declares. - `docker volume prune` — on Docker Engine 23.0 and later this removes **only anonymous** unused volumes by default, precisely because that is the safe bulk cleanup; `--all` extends it to unused named volumes too. Symptom to recognise: `docker system df` shows a large VOLUMES row with a high count and near-100% reclaimable, and `docker volume ls` is a wall of hex. ## The build-time trap This is the part that separates people who have read the docs from people who have been bitten. Consider: ```dockerfile VOLUME /var/lib/data RUN mkdir -p /var/lib/data && ./seed-database.sh # writes 200 MB of seed data ``` During that `RUN`, the builder honours the declared volume by mounting a temporary volume at `/var/lib/data`. The script writes into that temporary volume, the layer is committed from the container's filesystem changes — and the volume's contents are **not** part of the layer. The image ships with an empty directory, the seed data is gone, and the error message is nothing at all. The rule: never write to a path after declaring it as a `VOLUME`; put `VOLUME` last if you keep it at all. The same metadata is also sticky through inheritance: a child image `FROM` a parent that declared a volume inherits the declaration, and there is no `UNVOLUME` instruction to take it back. You can only avoid the anonymous volume by always mounting something at that path yourself (`-v app-data:/var/lib/data` or a bind mount), which is the practical mitigation when you must consume such a base image. ## What to do instead Modern practice is to omit `VOLUME` from application Dockerfiles and document the persistent paths instead — in the README, in the Compose file, in the Helm chart. Persistence then becomes an explicit operator decision with a name attached, which means it can be inspected, labelled, backed up and deleted deliberately. `docker inspect --format '{{json .Config.Volumes}}' <image>` tells you whether an image you are about to adopt carries the declaration.
- How do you stop an image whose Dockerfile declares `VOLUME /var/lib/data` from creating an anonymous volume on every run?Cover the declared path with your own mount at run time: `--mount type=volume,src=app-data,dst=/var/lib/data` or a bind mount. Because the run command already supplies a mount for that path, the daemon skips creating an anonymous one. There is no way to remove the declaration from the image short of rebuilding it, since Dockerfile has no `UNVOLUME` instruction.
- Two containers are started from the same image that declares a `VOLUME`. Do they share state?No. Each container gets its own freshly created anonymous volume, so they see independent empty directories. Sharing requires an explicit named volume mounted into both containers at the same path.
- A colleague says `docker rm` cleans up after itself. What do you correct?`docker rm` removes the container and its writable layer but leaves anonymous volumes behind; you need `docker rm -v`, or `--rm` at run time. Named volumes are never removed by either — only `docker volume rm`, `docker volume prune --all` or `docker compose down -v` touch those.
saying these in an interview costs you the question
- Claiming `VOLUME` creates a directory or copies data at build time — it is metadata consumed at container creation.
- Thinking all containers from the image share one anonymous volume; each run gets its own.
- Assuming `docker rm` cleans anonymous volumes (it needs `-v`), or that stopping the container is enough.
- Writing files into a path after declaring it with `VOLUME` and expecting them in the image.
- Looking for an `UNVOLUME` instruction to cancel an inherited declaration.