skip to content

Volumes & Storage

Where container state lives and what happens to it when the container is gone: Docker-managed named volumes, bind mounts onto a host path, in-memory tmpfs, and moving data between hosts. Interviewers probe it because the container filesystem is disposable.

part ofDockeroverview, primer and where to startread it →
on this pageshow

questions

20

Walk through the lifecycle of a Docker named volume: how you create one, how you find out where its data physically lives, how it gets attached to a container, and what happens to the data when that container is deleted.

level: juniorimportance: must knowfreq 70%

answer

  1. create / ls / inspect / rm — four verbs
  2. Mountpoint = /var/lib/docker/volumes/<name>/_data
  3. named survives docker rm; anonymous dies with -v
  4. "volume is in use" → a stopped container still holds it
  5. docker system df -v for size

basics

~20 s

Create with docker volume create name, list with docker volume ls, see its host path and driver with docker volume inspect name, attach with -v name:/path. Named volumes outlive containers: deleting the container leaves the volume until you run docker volume rm.

solid answer

~50 s

A named volume is a storage object managed by the Docker daemon, independent of any container. `docker volume create app-data` creates it; `docker volume ls` lists volumes; `docker volume inspect app-data` shows its driver, options, labels and `Mountpoint` (for the built-in `local` driver, typically `/var/lib/docker/volumes/app-data/_data`). You attach it at run time with `-v app-data:/var/lib/postgresql/data` or the more explicit `--mount type=volume,src=app-data,dst=/var/lib/postgresql/data`; if it does not exist, Docker creates it implicitly. Lifecycle is the key point: the volume is **not** owned by the container. `docker rm` (even `docker rm -v`) does not delete a named volume — only anonymous ones. The data survives container removal, image upgrades and daemon restarts, which is exactly why databases use volumes. You delete it explicitly with `docker volume rm app-data`, and Docker refuses while any container (even a stopped one) still references it.

code

bash · 11 lines
bash
docker volume create app-data
docker volume inspect app-data --format '{{.Driver}} {{.Mountpoint}}'

docker run -d --name db \
  --mount type=volume,src=app-data,dst=/var/lib/postgresql/data \
  -e POSTGRES_PASSWORD=secret postgres:16

docker rm -f db          # volume and data still there
docker volume ls | grep app-data

docker volume rm app-data   # explicit delete

go deeper

for a junior

Know the four commands (create, ls, inspect, rm) and the headline rule: named volumes outlive containers and must be deleted explicitly.

for a middle

Add the distinctions — named vs anonymous under docker rm -v and --rm, -v vs --mount syntax, why volume is in use appears, and what inspect fields mean.

for a senior

Frame it operationally: volumes are the state boundary in a redeploy, disk creep is the real failure mode, docker system df -v and label conventions are how you keep a host inventory manageable.

for a principal

Talk policy — naming and labelling conventions so every volume traces to an owning service, who is allowed to delete state, and the fact that host-local volumes tie a stateful workload to one machine, which constrains your placement and DR story.

## What a named volume actually is A container's writable layer is throwaway: it is created with the container and destroyed with it. A **volume** is a separate, daemon-managed storage object with its own name and its own lifetime. Containers reference volumes; they do not own them. That single sentence explains almost every behaviour below. ## Creating ``` docker volume create app-data ``` This registers a volume with the default driver (`local`) and no options. You rarely have to run it explicitly: if you start a container with `-v app-data:/data` and no such volume exists, the daemon creates it on the spot with default settings. Explicit creation matters when you want to set something at creation time — labels (`--label app=billing`) or driver options — because those cannot be changed later; you would have to remove and recreate the volume. ## Listing and inspecting ``` docker volume ls docker volume ls -f dangling=true # volumes no container references docker volume inspect app-data ``` `inspect` returns JSON with the fields that matter operationally: - `Driver` — which storage plugin manages it (`local` unless you chose otherwise). - `Options` — the driver options captured at creation. - `Mountpoint` — where the data sits for the `local` driver on Linux: `/var/lib/docker/volumes/app-data/_data`. - `Scope` — `local` (this host) or `global` (a cluster-wide plugin). - `Labels`, `CreatedAt`. Treat `Mountpoint` as *diagnostic*, not as an interface. It is a real host path on Linux and you can read it as root, but on Docker Desktop for macOS/Windows the daemon runs inside a Linux VM, so that path does not exist on your laptop. Tooling that pokes at `/var/lib/docker/volumes` directly breaks the moment the driver changes, so it is not the supported way to move data in and out. ## Attaching to a container Two syntaxes, same underlying object: ``` docker run -d -v app-data:/var/lib/postgresql/data postgres:16 docker run -d --mount type=volume,src=app-data,dst=/var/lib/postgresql/data postgres:16 ``` The short `-v` form is terse and ambiguous (the first field is a volume name if it has no slash, a host path if it does). `--mount` is explicit key/value, fails loudly on typos, and is the only way to pass some options (such as `volume-nocopy` or `readonly` in a self-documenting way). Prefer `--mount` in anything you commit. At container start the daemon mounts the volume's directory over the target path inside the container's mount namespace. Anything the image had at that path is hidden by the mount — though for an *empty* volume Docker first copies the image content into it ("copy-up"), which is a separate behaviour worth knowing. Many containers may mount the same volume simultaneously; Docker does no locking, so concurrent writers need an application that tolerates it (two Postgres instances on one data directory will corrupt it). ## Removal — the part interviews test - `docker stop` / `docker rm` on the container: the volume and its data **survive**. - `docker rm -v` (or `docker run --rm`): removes *anonymous* volumes attached to that container; a **named** volume is untouched. - `docker volume rm app-data`: the explicit delete. It fails with "volume is in use" if any container — including a stopped or created-but-never-started one — still references it. `docker ps -a --filter volume=app-data` finds the holder. - `docker compose down`: removes containers and networks but keeps named volumes; `docker compose down -v` removes the volumes the Compose file declares. - `docker volume prune`: bulk-removes unused volumes; on modern Engine versions that means anonymous ones only unless you pass `--all`. Because deletion is always explicit for named volumes, the usual production failure is not data loss but disk creep: volumes from long-gone containers piling up under `/var/lib/docker/volumes`. `docker system df -v` shows per-volume size and is the right first command when a host fills up. ## Why this matters Volumes are the supported way to keep state across the whole reason containers are attractive: replacing a container is routine (new image tag, new config, crash restart), and every replacement destroys the writable layer. Putting the database directory, upload directory or cache on a named volume decouples "the process I redeploy constantly" from "the bytes I must never lose".

  • Why does `docker volume rm` fail for a volume whose container was stopped hours ago?
    A stopped container still exists as a Docker object and still holds its mount references, so the daemon refuses to delete the volume underneath it. Remove or recreate the container first (`docker rm <id>`), then delete the volume. `docker ps -a --filter volume=<name>` lists every container, running or not, that references it.
  • Can you rename or resize a named volume in place?
    No. There is no `docker volume rename`, and driver options such as size or NFS parameters are fixed at creation. The supported path is to create a new volume with the settings you want and copy the data across with a helper container that mounts both, then repoint the workload and delete the old volume.
  • Is reading `/var/lib/docker/volumes/<name>/_data` from the host a legitimate way to work with volume data?
    It is fine for a quick look on a Linux host as root, but it is not a supported interface. The path only exists for the built-in `local` driver, it does not exist at all on Docker Desktop where the daemon runs inside a VM, and writing there bypasses the daemon's bookkeeping. Use a helper container that mounts the volume instead.

A container is a rented apartment; a named volume is the storage unit down the street. Moving out of the apartment (removing the container) does not empty the storage unit — you have to go cancel that contract yourself.

saying these in an interview costs you the question

  • Saying that removing the container removes its named volume — data loss is the opposite of what happens.
  • Believing `docker run --rm` cleans up named volumes; it only removes anonymous ones.
  • Treating `/var/lib/docker/volumes/...` as the API and building backup scripts around it.
  • Thinking a volume must be created before use — Docker auto-creates it on first reference, which is exactly how stray volumes appear.
  • Assuming `docker compose down` wipes the database volume, or that it never does (it does with `-v`).

context

open as a page

Docker can attach storage to a container as a bind mount, a named volume, or a tmpfs mount. How do these three differ, and when would you reach for each?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A bind mount maps a host path into the container — good for local source code and config. A named volume is Docker-managed storage that outlives containers — good for databases and app data. A tmpfs mount lives only in RAM — good for scratch files and secrets.

open as a page

You bind-mount a host directory into a Docker container, the process inside writes files there, and afterwards those files on the host are owned by root and your normal user cannot delete them. Explain why this happens and what determines the owner recorded on the host.

level: juniorimportance: must knowfreq 70%

basics

~20 s

A bind mount is a passthrough to the same host inodes; Docker translates nothing. The kernel records the numeric UID of the writing process, and containers run as UID 0 by default, so the host sees root-owned files.

open as a page

A Dockerfile contains the instruction `VOLUME /var/lib/data`. What does the Docker daemon do when someone runs a container from that image, and why do hosts running such images accumulate dozens of unnamed volumes?

level: middleimportance: must knowfreq 58%

basics

~20 s

Every container started from that image gets a fresh anonymous volume (a random 64-hex-character name) mounted at /var/lib/data unless the run command supplies its own mount. Each docker run creates another one, and they are only cleaned up by --rm, docker rm -v or a prune — hence the pile-up.

open as a page

A team ships a container image that runs its process as a non-root user and bind-mounts the source tree from each developer's machine. Describe the available strategies for making the container user's numeric ID line up with the host user's, and the tradeoffs of each.

level: middleimportance: must knowfreq 55%

basics

~20 s

Options: pass --user with the host's numeric UID/GID at run time; bake the UID in at build time with a build argument; or start as root, chown the mount in an entrypoint, then drop privileges with gosu. Each trades reproducibility against portability.

open as a page

Why is tarring a running database's Docker volume not a trustworthy backup, and what makes it one?

level: seniorimportance: must knowfreq 58%

basics

~20 s

A tar walk copies files one at a time over minutes while the engine keeps writing, so the archive is torn across instants and the database's on-disk invariants do not hold. Trustworthy copies come from stopping the writer, using the engine's own online backup, or an atomic storage snapshot.

open as a page

How do you read the contents of a Docker named volume when no container is running?

level: juniorimportance: should knowfreq 52%

basics

~20 s

Start a throwaway container with the volume mounted: docker run --rm -v digest-data:/data alpine ls -l /data. A named volume is only reachable through a mount, and the viewing image need not be the application's own image.

open as a page

A Docker named volume holds a service's data directory and you need an off-host backup of it, plus a way to restore that backup onto a different machine. How do you take and restore that backup, and what makes the copy trustworthy?

level: middleimportance: should knowfreq 42%

basics

~20 s

Run a throwaway helper container that mounts the volume read-only plus a host directory, and tar the data out: docker run --rm -v vol:/data:ro -v "$PWD":/backup alpine tar czf /backup/vol.tgz -C /data .. Restore by untarring into a fresh volume the same way. Quiesce or dump the service first — a live database copy is not consistent.

open as a page

You mount a brand-new, empty Docker named volume at `/usr/share/nginx/html`, a path that already contains files baked into the image. What does the container see on the first start, what does it see on the second start, and how would the result differ if you mounted a host directory at that path instead?

level: middleimportance: should knowfreq 48%

basics

~20 s

On the first start Docker copies the image's content at that path into the empty volume ("copy-up"), so the container sees the image files. On the second start the volume is no longer empty, so it is mounted as-is — stale content persists. A host bind mount never copies; it simply hides the image files.

open as a page

A Docker image already contains files at `/app/node_modules`. What does the container see at that path if you mount an empty named volume there, versus if you mount an empty host directory there as a bind mount?

level: middleimportance: should knowfreq 50%

basics

~20 s

An empty named volume is pre-populated: Docker copies the image's files at that path into the volume on first use, so the container still sees them. An empty bind mount is not — it simply hides the image content, and the container sees an empty directory.

open as a page

When would you attach a `tmpfs` mount to a Docker container instead of a volume or a bind mount, and what limitations does a tmpfs mount have?

level: middleimportance: should knowfreq 40%

basics

~20 s

Use tmpfs when data must never touch disk or persist: scratch files, caches, short-lived credentials, and the writable paths a --read-only container still needs. Limits: Linux containers only, not shareable, gone on stop, and it consumes the container's memory — always set a size.

open as a page

The `docker run` command accepts both the `-v`/`--volume` flag and the `--mount` flag for attaching storage. What are the practical differences between the two, and which would you use?

level: middleimportance: should knowfreq 55%

basics

~20 s

-v is a short colon-separated triple; --mount is explicit key=value pairs with type=. The dangerous difference: -v silently creates a missing host directory (empty, root-owned), while --mount type=bind fails fast. --mount is also more readable and exposes more options.

open as a page

A container that runs as a non-root user gets "permission denied" writing to a freshly created Docker named volume mounted at /data, but the same image works fine when the volume is mounted over a path the image already contains. Explain the ownership rules Docker applies when a volume is first mounted, and how to fix it.

level: middleimportance: should knowfreq 40%

basics

~20 s

On first use, Docker copies the image's content at that path into an empty named volume, ownership and modes included. If the path does not exist in the image, Docker creates it root-owned 0755, so a non-root process cannot write.

open as a page

An engineer ran `docker volume prune` on a busy production host to reclaim disk space. Which volumes does that command actually delete, which does it spare, and how would you reclaim volume space safely on such a host?

level: seniorimportance: should knowfreq 45%

basics

~20 s

It deletes volumes no container references. On Docker Engine 23.0+ that means anonymous ones only by default; --all extends it to unused named volumes — the dangerous flag, since a stopped service's database volume counts as unused. Safe practice: inspect with docker volume ls -f dangling=true and docker system df -v, prune by label, back up first.

open as a page

A team wants edits made in a developer's editor to appear instantly inside a running Docker container so the app hot-reloads. How would you set that up, and what problems does that arrangement typically cause?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Bind-mount the project directory into the container and run the framework's watch mode. Then handle the fallout: shadow build output and dependency directories with volumes, expect slow IO and missed file-change events on Docker Desktop, and never ship this setup — production images copy the code in.

open as a page

How would you run a Docker container whose filesystem is immutable except for the few paths that genuinely need to be written, and what does marking an individual mount read-only actually enforce?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Run with --read-only to make the root filesystem immutable, then add back exactly what is needed: sized tmpfs mounts for /tmp and /run, a named volume for real data, and :ro bind mounts for config and certs. A read-only mount blocks writes from that container only — the host and other containers can still change the data.

open as a page

On a RHEL or Fedora host with SELinux in enforcing mode, a container gets "Permission denied" reading a bind-mounted directory even though the files are mode 0777 and owned by the container's user. Explain the cause, and what the `:z` and `:Z` bind-mount options do.

level: seniorimportance: should knowfreq 35%

basics

~20 s

SELinux denies it on top of the file mode: the container runs as a confined type that may only touch files labelled container_file_t. The :z suffix relabels the mount as shared between containers; :Z relabels it private to one container.

open as a page

You own the container developer experience for a team split across macOS with Docker Desktop, native Linux workstations, and Linux CI runners, and bind-mount file ownership behaves differently on each. How would you design a strategy that holds up across all three?

level: principalimportance: should knowfreq 28%

basics

~10 s

Decide first which paths humans must edit; everything else moves to named volumes. For the remaining bind mounts, standardise one identity mechanism, and treat CI as the reference environment because Docker Desktop hides mismatches.

open as a page

Your tar restore into a Docker named volume leaves the app's data directory empty. Why?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

Usually the archive's paths are wrong or the volume is not the one the app reads. An archive made with absolute paths restores into /data/data, and a service recreated with a different volume name reads a fresh empty volume instead of the restored one.

open as a page

Docker's built-in `local` volume driver can be given mount options at creation time, and third-party volume plugins can be installed alongside it. Explain what a volume driver is responsible for, what the `local` driver's options let you do, and how you would decide whether a workload needs a plugin driver instead.

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

A volume driver is the plugin that creates, mounts and removes the storage behind a volume. The built-in local driver stores data on the host and accepts mount(8)-style options (type=nfs, type=tmpfs, bind). A plugin driver adds remote or cluster-scoped storage — worth it only when containers must move between hosts and keep their data.

open as a page