skip to content

A Docker named volume holds a service's data directory and you need an off-host backup of it, plus a way to restore that backup onto a different machine. How do you take and restore that backup, and what makes the copy trustworthy?

level: middleimportance: should knowfreq 42%

answer

  1. helper container: -v vol:/data:ro + -v $PWD:/backup + tar
  2. -C /data . → relative paths in the archive
  3. restore into a fresh, never-started volume
  4. quiesce or pg_dump — live tar is torn
  5. --numeric-owner, -p; untested restore = no backup

basics

~20 s

Run a throwaway helper container that mounts the volume read-only plus a host directory, and tar the data out: docker run --rm -v vol:/data:ro -v "$PWD":/backup alpine tar czf /backup/vol.tgz -C /data .. Restore by untarring into a fresh volume the same way. Quiesce or dump the service first — a live database copy is not consistent.

solid answer

~50 s

There is no `docker volume backup`, so the idiom is a **helper container**: a short-lived container that mounts the volume and a bind mount to somewhere you can reach. ``` docker run --rm -v app-data:/data:ro -v "$PWD":/backup alpine \ tar czf /backup/app-data.tgz -C /data . ``` Restore is the mirror image into a new volume on the target host, which copy-up leaves alone because the tar makes it non-empty: ``` docker run --rm -v app-data:/data -v "$PWD":/backup alpine \ sh -c 'cd /data && tar xzf /backup/app-data.tgz' ``` Trustworthiness comes from three things. **Consistency**: stop the writer, or better, use the application's own dump (`pg_dump`) — tarring a live database directory gives you a torn copy. **Fidelity**: preserve numeric uid/gid and permissions (`tar --numeric-owner -p`) so the service can still read its files. **Verification**: restore into a scratch volume and start the service against it; an untested backup is a hypothesis.

code

bash · 12 lines
bash
# backup (source mounted read-only)
docker run --rm -v app-data:/data:ro -v "$PWD":/backup alpine \
  tar czf /backup/app-data-$(date +%F).tgz -C /data .

# restore onto another host, into a fresh volume
docker volume create app-data
docker run --rm -v app-data:/data -v "$PWD":/backup alpine \
  sh -c 'cd /data && tar xzpf /backup/app-data-2026-08-12.tgz'

# migrate one volume to another (e.g. new driver options)
docker run --rm -v old-vol:/from:ro -v new-vol:/to alpine \
  sh -c 'cp -a /from/. /to/'

go deeper

for a junior

Show the two helper-container commands (tar out, tar in) and know the volume must be mounted into a container to be read portably.

for a middle

Add the details that make it correct: read-only source mount, -C /data . for relative paths, restore into a fresh volume, and preserving ownership and permissions.

for a senior

Lead with consistency — quiesce or use the engine's dump — then talk verification, off-host storage, encryption, and label-driven enumeration of what to back up.

for a principal

Frame it as a data-durability policy: RPO/RTO targets decide dump-vs-snapshot-vs-file-copy, restore drills are scheduled not aspirational, and host-local volumes without a shipped-off-host copy are an accepted single point of failure that must be named.

## Why a helper container at all A named volume with the built-in `local` driver does live on the host at `/var/lib/docker/volumes/<name>/_data`, so root can tar it directly — but that path is an implementation detail. It does not exist for plugin drivers, and on Docker Desktop the daemon runs inside a Linux VM, so nothing on your Mac or Windows filesystem corresponds to it. The portable interface is the one Docker gives every container: mount the volume. A **helper container** is a container whose only job is to move bytes. It mounts the volume at one path, a host directory (bind mount) at another, and runs an archiver. It uses `--rm` so it evaporates when done, and a tiny image (`alpine`, `busybox`) so it costs nothing to start. ``` docker run --rm \ -v app-data:/data:ro \ -v "$PWD":/backup \ alpine tar czf /backup/app-data-$(date +%F).tgz -C /data . ``` Points of craft in that one line: `:ro` on the source so a mistake cannot damage the live data; `-C /data .` so the archive holds relative paths (`./file`) and unpacks cleanly anywhere instead of carrying an absolute `/data/` prefix; the date in the filename so backups do not overwrite each other. ## Restore ``` docker volume create app-data docker run --rm -v app-data:/data -v "$PWD":/backup alpine \ sh -c 'cd /data && tar xzf /backup/app-data-2026-08-12.tgz' ``` On a different host, ship the tarball there first (scp, object storage) and run the same command. Note the interaction with copy-up: if you instead started the service with an empty volume, Docker would seed the volume from the image and the volume would no longer be empty — restoring on top of that mixes two datasets. Restore into a **freshly created, untouched** volume before the service ever starts. ## Consistency: the part that actually decides whether the backup works Copying files out from under a running process gives you a *crash-consistent at best, torn at worst* image: the archiver walks the tree over seconds while the process rewrites pages, so different files come from different moments. In order of preference: 1. **Application-native dump.** `pg_dump`, `mysqldump`, `redis-cli --rdb`, `mongodump`. These produce a logically consistent snapshot from a running service and are usually smaller and version-portable. For a database this is the right answer and file-level backup is the fallback. 2. **Quiesce, then copy.** `docker stop svc`, run the helper, `docker start svc`. Simple and correct; costs downtime. 3. **Filesystem or storage snapshot** (LVM, ZFS, EBS) then tar the snapshot. Keeps downtime near zero but only reaches crash consistency unless the application flushes first. If you must copy live, say so out loud in the interview and pair it with a restore test — for many engines a crash-consistent copy does recover, because they replay their write-ahead log, but that is a property of the engine, not of your backup. ## Fidelity The helper runs as root by default, which is what lets it read files owned by the service's uid. Preserve that ownership on the way back: - `tar --numeric-owner` records raw uids/gids rather than names, which matter because `postgres` may be uid 999 in one image and 70 in another. - `tar -p` (`--preserve-permissions`) on extract keeps the modes; a Postgres data directory refuses to start unless it is mode 0700 and owned by the right uid. - Symlinks, hard links and sparse files: plain `tar` handles the first two; add `-S` if the data set is sparse. - Extended attributes and ACLs need `--xattrs --acls` and a tar build that supports them; GNU tar in a Debian-based helper is safer here than BusyBox tar in Alpine. BusyBox tar (what `alpine` ships) is fine for ordinary trees; reach for `debian:stable-slim` when you need the GNU flags. ## Making it operational - **Verify**: restore into a scratch volume, start the real image against it, run a health query. Schedule this, do not do it once. - **Encrypt and ship off-host**: pipe through `gpg` or upload to object storage from the helper. A backup on the same disk as the data is not a backup. - **Label your volumes** (`docker volume create --label backup=daily app-data`) so a script can enumerate what to back up via `docker volume ls -f label=backup=daily` instead of a hand-maintained list. - **Record the shape**: which service, which image tag, which volume, whether it was quiesced. Restores happen under stress and the person doing it may not be you. ## The reflex to build Any time you need to read, seed, migrate or resize volume data, the same helper-container pattern applies — mount the volume, mount whatever else you need, run a one-shot tool, `--rm`. Copying between two volumes is the same trick with both mounted: `docker run --rm -v old:/from:ro -v new:/to alpine sh -c 'cp -a /from/. /to/'`.

  • Why not just `tar` the volume's host mountpoint directly as root?
    It works only for the built-in `local` driver on a Linux host: plugin drivers may keep data on a remote system with no local path, and on Docker Desktop the mountpoint lives inside the daemon's VM and is unreachable from the host. Reading through a helper container is the interface every driver supports, so the same script keeps working when the storage changes underneath.
  • You restore a Postgres volume and the container refuses to start with a permissions error. What went wrong?
    The archive was extracted without preserving ownership and mode, so the data directory is root-owned or too permissive, and Postgres requires it to be owned by its own uid and to be mode 0700. Re-extract with `tar -xpf --numeric-owner`, or fix it with a helper container running `chown -R 999:999 /data && chmod 700 /data` using the uid the image actually uses.

saying these in an interview costs you the question

  • Tarring a running database's data directory and calling it a backup without mentioning consistency at all.
  • Assuming a `docker volume backup` command exists.
  • Depending on `/var/lib/docker/volumes/<name>/_data` in scripts that must also work on Docker Desktop or with plugin drivers.
  • Restoring into a volume the service has already started against, so restored data mixes with what is already there.
  • Extracting without preserving uid/gid and permissions, then being surprised the service cannot read its own files.

context