skip to content

questions

3

Why is tarring a running database's Docker volume not a trustworthy backup, and what makes it one?

level: seniorimportance: must knowfreq 58%

answer

  1. The copy takes minutes; so does the damage
  2. First file and last file, different instants
  3. Crash-consistent is not the same as torn
  4. Stop it, dump it, or snapshot it atomically
  5. A backup nobody restored is not a backup

basics

~20 s

A tar walk copies files one at a time over minutes while the engine keeps writing, so the archive is torn across instants and the database's on-disk invariants do not hold. Trustworthy copies come from stopping the writer, using the engine's own online backup, or an atomic storage snapshot.

solid answer

~60 s

Copying a live data directory — with `tar` from a helper container or with `docker cp`, it makes no difference — is a sequential read of a tree that is changing underneath it. File A is captured at 03:00:01 and file Z at 03:04:12, and a database's consistency rules span those files, so the result is neither a point-in-time image nor something the engine promises to recover from. There are three honest options. Stop the container so the engine shuts down cleanly and flushes, then tar the now-static volume; that is the simplest and costs downtime. Or use the database engine's own online backup or dump tool, which knows how to produce a consistent copy while serving traffic. Or take an atomic snapshot at the storage layer and let the engine replay its log on restore — crash-consistent rather than clean. Docker itself offers no volume snapshot or quiesce command. Whichever you pick, a backup is only proven once you have restored it into a fresh volume and started the service against it.

code

bash · 4 lines
bash
docker stop -t 60 digest
docker run --rm -v digest-data:/data:ro -v "$PWD":/backup alpine \
  tar czf /backup/digest-$(date +%Y%m%dT%H%M).tgz -C /data .
docker start digest

go deeper

for a junior

Know that copying files out of a volume while a database is writing to it produces an unreliable copy, and that stopping the container first is the simple fix worth naming.

for a middle

Explain why a sequential copy spans instants and how that differs from a crash-consistent snapshot, and describe what docker stop actually causes the engine to do before the volume goes still.

for a senior

Demonstrate a defensible procedure end to end — quiesce choice with its downtime cost, the copy, restore into a fresh volume, and a scheduled verification — and be able to say why docker pause is not a shortcut.

for a principal

Own the policy: which datasets get engine-native online backups versus stop-and-copy, what recovery point and recovery time each buys, where snapshot atomicity actually comes from, and how restores are rehearsed rather than assumed.

## The defect: a copy that spans instants Backing up a Docker named volume means reading its files. The standard idiom mounts the volume into a throwaway container next to a bind mount of a host directory and runs `tar` between them. That is a perfectly good way to move bytes, and it is exactly the wrong thing to point at a database that is running. `tar` walks the tree. On a 412 MB volume it might take four minutes. The first file is read at one instant and the last at another, and between those two instants the database engine has been writing: appending to its log, checkpointing pages, renaming files, deleting others. The archive is a mosaic of different moments — often called a *torn* or *fuzzy* copy. Nothing in it is guaranteed to satisfy the engine's own on-disk invariants, because those invariants span files. A concrete flavour of the problem: an engine that keeps a write-ahead log alongside its main data file — SQLite in WAL mode keeps a `-wal` sidecar, and most server engines have an equivalent — can easily have the data file captured before a checkpoint and the log captured after it, producing a pair that never coexisted. Two clarifications matter because candidates often conflate them: * **Crash consistency** is what you get from an *atomic* copy taken at a single instant — a filesystem or block-device snapshot. It looks to the engine exactly like a power cut: recoverable, because every serious engine is built to survive that. * **A sequential file-by-file copy of a live tree gives you neither** a clean shutdown nor crash consistency. It is a state the machine was never actually in, so no recovery procedure is promised to fix it. `docker cp` of a live data directory has precisely the same defect. It is a different command, not a different guarantee; it too reads the tree file by file while the writer runs. The reason "just `docker cp` the data dir" fails a backup review is consistency, not convenience. ## Quiescing: the three honest options **1. Stop the writer.** `docker stop <container>` sends SIGTERM, the engine performs its shutdown sequence — flush buffers, checkpoint, close cleanly — and after the container exits the volume is static. Now tar it from a helper container; nothing is changing, so the archive is a faithful copy of a clean shutdown state, the easiest thing in the world to restore. The cost is downtime: for the email-digest builder that is the stop plus the copy plus a six-second cold start on the way back. Give the engine a stop timeout long enough to finish (`docker stop -t 60`), or SIGKILL after ten seconds turns your clean shutdown back into a crash. **2. Use the engine's own online backup.** Databases ship tools that produce a consistent copy while serving traffic, because they can coordinate with the internals — take a checkpoint, note a log position, stream a copy, note the end position. Run that tool and back up *its output* rather than the raw data directory. This is the standard answer for anything that cannot take downtime, and it moves the correctness burden onto the engine's authors, where it belongs. **3. Snapshot atomically at the storage layer.** A filesystem or cloud-disk snapshot captures the whole volume at one instant, giving crash consistency; the engine then recovers on first start, and you copy the snapshot at leisure without holding the service still. Note what this requires: the volume's data must live on a storage layer that can do that. Docker's own tooling cannot — there is no `docker volume snapshot` — so this is a property of the host filesystem, the block device, or a network-backed volume, not of Docker. ## What `docker pause` does and does not buy `docker pause` freezes the container's processes with the kernel's cgroup freezer. Files do stop changing, so the tar is no longer torn — but the on-disk state is still whatever the engine happened to have written when it was frozen, with any in-memory state unflushed. That is crash consistency at best, never a clean shutdown, and the pause lasts the entire copy, so it is an outage that also produces a weaker artifact than simply stopping the container. It is not a shortcut. ## Restore is part of the design Two properties separate a backup from a directory of tarballs. * **Restore into a fresh, empty volume.** Extracting over a volume that already holds data leaves a mix of old and new files — an engine can fail in inventive ways when its data directory contains files from two eras. Create a new volume, restore into it, point a container at it. * **Prove it.** Restore into that fresh volume on a scratch host, start the service against it, and check that it comes up without recovery errors and that the data is the size and shape you expect. Do it on a schedule, not during the incident. An unverified backup is an untested code path, and the failures described above are all silent — the tar exits 0 every time. ## Choosing under pressure If downtime is cheap, stop the container: it is the least machinery and the strongest guarantee. If it is not, use the engine's online backup tool. Reach for storage snapshots when the volume is large enough that copying it is the bottleneck, and accept that you are then relying on the storage layer's atomicity claim. Reserve the plain live tar for volumes holding data that is not a database at all — static assets, generated reports, an append-only spool — where a torn copy has no invariants to break.

  • Does pausing the container with `docker pause` before the copy make a live tar safe?
    It stops the files changing under tar, so the copy is no longer torn, but the on-disk state is only whatever the engine had already written — crash consistency at best, with in-memory state lost. The freeze also lasts the whole copy, so you pay an outage and still get a weaker artifact than `docker stop` would have given you.
  • Why restore into a fresh empty volume rather than over the existing one?
    Extracting over a populated volume merges two generations of files: anything present in the old data directory but absent from the archive survives, and an engine that finds stale files beside restored ones can fail in confusing ways. A new volume makes the restored state the only state, and it leaves the original intact if the restore turns out to be bad.
  • The service cannot take downtime and the engine has no online backup tool. What is left?
    Take the backup from a replica rather than the primary — stop or quiesce the replica, copy its volume, bring it back and let it catch up — or take an atomic snapshot at the storage layer and accept crash-consistent recovery on restore. Both move the pause off the serving path rather than pretending a live tar is safe.
  • How long a stop timeout should the backup script give the container?
    Long enough for the engine's own shutdown to finish. `docker stop` sends SIGTERM and SIGKILLs after ten seconds by default, which can cut a checkpoint in half and turn a clean shutdown into a crash. Measure a normal shutdown, set `-t` comfortably above it, and check the exit code rather than assuming.

saying these in an interview costs you the question

  • Claims tar of a live data directory is atomic
  • Thinks docker cp is safer than tar for a live volume
  • Treats docker pause as equivalent to a clean shutdown
  • Believes Docker has a volume snapshot command
  • Restores over a populated volume and calls it done
  • Never restores a backup to verify it works

context

open as a page

How do you read the contents of a Docker named volume when no container is running?

level: juniorimportance: should knowfreq 52%

basics

~20 s

Start a throwaway container with the volume mounted: docker run --rm -v digest-data:/data alpine ls -l /data. A named volume is only reachable through a mount, and the viewing image need not be the application's own image.

open as a page

Your tar restore into a Docker named volume leaves the app's data directory empty. Why?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

Usually the archive's paths are wrong or the volume is not the one the app reads. An archive made with absolute paths restores into /data/data, and a service recreated with a different volume name reads a fresh empty volume instead of the restored one.

open as a page