skip to content

Union Filesystems & overlay2

overlay2 stacks the read-only image layers under one writable container layer, so the first write to an existing file copies the whole file up. Interviewers ask because it explains slow container writes, whiteout files, and why the layer is disposable.

part ofDockeroverview, primer and where to startread it →
on this pageshow

questions

4

A running container writes a 2 GB file into a path inside itself. Where does that data physically live on the host, what happens to it when the container is removed, and why does the image on disk stay unchanged?

level: juniorimportance: must knowfreq 55%

answer

  1. image layers read-only, one writable layer per container
  2. writes land in container diff dir
  3. docker rm deletes it; stop/start does not
  4. many containers share one image copy
  5. persist or write-heavy → volume/bind mount

basics

~20 s

It lands in the container's own thin writable layer on the host, on top of the read-only image layers. Removing the container deletes that layer and the data. Image layers are never written to, which is why many containers can share one image.

solid answer

~50 s

An image is a stack of **read-only** layers. When a container starts, the runtime adds one thin **writable layer** on top and presents the union of all of them as the container's root filesystem. Every write — new files, modified files, deletions — is recorded only in that top layer, under the storage driver's directory on the host (for overlay2, the container's `diff` directory beneath the Docker data root). So the 2 GB file consumes host disk in the container's writable layer, not in the image. The image is untouched, which is exactly what lets ten containers from the same image share one copy of its layers. `docker rm` deletes the writable layer, and the data is gone. `docker stop`/`start` does not — a stopped container keeps its layer until removed. Data that must outlive the container, or that is write-heavy, belongs on a volume or bind mount, which bypasses the union filesystem entirely.

code

bash · 7 lines
bash
docker system df -v

# reclaim the writable layers of stopped containers
docker container prune

# keep data outside the union filesystem
docker run -d -v appdata:/var/lib/app myimage:1.0

go deeper

for a junior

Name the three pieces — read-only image layers, per-container writable layer, union view — and say the data dies with the container.

for a middle

Add where it lives on the host under the storage driver, the stop-versus-rm distinction, and why sharing read-only layers keeps disk usage low.

for a senior

Separate lifetime, performance and operability arguments for using mounts, and mention hosts filling up with unremoved stopped containers.

for a principal

Frame it as a platform rule: containers hold no state, all mutable data is on declared mounts, and image immutability is the property that makes rollout and rollback deterministic.

## The layered picture A container image is a stack of layers, each a set of filesystem changes, and every one of them is **read-only** once built. Starting a container does not copy that stack. Instead the runtime creates one new, initially empty directory — the **writable container layer** — and asks the storage driver to present the image layers plus this new directory as a single filesystem tree. The container sees an ordinary root filesystem at `/`; underneath, it is a union of directories on the host. ## Where writes go All modifications are captured in the writable layer: - **Creating a file** writes it into the writable layer only. - **Modifying a file that came from an image layer** copies it up into the writable layer first, then edits the copy. The original in the image layer is untouched and simply hidden. - **Deleting a file that came from an image layer** records a marker in the writable layer that hides it. The bytes remain in the lower layer. So the 2 GB file physically occupies 2 GB in the container's writable layer, under the Docker data root — with overlay2, `/var/lib/docker/overlay2/<id>/diff` — and the image's own layer directories are byte-identical to what they were before. ## Why images stay pristine This immutability is what makes containers cheap. Ten containers started from the same 400 MB image do not consume 4 GB: they share the same read-only layer directories on disk and each adds a writable layer that starts empty. It is also why an image is a reproducible artifact — nothing a running container does can alter it, so the same image started tomorrow behaves the same way. ## Lifecycle of the writable layer - `docker stop` then `docker start`: the writable layer persists, so the file is still there. - `docker restart`: same — the container object survives. - `docker rm` (or `docker run --rm` on exit): the writable layer is deleted with the container, and the 2 GB goes with it. - Recreating a container from the same image gives a fresh, empty writable layer. There is no continuity. A common surprise: `docker rm` frees the space, but only then. Stopped containers still hold their writable layers, which is why hosts fill up with old containers nobody removed. `docker ps -a` and `docker system df -v` reveal this, and `docker container prune` reclaims it. ## Why you should not store real data there Three separate reasons, and candidates should distinguish them: 1. **Lifetime.** The data is tied to a container instance, and containers are meant to be replaceable. Any redeploy loses it. 2. **Performance.** Writes through the union filesystem pay extra cost, most visibly the copy-up of an entire existing file on its first modification. Databases and other write-heavy workloads are noticeably slower. 3. **Portability and inspection.** The writable layer is storage-driver-specific and awkward to back up. A volume is a plain directory the host can snapshot and a new container can re-attach. The answer is to mount a **volume** or **bind mount** at the path that holds mutable data. Mounts are not part of the union: reads and writes go straight to the host filesystem, so they survive `docker rm` and avoid copy-up. `tmpfs` mounts serve the opposite case — data that must be fast and must never touch disk. ## What to say when asked Name the three parts: read-only image layers, one writable layer per container, union view on top. Then state the consequence — writes are container-scoped and disappear with the container — and finish with the remedy, mounts for anything that must persist or is write-heavy.

  • Does stopping a container free the disk space its writes consumed?
    No. A stopped container still exists as an object with its writable layer intact on the host, which is why the data is still there after `docker start`. Only removing the container — `docker rm`, `docker container prune`, or `--rm` at run time — deletes the layer and reclaims the space.
  • Ten containers run from the same 400 MB image. How much disk do the image layers use?
    About 400 MB in total, not 4 GB. The image's layer directories are read-only and shared by every container started from that image; each container adds only its own initially empty writable layer. Disk usage grows with what the containers actually write, not with how many of them there are.

The image layers are a printed book you may not write in; the writable layer is a sheet of tracing paper laid on top. You write on the tracing paper, and reading through it you see your edits over the print. Throw the sheet away and the book is untouched.

saying these in an interview costs you the question

  • Saying writes go into the image, so the image changes as the container runs
  • Believing `docker stop` frees the space consumed by container writes
  • Assuming each container gets its own full copy of all image layers
  • Treating the writable layer as durable storage for databases or uploads
  • Confusing a named volume with the writable layer — volumes deliberately bypass the union filesystem

context

open as a page

Docker's default storage driver on Linux is overlay2, which builds a container's root filesystem with an overlay mount. Explain what the lowerdir, upperdir, workdir and merged directories are and how they combine.

level: middleimportance: must knowfreq 48%

basics

~20 s

lowerdir is the stack of read-only image layers, upperdir is the container's writable layer, merged is the unified view the container uses as /, and workdir is a private scratch directory the kernel needs for atomic operations. Upper entries shadow lower ones.

open as a page

Inside a running container you delete a directory such as /usr/share/doc that came from a read-only image layer. How does the union filesystem record that deletion, and why does the deleted data still exist on the host?

level: middleimportance: must knowfreq 44%

basics

~20 s

Lower layers cannot be modified, so the deletion is recorded as a whiteout in the writable layer that hides the name. The original bytes stay in the read-only image layer, still on disk and still extractable, just invisible through the merged view.

open as a page

A process inside a container opens a 4 GB file that came from a read-only image layer and writes one byte near the start. What does the union filesystem do, and what does that imply for write-heavy workloads?

level: seniorimportance: should knowfreq 40%

basics

~20 s

overlayfs copies the whole file from the image layer into the container's writable layer before applying the write — a copy-up. One byte costs 4 GB of I/O and 4 GB of disk. Write-heavy or large-file workloads belong on volumes, which skip this.

open as a page