A process inside a container opens a 4 GB file that came from a read-only image layer and writes one byte near the start. What does the union filesystem do, and what does that imply for write-heavy workloads?
answer
- copy-up = whole file, not blocks
- one byte → full-size read + write + disk
- chmod/chown/touch trigger it too
- chown -R at start = mass copy-up
- reads cheap, page cache shared; mounts bypass
basics
~20 soverlayfs copies the whole file from the image layer into the container's writable layer before applying the write — a copy-up. One byte costs 4 GB of I/O and 4 GB of disk. Write-heavy or large-file workloads belong on volumes, which skip this.
solid answer
~50 sOn first modification of a file that lives in a lower (image) layer, overlayfs performs a **copy-up**: the entire file — data, permissions, timestamps, xattrs — is copied into upperdir via workdir, and only then is the write applied to the copy. Subsequent writes to that file are ordinary writes, since it now exists in the writable layer. So a single-byte write to a 4 GB file costs a 4 GB read plus a 4 GB write, a large first-write latency spike, and 4 GB of extra host disk — the original still occupies space in the shared image layer. Implications: append-heavy logs, database files and anything large and mutable perform badly on the union filesystem, and their first touch is the worst. The remedy is to place such paths on a **volume, bind mount or tmpfs**, which are mounted over the merged view and bypass overlayfs entirely. Reads from lower layers are cheap and share host page cache across containers.
code
bash · 8 lines# bad: 4 GB file lives in an image layer and is modified at runtime
docker run -d myimage:1.0
# better: mutable path is a volume, so writes never enter overlayfs
docker run -d -v pgdata:/var/lib/postgresql/data postgres:17
# inspect how large a container's writable layer has grown
docker inspect --format '{{.GraphDriver.Data.UpperDir}}' mycontainer | xargs du -shgo deeper
Know that the whole file is copied into the writable layer before the first write, so modifying big image files is expensive.
Quantify the cost, note that later writes are normal, and place mutable data on volumes.
Diagnose from symptoms — first-write stalls, oversized writable layers, slow chown -R startups — and name non-obvious triggers such as metadata changes and hard-link breakage.
Set the platform rule: images carry immutable content only, mutable and large data is declared on mounts, and image build steps establish ownership so runtime never rewrites the tree.
## Copy-on-write at file granularity overlayfs is copy-on-write, but the unit of copying is a **whole file**, not a block or a page. The first time a container modifies a file that exists only in a read-only lower layer, the kernel: 1. Creates a copy of the file in workdir — contents, mode, ownership, timestamps and extended attributes. 2. Atomically moves it into upperdir. 3. Reopens the write against that upper copy. Only then does the one-byte write proceed. From that moment the file exists in the writable layer and behaves like an ordinary file; the copy-up cost is paid exactly once per file per container. ## Costs of a copy-up - **Latency spike on first write.** A 4 GB copy-up can take seconds to minutes depending on the disk. Applications experience this as an inexplicable stall on the first modification, then normal speed afterwards — a signature symptom worth recognising during diagnosis. - **Disk amplification.** The lower copy is not freed; it remains part of the shared image layer. The host now stores both. - **Metadata-only operations still trigger it.** `chmod`, `chown` and `touch` on a lower-layer file cause a copy-up too, because the attributes belong to the file. A recursive `chown -R` over an image's tree at container start can copy up gigabytes for no functional reason — a classic cause of slow startup. - **Hard links do not survive.** Files hard-linked in a lower layer become independent copies once copied up, so link counts observed inside the container can differ from the image's, and space usage grows. - **Rename across layers.** Renaming a lower-layer file or directory may require copy-up of the source, so "just a rename" can be expensive; directory renames in particular have historically been restricted. ## What is cheap - **Reads** from lower layers are near-native: overlayfs passes through to the underlying filesystem, and the host page cache is shared. Fifty containers reading the same library from the same image layer share one set of cached pages, which is a real memory advantage of containers over full VMs. - **New files** created by the container go straight into upperdir with no copy-up. - **Second and later writes** to an already-copied-up file are ordinary writes. ## Consequences for workload placement The rule that follows is simple: **large or write-heavy mutable data must not live on the union filesystem.** - **Databases** — data directories perform poorly, both from copy-up on the initial files and from the general overhead of the union. Mount a volume at the data directory. - **Logs** — a log file shipped inside the image and appended to will copy up on first append. Prefer writing to stdout/stderr, or to a mounted path. - **Caches and scratch** — use a volume, or `tmpfs` when the data is disposable and speed matters more than durability. - **Application-writable trees** — declare them as mounts rather than relying on the writable layer. Mounts of any kind are attached over the merged view, so I/O under those paths never enters overlayfs. That is why the same database is dramatically faster with a volume than without. ## Diagnosis in the field Symptoms that point at copy-up: a container that is fast on reads and pathologically slow on the first write to a specific big file; container start times dominated by a permission-fixing step; disk usage under the Docker data root growing far beyond what the container appears to have written. Check the container's `GraphDriver.Data.UpperDir` and measure its size — a writable layer far larger than the application's own output implies copies of image files. Design-side mitigations exist too: shipping large mutable files in the image at all is usually the deeper mistake, and initialising ownership at build time avoids the startup `chown -R`. ## Answering well Say "copy-up of the whole file, then write", quantify it (one byte costs the file's full size in I/O and disk), name at least one non-obvious trigger such as `chown`, and finish with the placement rule: mounts for anything large and mutable.
- Why does an entrypoint that runs `chown -R app:app /opt/app` make container startup slow?Changing ownership is a metadata change on each file, and any file that lives in a read-only image layer must be copied up in full before its attributes can change. A recursive chown therefore copies the entire tree into the writable layer, costing time and disk proportional to the image content. Setting ownership at build time with COPY --chown avoids it.
- Do fifty containers running the same image each cache that image's libraries separately in RAM?No. Reads from lower layers pass through to the underlying filesystem, so all fifty containers share the same host page cache entries for those files. This is a genuine density advantage: memory usage scales with what containers write and allocate, not with how many share the same image.
- Does copy-up happen again every time the file is written?No, only on the first modification. Once the file has been copied into the container's writable layer, it exists in upperdir and subsequent writes are ordinary writes with no union overhead. The cost is once per file per container, which is why the symptom is a one-off stall rather than sustained slowness.
Correcting one word on a page of a reference book you may not write in: you must photocopy the entire page first, then edit the photocopy. The cost is per page, not per word.
saying these in an interview costs you the question
- Saying overlayfs copies only the changed block or page rather than the whole file
- Believing read-only access to a lower-layer file triggers a copy
- Assuming metadata operations such as chown or touch are free on image files
- Claiming a volume mount still goes through the union filesystem
- Thinking the original copy in the image layer is freed after copy-up