Inside a running container you delete a directory such as /usr/share/doc that came from a read-only image layer. How does the union filesystem record that deletion, and why does the deleted data still exist on the host?
answer
- delete = tombstone in upper, not removal below
- whiteout = char device 0/0
- directory-wide = trusted.overlay.opaque=y
- .wh.<name> and .wh..wh..opq in layer tars
- hidden ≠ erased; secrets stay extractable
basics
~20 sLower layers cannot be modified, so the deletion is recorded as a whiteout in the writable layer that hides the name. The original bytes stay in the read-only image layer, still on disk and still extractable, just invisible through the merged view.
solid answer
~50 sRead-only lower layers cannot be changed, so a delete is recorded as a **marker in the writable layer**. At the kernel level, overlayfs creates a **whiteout**: a character device with device number 0/0 carrying the deleted name in upperdir. Lookup sees it and reports the name as absent even though a lower layer still has the file. Removing a whole directory whose lower version has entries is recorded by marking the upper directory **opaque** — the extended attribute `trusted.overlay.opaque="y"` — so none of the lower entries show through. In the OCI image-layer *tar* format the same idea is encoded differently: a zero-length file named `.wh.<name>` deletes that entry, and `.wh..wh..opq` marks the directory opaque. The consequence that matters: the data is hidden, not erased. It remains in the lower layer on disk and in the image's blob, so anything deleted this way — including a secret removed in a later build step — is still recoverable from the image.
code
bash · 5 linesUPPER=$(docker inspect --format '{{.GraphDriver.Data.UpperDir}}' mycontainer)
ls -l "$UPPER/usr/share"
# c--------- 1 root root 0, 0 doc <- whiteout tombstone
getfattr -n trusted.overlay.opaque "$UPPER/usr/share/doc" 2>/dev/nullgo deeper
Say that the delete is recorded as a marker in the writable layer and the original file still exists underneath, hidden.
Name the mechanisms precisely — character-device whiteout, opaque directory xattr, and the .wh. conventions in layer tars.
Draw out the consequences: secrets deleted in a later step remain extractable and must be rotated, and deletions never reclaim image size.
Treat any material that ever entered a shipped layer as disclosed, and set build-time rules and scanning so that credentials never reach a layer rather than being deleted afterwards.
## The problem deletion poses Image layers are read-only and shared between containers, so a container cannot remove a file from one — that would change what every other container sees. The union filesystem therefore needs a way to say "this name is gone" without touching the layer that holds it. ## Whiteouts in the kernel overlayfs uses two mechanisms, both recorded in upperdir: **Whiteout entries.** Deleting a file that exists in a lower layer creates, in the corresponding upper directory, a **character device with major/minor 0/0** bearing that filename. Overlayfs treats such a node as a tombstone: during lookup the name is reported as not existing, and it never appears in a directory listing. The device node itself is tiny — the tombstone costs bytes, while the hidden content keeps its full size in the lower layer. **Opaque directories.** Deleting a directory that exists below, or replacing it, is recorded by creating a directory in upperdir with the extended attribute `trusted.overlay.opaque="y"`. Opaque means "do not merge the lower versions of this directory at all", so every lower entry disappears at once instead of needing one tombstone each. This is also how a directory that is deleted and then recreated behaves correctly rather than resurrecting old entries. Because the xattr lives in the `trusted.` namespace, setting it requires privilege — which is why unprivileged handling of layers is done by userspace tooling rather than by direct overlay mounts. ## Whiteouts in the image format When changes are captured into an image layer, the layer is a tar archive and tar has no concept of a character-device tombstone that ordinary tools would honour. The OCI image-layer specification therefore encodes the same semantics with filename conventions: - `.wh.<name>` — a zero-length entry meaning "delete `<name>` from the layers below". - `.wh..wh..opq` — an entry inside a directory meaning "this directory is opaque; ignore lower entries". When a runtime unpacks layers, it translates these into real whiteouts for the storage driver. So the same idea exists twice: once as an on-disk kernel representation, once as a portable on-the-wire encoding. ## Why the data is still there A whiteout hides; it does not erase. The lower layer's copy is untouched: - On the host, the file is still present in the image layer's directory under the Docker data root. - In the registry, it is still inside that layer's blob, and anyone who can pull the image can extract it. The security consequence is the one interviewers chase. If a build step adds a credential, a private key or a `.git` directory, and a later step deletes it, the delete only writes a whiteout into a *new* layer. Both layers ship. Extracting the earlier layer recovers the secret in full, with no special tooling — unpack the tar and read it. Treat any secret that ever existed in a layer as disclosed and rotate it; making the file invisible in the final filesystem changes nothing about who can read it. Avoiding the situation in the first place is a build-structure question, handled by never introducing the material into a shipped layer. The same mechanism explains a size effect people meet constantly: removing files in a later step never shrinks the image, because both the original content and the tombstone are shipped. It can only add. ## Observing it On the host, listing the container's upperdir shows tombstones as `c---------` character devices with the deleted names, and `getfattr -n trusted.overlay.opaque` on a directory reveals opacity. Extracting an image's layer tars — for example with `docker save` and unpacking — shows `.wh.` entries directly, and reveals content that the running container reports as absent. ## Answering well Give the mechanism first (tombstone in the writable layer, opaque xattr for directories), then the tar-level encoding as the portable equivalent, then the two consequences: content remains recoverable from earlier layers, and deleting never reclaims image size.
- A Dockerfile copies a private key, uses it, then deletes it in a later RUN step. Is the key safe?No. The deletion only writes a whiteout into a later layer; the key remains intact inside the earlier layer's blob and ships with the image. Anyone who can pull the image can unpack that layer and read it, so the key must be treated as compromised and rotated. Keeping it out of any shipped layer in the first place is the only real fix.
- Why does deleting files in a later build step never reduce the image size?Layers are additive: a later layer can add a tombstone but cannot remove bytes from an earlier one. The image is the sum of all layer blobs, so the original content plus the tombstone both ship, and the total can only grow. Size must be addressed by not adding the content to a shipped layer at all.
Painting over a word with correction fluid on the top transparency sheet: looking down, the word is gone, but the sheet underneath still carries it, and lifting the stack off reveals it intact.
saying these in an interview costs you the question
- Believing a delete removes the file from the underlying image layer
- Assuming files deleted in a later Dockerfile step are unrecoverable from the image
- Claiming that deleting files late in a build shrinks the resulting image
- Thinking a whiteout is a normal empty file, or that it stores the deleted content
- Ignoring the directory-level opaque marker and assuming every hidden entry needs its own tombstone