skip to content

In a Dockerfile, someone installs packages in one RUN instruction and deletes the package cache in a separate, later RUN instruction, but the built image is no smaller. Explain why, and how you would fix it.

level: juniorimportance: must knowfreq 70%

answer

  1. layers additive, image = sum of blobs
  2. whiteout hides, does not delete
  3. chain install && rm in one RUN
  4. apk --no-cache, rm /var/lib/apt/lists/*
  5. deleted secret still extractable

basics

~20 s

Every Dockerfile instruction commits its own read-only layer. A later layer can only mark files as deleted; the bytes still sit in the earlier layer and still ship. Delete in the same RUN that created the files.

solid answer

~50 s

An image is a stack of immutable layers, one per filesystem-changing instruction. A later `RUN rm -rf` writes a whiteout entry that hides the path at runtime, but the original blob still exists in the earlier layer, and the image is the sum of all layers. The delete adds bytes instead of removing megabytes. Fix: create and clean up in one instruction, so the layer is committed only after cleanup. ```dockerfile RUN apt-get update && apt-get install -y --no-install-recommends curl \ && rm -rf /var/lib/apt/lists/* ``` Same rule for language caches (`pip --no-cache-dir`, `npm cache clean --force`) and downloaded tarballs: fetch, extract, delete in one RUN. The security corollary matters more than size: a credential copied in and deleted later is still extractable from the image, so use BuildKit secret mounts or a multi-stage build.

code

dockerfile · 9 lines
dockerfile
# bad: cleanup in its own layer
RUN apt-get update
RUN apt-get install -y curl build-essential
RUN rm -rf /var/lib/apt/lists/*

# good: one layer, committed after cleanup
RUN apt-get update \
 && apt-get install -y --no-install-recommends curl \
 && rm -rf /var/lib/apt/lists/*

go deeper

for a junior

State the rule: layers are additive, a later delete only hides, so clean up in the same RUN. Show the apt-get chain.

for a middle

Add the mechanism (whiteout entries in the union filesystem) and the cache-granularity tradeoff of chaining.

for a senior

Lead with the secret-leak consequence, then reach for multi-stage builds and BuildKit cache/secret mounts rather than heroic one-liners.

for a principal

Frame it as build hygiene policy: enforce multi-stage and secret mounts in shared templates and CI linting so authors never have to remember the rule.

## Layers are additive An image is an ordered stack of read-only layers. Each `RUN`, `COPY` or `ADD` commits one layer holding the filesystem changes that step made; metadata-only instructions (`ENV`, `LABEL`, `CMD`, `WORKDIR`) add no filesystem data. The size a registry stores is the sum of all layer blobs, not the size of the final merged filesystem. ## Why deleting later does nothing The container filesystem is a union of layers, and a layer cannot mutate what is below it. A deletion is therefore expressed as a whiteout: overlay2 writes a character device with major/minor 0/0 at that path, or an opaque-directory marker for a whole tree. Anyone reading the merged filesystem sees the file as gone, but `docker pull` still downloads the earlier layer containing the real bytes, and `docker save` plus `tar -x` recovers the file. So a standalone `RUN rm -rf /var/lib/apt/lists/*` actually makes the image slightly bigger. ## The fix: one instruction Chain creation and cleanup with `&&` (or a BuildKit heredoc) so the intermediate state is never committed: - apt: `apt-get update && apt-get install -y --no-install-recommends X && rm -rf /var/lib/apt/lists/*` - apk: `apk add --no-cache X` - source builds: download, `make install`, `rm -rf /tmp/src` in one RUN The tradeoff is cache granularity and readability: a giant RUN invalidates wholesale when any part of it changes. Group commands that change together, not everything. ## Better alternatives - **Multi-stage builds**: do the messy work in a builder stage and `COPY --from=builder` only the artifact. Nothing from the builder ships, so intermediate junk cannot leak at all. - **BuildKit cache mounts**: `RUN --mount=type=cache,target=/var/cache/apt ...` keeps caches on the builder for speed while writing nothing into a layer. - **Squashing** (`docker build --squash`, a legacy experimental daemon feature, or a `FROM scratch` + `COPY --from` flatten) collapses layers so deleted files really disappear. It works but destroys layer sharing between images and pull-time reuse, so a fleet of squashed images can move more bytes overall. Build clean rather than flattening dirty. ## Security corollary Because earlier layers stay retrievable, any secret that was ever written into a layer is compromised even though `docker run` shows nothing at that path. Use `RUN --mount=type=secret`, or fetch private dependencies in a builder stage whose layers never ship.

  • Does this mean every Dockerfile should collapse into a single RUN instruction?
    No. Chaining only helps where an instruction creates files a later instruction removes. Collapsing everything makes any small change invalidate the whole build and produces an unreadable Dockerfile. Group commands that share a lifecycle and change together, and keep stable steps separate so they stay cached.
  • A build copies a private SSH key, uses it, then deletes it in a later RUN. Is the key safe?
    No. The COPY layer still contains the key, and anyone with the image can extract it with docker save or by pulling and unpacking layers. Use BuildKit secret mounts, which expose the file only for that RUN and never commit it, or fetch dependencies in a builder stage whose layers are discarded.

Like sticky notes stacked on a page: covering a word with a new note hides it, but the page still weighs the same and peeling the notes off reveals it.

saying these in an interview costs you the question

  • Thinking rm in a later RUN reclaims space
  • Believing docker system prune shrinks an already-built image
  • Assuming a deleted secret is unrecoverable from the image
  • Collapsing an entire Dockerfile into one RUN to 'optimize'
  • Confusing merged filesystem size with the sum of layer sizes

context