skip to content

In a Dockerfile, what is the difference between the COPY and ADD instructions, and why do most style guides say to default to COPY?

level: juniorimportance: must knowfreq 75%

answer

  1. COPY = boring, predictable
  2. ADD = local tar extract + URL fetch
  3. Remote archives are downloaded, NOT extracted
  4. Zip is never extracted
  5. ADD --checksum for verified downloads

basics

~20 s

COPY just copies files from the build context into the image. ADD does that plus two magic behaviours: it auto-extracts local tar archives and can fetch remote URLs. Default to COPY because its behaviour is predictable; use ADD only when you want extraction.

solid answer

~50 s

Both put files into the image. `COPY <src> <dest>` copies files and directories from the build context, and that is all it does — which is exactly why it is the default recommendation: the behaviour is obvious from reading the line. `ADD` adds two implicit behaviours: 1. **Local tar auto-extraction.** If the source is a local archive in a recognised compression format, ADD extracts it into the destination instead of copying the file. A `.zip` is *not* extracted — only tar-family archives. 2. **Remote fetch.** A URL source is downloaded. Historically this was the worst part: no extraction of downloaded archives, no checksum verification, and the fetch became a cached layer, so a changed remote file could go unnoticed. Modern Dockerfile syntax adds `--checksum` (and Git-repo sources), which makes ADD defensible for remote fetches. Rule of thumb: `COPY` for everything from the context, `ADD` only to extract a local tarball, and `ADD --checksum=sha256:…` (or an explicit `RUN curl` with verification) for remote artifacts.

code

dockerfile · 16 lines
dockerfile
# Predictable: files from the context
COPY package.json package-lock.json ./
COPY src/ ./src/

# ADD's one clear win: unpack a local tarball
ADD vendor/toolchain.tar.gz /opt/

# Verified remote fetch (Dockerfile 1.6+)
ADD --checksum=sha256:3f7a1c9b0d2e5f6a7b8c9d0e1f2a3b4c5d6e7f8091a2b3c4d5e6f708192a3b4c \
    https://example.com/tool-1.4.0.bin /usr/local/bin/tool

# Alternative: verify, extract and clean up in one layer
RUN curl -fsSL https://example.com/tool-1.4.0.tgz -o /tmp/t.tgz \
 && echo "3f7a...4c  /tmp/t.tgz" | sha256sum -c - \
 && tar -xzf /tmp/t.tgz -C /opt \
 && rm /tmp/t.tgz

go deeper

for a junior

State the core split — COPY copies, ADD also extracts local tars and fetches URLs — and that COPY is the default because it is predictable.

for a middle

Add the traps: only tar-family archives, detection by content, remote sources are not extracted, and archive metadata determines resulting paths and ownership.

for a senior

Discuss verification and layer hygiene — --checksum versus a single RUN that downloads, verifies, extracts and deletes — plus path-traversal risk when extracting untrusted archives.

for a principal

Position it as supply-chain policy: pinned, digest-verified external artifacts, banning unverified remote fetches in CI lint rules, and where third-party binaries enter the image at all.

## The two instructions Both `COPY` and `ADD` add files to the image filesystem and both create a new layer. They share the same destination semantics: a trailing `/` means "into this directory", the destination is created if missing, relative destinations resolve against the current `WORKDIR`, and multiple sources require the destination to be a directory. Sources are interpreted relative to the build context root, and you cannot reference paths outside the context (`COPY ../secrets .` is an error). Where they differ is that `ADD` does extra work implicitly. ## ADD behaviour 1: tar auto-extraction If the source is a **local** file that Docker recognises as a tar archive — plain, gzip, bzip2 or xz compressed — `ADD` unpacks it into the destination instead of placing the archive there. This is how base images are traditionally built: `ADD rootfs.tar.xz /` is the canonical `FROM scratch` pattern. The pitfalls are the surprise cases: - **Only tar-family archives.** A `.zip`, `.7z` or `.rar` is copied as a file, not extracted. Candidates routinely assume ADD extracts "archives" generally. - **Detection is by content, not extension.** A file named `data.bin` that happens to be a gzipped tar gets extracted. This is a genuine surprise when shipping a binary blob. - **The archive's own metadata wins.** Ownership, permissions and paths come from inside the tar, including any leading directory, so `ADD app.tar.gz /opt/` may produce `/opt/app/...` rather than `/opt/...`. A maliciously crafted archive with `../` entries is a path-traversal concern when you extract untrusted input. - **Remote archives are NOT extracted.** `ADD https://…/x.tar.gz /tmp/` downloads the file and leaves it as a file. Only local sources extract. This asymmetry catches people constantly. ## ADD behaviour 2: remote sources A URL source is downloaded during the build. Classic problems: - **No integrity check** (historically). You got whatever the server served, forever baked into a layer. - **Cache staleness.** The instruction's cache key did not follow the remote content in the old builder, so a changed upstream file could be silently skipped on rebuild. - **No decompression or cleanup.** You still need a `RUN` to unpack, and the downloaded file persists in that layer even if a later `RUN` deletes it. Modern Dockerfile syntax (1.6+) added `ADD --checksum=sha256:<digest> <url> <dest>`, which verifies the download and makes the digest part of the cache key. That removes the strongest historical objection, so "ADD is always evil" is now an outdated absolute — the accurate statement is that unverified remote ADD is bad. Dockerfile 1.7+ also allows Git repository URLs as ADD sources. ## Why default to COPY Predictability. When a reviewer reads `COPY dist/ /app/`, there is exactly one possible outcome. `ADD dist/ /app/` might behave identically today and differently tomorrow if a file in `dist/` becomes a tarball. Reviewers should not have to reason about the *content type* of the sources to know what an instruction does. Both instructions also support `--chown`, `--chmod`, `--link` and (for COPY) `--from`, so switching to COPY costs no capability except extraction and fetch. The practical decision table: | Need | Use | |---|---| | Files from the build context | `COPY` | | Files from another build stage or image | `COPY --from=…` (ADD has no `--from`) | | Unpack a local tarball | `ADD` (its one clear win) | | Fetch a remote artifact | `ADD --checksum=sha256:…`, or `RUN` with an explicit download + verify | ## The RUN-download alternative `RUN curl -fsSL <url> -o /tmp/x.tgz && echo "<sha> /tmp/x.tgz" | sha256sum -c - && tar -xzf /tmp/x.tgz -C /opt && rm /tmp/x.tgz` gives verification, extraction and cleanup in a single layer, so the archive never persists in the image. Its cost is that it requires `curl`/`wget` in the image (fine in a builder stage, unwelcome in a lean runtime stage) and the layer's cache key is the command string, so the URL must be versioned rather than "latest". Choosing between this and `ADD --checksum` is a reasonable thing to discuss in an interview: ADD needs no tools in the image, RUN gives you cleanup within the same layer. ## Small things interviewers probe - Neither instruction preserves the host's user *names*; ownership defaults to UID/GID 0 unless `--chown` says otherwise. - `COPY` directory semantics: `COPY src/ /dest/` copies the *contents* of `src`, not the directory itself. - Both are affected by `.dockerignore`, which is why an excluded file produces a "not found" error that looks like a typo. - `ADD` cannot copy from another build stage; there is no `--from` on ADD.

  • What happens with `ADD https://example.com/app.tar.gz /opt/` — is the archive unpacked?
    No. Auto-extraction applies only to local sources; a remote source is downloaded and left as a file at the destination. You would still need a RUN to unpack it, and unless you pass `--checksum` there is no integrity verification, so you have all of ADD's downsides and none of its convenience.
  • Given that `ADD --checksum` exists, is there still a reason to prefer RUN with curl for remote artifacts?
    Yes, when you need the downloaded file gone from the final image. ADD writes the artifact into a layer, so removing it later only hides it. A single RUN that downloads, verifies the digest, extracts and deletes the archive leaves only the extracted result in that layer. The tradeoff is that RUN requires curl or wget to exist in the stage, which is fine in a builder but undesirable in a minimal runtime stage.

saying these in an interview costs you the question

  • Believing ADD extracts remote archives too
  • Thinking ADD unpacks .zip files
  • Claiming ADD is always forbidden, unaware of --checksum
  • Assuming a later RUN rm of an ADDed archive removes it from the image
  • Trying COPY with a source path outside the build context, like ../secrets

context