skip to content

questions

5

In `docker pull` output, what do the per-layer lines `Already exists`, `Downloading` and `Extracting` each mean?

level: juniorimportance: must knowfreq 68%

answer

  1. an image is a list, not a file
  2. each layer has its own name
  3. that name is a hash of the bytes
  4. two phases per missing layer: move, then unpack
  5. compressed on the wire, expanded on disk

basics

~20 s

Already exists means that layer, identified by its sha256 digest, is already in the host's local image store, so nothing is transferred. Downloading is the compressed layer blob arriving from the registry. Extracting is that blob being decompressed and unpacked onto disk.

solid answer

~40 s

A `docker pull` fetches the image manifest first, then works layer by layer. Every layer is a separate blob addressed by a sha256 digest, so the daemon can check what it already has: a layer already in the local store prints **Already exists** and costs nothing. A missing layer prints **Downloading** while its compressed blob transfers, then **Extracting** while the daemon decompresses it and unpacks the files onto the storage driver, and finally **Pull complete**. Because layers are keyed by digest and not by image name, a brand-new image can print `Already exists` for most of its layers when it shares a base with something the host already ran. That is also why the second pull of the same tag finishes in a second: nothing is missing, so nothing is transferred or unpacked.

code

bash · 4 lines
bash
docker pull myrepo/tileserver:1.9
docker images myrepo/tileserver
docker system df -v | head -20
docker image inspect myrepo/tileserver:1.9 --format '{{range .RootFS.Layers}}{{println .}}{{end}}'

go deeper

for a junior

Be ready to read a pull transcript out loud: which lines mean network work, which mean local work, and which mean no work at all. Knowing that layers are matched by digest is the one fact the rest hangs on.

for a middle

An interviewer expects you to explain the mechanics: manifest first, per-layer digest check, bounded parallel downloads, then ordered extraction and decompression. Be able to say why the on-disk size and the transferred size differ.

for a senior

Show that you use this to reason about cold starts: a fresh host shares nothing and pays for every layer, while a host that has run a sibling image pays only for the top layers. Point at docker system df rather than guessing.

for a principal

Own the consequence for the platform: how much of your fleet's start-up cost is base-image reuse you could be getting and are not, and whether standardising base images across services is worth the coordination it takes.

### The unit of transfer is a layer, not an image An image is not one file. It is a *manifest* — a small JSON document listing a config blob and an ordered list of layer blobs — plus those blobs themselves. Every blob is content-addressed: its name is the sha256 digest of its own bytes. `docker pull myrepo/tileserver:1.9` first resolves the tag to a manifest, reads the layer list out of it, and only then decides what to move over the network. That design is the whole reason the progress output looks the way it does. Because a layer is named by a digest of its content, the daemon can ask a purely local question for each entry in the manifest — *do I already have this exact blob?* — before touching the network. ### What each status line means - **Already exists** — the layer is already in the local image store. Zero bytes cross the network. This is per layer, not per image: pulling an image you have never seen before will still print `Already exists` for every layer it shares with an image already on the host, most commonly the base-image layers. - **Waiting** — the layer is queued. The daemon transfers only a bounded number of layer blobs at a time (three by default), so the fourth and later missing layers sit in `Waiting` until a slot frees up. - **Downloading** — the *compressed* blob is arriving. The percentage and byte counter track compressed bytes, which is why the numbers here are smaller than what the image occupies on disk. - **Verifying Checksum** — the daemon hashes what it received and compares it with the digest from the manifest. A mismatch fails the pull; this is the integrity check that makes digests meaningful. - **Extracting** — the blob is being decompressed (gzip, for the usual media type) and its tar stream unpacked onto the storage driver as a new read-only layer directory. This step is CPU and disk work, not network work, and it happens in layer order because each layer is applied on top of its parent. - **Pull complete** — that layer is downloaded, verified and extracted. After the last layer the client prints the digest of the manifest it actually pulled, plus the image reference. Recording that `sha256:` digest is how you name the exact bytes you got, independent of where the tag points later. ### Why the second pull is instant Run the same `docker pull` twice and the second run prints `Already exists` for every layer and finishes in well under a second. Nothing about the tag is cached — the daemon still contacts the registry to resolve the tag and fetch the manifest — but every blob the manifest names is already local, so there is nothing to transfer or unpack. The same effect explains why a rebuild of a Kotlin service where only the last application layer changed pulls a few megabytes rather than the whole image: only the changed blob has a digest the host has never seen. ### The sizes do not line up, and that is expected `docker images` reports the *uncompressed* size of the layers as they sit on disk, summed. The registry stores and serves them compressed. A 1.87 GB entry in `docker images` may have moved only around 700 MB over the wire. Worse, the sum across several images double-counts nothing but also hides sharing: two images that share a 480 MB base occupy 480 MB once on disk, while `docker images` shows the base size in both rows. `docker system df` is the honest view of what the images actually cost on the host. ### Practical consequences Because reuse is by digest, a host that has run *anything* on the same base image pays a fraction of the cost of a host that has run nothing. A freshly created machine is the worst case in existence: no layer of any image is local, so every byte of every layer is transferred and unpacked before the first container can start. That gap between a warm host and a fresh one is the entire subject of image pull performance, and it starts here, in the difference between an `Already exists` line and a `Downloading` one. ``` $ docker pull myrepo/tileserver:1.9 1.9: Pulling from myrepo/tileserver 2d473b07cdd5: Already exists 9f54eef41275: Downloading [====> ] 62.4MB/271.3MB 0d3a1e5cd0ab: Waiting ... Digest: sha256:7c1f... Status: Downloaded newer image for myrepo/tileserver:1.9 ```

  • Why can a pull of an image you have never pulled before still print `Already exists` for most of its layers?
    Layers are stored and matched by their sha256 digest, not by the image name they came from. Any image built on the same base contributes identical base layer blobs, so those digests are already local. Only the layers unique to the new image — usually the application layers on top — have to be fetched.
  • The size in `docker images` is much larger than the bytes the pull transferred. Why?
    The registry stores layers compressed, typically as gzipped tar streams, and the download counter tracks those compressed bytes. `docker images` reports the uncompressed size the layers occupy after extraction. A rough 2-3x difference is normal, and it is also why extraction can take longer than the download on a slow disk.
  • Does a second `docker pull` of the same tag contact the registry at all?
    Yes. The daemon always resolves the tag and fetches the manifest, because a tag is mutable and may now point at different content. What it skips is the blob transfer: every layer digest in the manifest is already local, so the pull prints `Already exists` throughout. Pulling by digest instead of tag makes that resolution deterministic.

Pulling an image is like restocking a shelf from a parts list: you walk the list, skip every bin that already holds the exact part number, order only the missing ones, and unpack each box in order.

saying these in an interview costs you the question

  • Thinks the image is downloaded as one single file
  • Believes layers are cached by tag name rather than by digest
  • Says a repeated pull re-downloads everything
  • Confuses the size in docker images with bytes transferred
  • Assumes Extracting is still network work
  • Thinks Already exists means the layer exists in the registry

context

open as a page

What does the Docker daemon's `max-concurrent-downloads` setting control, and why does raising it often not speed a pull up?

level: middleimportance: should knowfreq 44%

basics

~20 s

It caps how many layer blobs the Docker daemon downloads in parallel for a pull, defaulting to 3 and set in daemon.json. Raising it adds streams, not bandwidth, and does nothing for extraction, which decompresses layers one at a time in order.

open as a page

A `docker pull` of a 1.87 GB image takes 4m12s on a fresh host and 38s on a warm one. How do you find where the time goes?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Split the pull into transfer and extraction and measure each: whether seconds pile up under Downloading or Extracting, whether one layer dominates the manifest, the real throughput to the registry, and CPU and disk during unpacking. The warm host only reuses layers it already holds.

open as a page

Container image pulls dominate cold-start time across your fleet. Which levers do you pull, and what does each cost?

level: principalimportance: should knowfreq 38%

basics

~20 s

Measure pull time as a first-class signal, then pick among four levers: move fewer bytes, move them a shorter distance, move them before they are needed, or start before they all arrive. Each trades a different cost — staleness, a dependency, coupling, or money.

open as a page

What do lazy-pull container image formats such as eStargz and SOCI change about container start-up?

level: seniorimportance: nice to knowfreq 17%

basics

~20 s

They make image layers seekable, so a container starts before its image has finished transferring and file content is fetched on demand. Both need a containerd snapshotter on the node, and they trade a faster start for a live registry dependency while running.

open as a page