Why does the same image reach one host in seconds and take minutes on a host that never ran it?
answer
- only the missing bytes move
- every blob asked for by digest
- local store checked before the network
- layers transfer whole, never diffed
- compressed on the wire, expanded on disk
basics
~20 sA pull moves only the layer blobs a host is missing. A host that already holds most of them reads a manifest and little else, while a host with an empty content store has to transfer every layer and expand it to disk.
solid answer
~40 sAn image is a manifest plus a set of blobs — one configuration blob and one per layer — and every blob is named by a digest computed from its own bytes. The host keeps a local content store keyed by those digests. On a pull it resolves the reference, reads the manifest, and asks only for the digests it does not already hold. A host that ran this image before is missing nothing, so essentially no payload crosses the network. A host with an empty store transfers everything, compressed, and then expands it onto disk. The interesting middle case is a host running a different image built on the same base: those base digests are already present, so only the unique top layers move.
go deeper
Recall that an image is a set of layers named by digests, and that a pull fetches only the ones the host is missing. That single sentence answers the question.
Explain the mechanics: the manifest lists blobs by digest, the host checks its local store per digest, and blobs move whole because a changed byte means a changed digest.
Show that you plan for the cold case — measuring start-up on an empty store, keeping a shared base so scale-out is partly warm, and not reading pull time as image size.
Frame it as a fleet property: how much of your image surface is shared across teams decides how much traffic and cold-start latency the whole estate pays when it grows.
## What a pull actually moves An image is not one file. It is a small document — the **manifest** — plus a set of independent, immutable **blobs**: one configuration blob holding runtime defaults, and one blob per **layer**. Each blob is named by a **digest**, a cryptographic hash of its own bytes. Two blobs with the same digest are the same bytes anywhere in the world, and two different byte sequences cannot share a digest. That one property is what makes a pull incremental. The host keeps a local content store keyed by digest. Told to run a reference, it resolves that reference to a manifest, reads the list of blobs the manifest names, and asks one question per entry: *do I already hold this digest?* Only the digests it does not hold are requested over the network. ## Why one host is fast and the other is slow - A **cold host** — an empty store, a host just added to the pool, a host whose store was recently reclaimed — holds none of those digests. It transfers the configuration blob and every layer, compressed, then expands them onto disk. That is the minutes. - A **warm host** that has run this exact image holds every digest already. It resolves the reference, reads the manifest, finds nothing missing, and is ready. Essentially no payload crosses the network; what remains is a round trip or two. That is the seconds. - A **partly warm host** — one running a *different* image built on the same base — is the common case in a real fleet. The base layers sit in the store under the same digests, so only the layers unique to this image move. A large image can arrive as a small transfer. | Host state | What is requested | What dominates the clock | |---|---|---| | empty store | manifest, then every blob | compressed transfer, then expansion | | ran this image before | manifest only | one or two round trips | | shares a base with a present image | manifest, then the unique top layers | the size of those unique layers | ## Layers move whole, never diffed A pull has no sub-layer difference mechanism. A blob is fetched entirely or not at all. If one file inside a layer changes, that layer's bytes change, so its digest changes, so it is a blob the host has never seen — and the whole thing transfers, however small the edit was. A layer's digest depends only on **its own** bytes, so a layer sitting above a changed one is re-transferred only if its own bytes also changed. In practice they usually did, because most builders regenerate everything downstream of a change. The useful mental model is therefore: *the stable bottom of the stack is what you get for free, and everything from the first change upward is paid for in full.* ## What else the clock is measuring 1. **Compressed on the wire, expanded on disk.** Blobs travel compressed; the host writes out the expanded form. Disk consumption is reliably larger than the download figure, sometimes by a wide margin. 2. **Expansion is work.** Writing out a large layer costs processor time and disk throughput on the host, and that cost does not shrink because the network was fast. 3. **Contention.** A cold pull while many other hosts are pulling the same thing is slower than a cold pull on its own, because they all queue behind the same source. ## What this changes in practice - **Never time a cold start on a machine that is already warm.** The number you get is the manifest round trip, not the transfer, and it will be wrong by two orders of magnitude when a fresh host joins. - **A shared base is the cheapest optimisation available.** If every image in the fleet stands on the same lower layers, almost every host is partly warm and only small top layers ever move. - **Pull time is not a proxy for image size.** A fast pull can mean a tiny image or a warm host, and only one of those still holds when the pool grows. - **A small edit is not a small transfer.** Changing one file low in the stack can cost the same bytes as a rebuild, which is why where a change lands in the layer order matters as much as how big it is. The whole behaviour follows from content addressing: because a blob's name is derived from its content, the host can answer "do I have this already?" locally, with certainty, before it opens a connection. Incremental transfer is a consequence of that, not a separate feature.
- Two images share a base layer. What does pulling the second one transfer?Only the blobs whose digests are absent from the local store. The shared base is already present under the same digests, so it is skipped entirely, and the transfer is just the layers unique to the second image — which can be a tiny fraction of its expanded size.
- One file changed in the top layer of an image. What crosses the network on a host that holds the previous version?That entire layer blob. Changing any byte changes the layer's digest, and a digest the host has never seen is fetched in full. There is no file-level difference inside a blob; the other layers, whose digests are unchanged, are not requested.
- Why is the on-disk footprint bigger than the number the pull reported?Blobs are transferred compressed and written out expanded, so disk always exceeds the transfer figure. Sizing host disk from download numbers is how a fleet runs out of space while every dashboard says the images are small.
Restocking a workshop by part number: you read the parts list, check the drawers, and order only the numbers that are missing. The number is stamped on the part itself, so a part that changed at all is a different number and arrives whole.
saying these in an interview costs you the question
- Thinks a pull sends only the files that changed inside a layer
- Believes a fast pull proves the image is small
- Assumes the same image takes the same time on every host
- Thinks re-running an image re-downloads layers already on disk
- Sizes host disk from the compressed download figure