A `docker pull` of a 1.87 GB image takes 4m12s on a fresh host and 38s on a warm one. How do you find where the time goes?
answer
- the warm host is not doing the same work
- split the clock into two phases
- look at the manifest before the network
- one layer can set the floor
- watch a core and a disk during unpack
basics
~20 sSplit the pull into transfer and extraction and measure each: whether seconds pile up under Downloading or Extracting, whether one layer dominates the manifest, the real throughput to the registry, and CPU and disk during unpacking. The warm host only reuses layers it already holds.
solid answer
~50 sTreat it as three suspects: **distance**, **shape** and **local cost**. First reproduce with timing and watch the progress lines — do the seconds pile up under `Downloading` or `Extracting`? Then look at the image's shape with `docker manifest inspect`: one 812 MB layer sets a floor that no concurrency setting removes, and many tiny layers behave differently from a few big ones. For distance, measure actual throughput from that host to the registry, remembering that registries commonly redirect blob requests to object storage, so the redirect target's location is what matters, not the registry hostname's. For local cost, watch CPU and disk during extraction: gzip decompression is single-threaded per layer, and a small burstable instance or a throughput-limited network disk shows up here. Finally, name the real asymmetry: the warm host prints `Already exists` for most layers, so it is not doing the same work at all.
code
bash · 5 linesdocker image rm myrepo/tileserver:1.9 2>/dev/null
docker system prune -a -f
docker system df
time docker pull myrepo/tileserver:1.9
docker system df -vgo deeper
Know that a fresh machine has no layers at all and must fetch everything, while a machine that ran a similar image reuses most layers. Being able to say that clearly is enough at this level.
Be able to break the pull into transfer and extraction, name what each is limited by, and read compressed layer sizes from the manifest without pulling the image.
Show a method, not a guess: reproduce cold on purpose, attribute the seconds to a phase, measure from the affected host including redirect targets, and map each finding to a specific lever instead of turning knobs hopefully.
Frame it as a fleet cost rather than one slow command: what pull time does to capacity that arrives late, whether the answer is proximity, pre-warming or image discipline, and what you are willing to spend to buy the minutes back.
### Frame the problem before touching a knob The 38 s warm figure is not a target — it is a different workload. A warm host already holds most of the image's layers by digest, so it transfers and unpacks only what changed. The 4m12s is the honest cost of materialising the whole image from nothing, which is exactly what every newly created node pays. The investigation is about decomposing that 4m12s. A pull decomposes into: manifest resolution (milliseconds), blob transfer (network), digest verification (CPU, cheap), and extraction — decompression plus unpacking onto the storage driver (CPU and disk). Attribute the time to one of those before proposing anything. ### Step 1 — watch which phase eats the clock Run the pull on a genuinely fresh host under `time`, with the progress output visible. The naive but decisive observation is whether the last minutes are spent with layers in `Downloading` or in `Extracting`. Take a `docker system df` before and after so you know what was actually materialised. Worth doing once with an empty store: `docker image rm` plus `docker system prune -a` on a scratch host reproduces the cold case on demand, which beats waiting for the next scale-out to observe it. ### Step 2 — read the image's shape ``` docker manifest inspect myrepo/tileserver:1.9 ``` This shows each layer's compressed size *without pulling the image*. For a Kotlin geospatial tile server built with a multi-stage Gradle build, a very common shape is a modest JRE base, a small application layer, and one enormous layer holding baked map assets — say 812 MB compressed out of a 1.87 GB image. That single layer is transferred by a single connection and extracted by a single thread, so it alone can account for most of the wall clock. Concurrency cannot split it; only changing what is in the image, or where the bytes come from, can. The opposite shape matters too: 40 small layers spend proportionally more time on per-blob overhead and benefit more from raising the daemon's `max-concurrent-downloads` above its default of 3. ### Step 3 — measure distance and throughput honestly Measure from *that host*, not from your laptop. Two facts routinely mislead people here: - Registries typically answer a blob request with a redirect to object storage or a CDN. The registry endpoint may be nearby while the blobs are served from another region. Follow the redirect and see where the bytes actually come from. - A single TCP stream on a long, fat path is latency-bound. If three concurrent streams together reach 300 Mbit/s on a 1 Gbit/s link with 140 ms of round-trip time, the link is not the constraint — the number of streams is, and this is the one case where raising download concurrency genuinely pays. If the bytes are simply far away, the durable fix is proximity: serve the image from a copy in the same region or the same network as the nodes. ### Step 4 — check the local cost During the `Extracting` phase, look at the host: - **CPU.** One core pinned at 100% while the network sits idle is gzip decompression, which is single-threaded per layer and cannot overlap across layers because layers apply in order onto their parent. Small burstable instances make this dramatic, especially once CPU credits are exhausted. - **Disk.** Unpacking writes an enormous number of small files. A throughput- or IOPS-limited network volume shows up as an extraction phase far longer than the download that fed it. Check the storage driver in `docker info` and the volume's limits. - **Store state.** Confirm the host really is cold. A node recycled from a warm pool that still has a partially populated image store will produce misleadingly good numbers. ### Step 5 — say what the fix class is Map the finding to the lever: | Finding | Lever | |---|---| | Time in Downloading, many layers, latency-bound | raise the daemon's download concurrency | | Time in Downloading, one dominant layer | move the bytes closer; split what the image carries | | Time in Downloading, link saturated | proximity or a nearer copy of the image; nothing local helps | | Time in Extracting, one CPU pinned | faster codec (zstd), fewer/smaller layers, bigger instance | | Time in Extracting, disk pegged | faster or larger volume, different storage driver setup | | Fresh hosts only | do not pull on the critical path: pre-pull or bake the image in | The last row is the one that most often ends the incident. If only brand-new hosts are slow, the fastest pull is the one that already happened: put the image on the host before it is asked to serve traffic. ### What not to conclude Do not benchmark against the warm host and call the gap a bug. Do not raise concurrency reflexively — on a saturated link it changes nothing, and across a whole fleet it is load you are aiming at your own registry. And do not report a number from one pull: run it several times on genuinely cold hosts, because a first-of-the-day pull can also be measuring cold caches on the serving side.
- The pull spends most of its time in `Extracting` with one CPU core saturated. What are your options?Decompression is single-threaded per layer and layers extract in order, so you cannot parallelise it. Options: give the host more per-core capacity or get off a credit-limited burstable instance, check whether the disk rather than the CPU is the real limit, move to a faster layer codec such as zstd if every engine in the fleet supports it, or reduce how much content has to be unpacked at all.
- You suspect the registry is far away. How do you confirm it from the node?Fetch a blob directly with curl from that host, following redirects, and record the effective URL, total time and download speed. Registries usually redirect blob requests to object storage, so the redirect target tells you where the bytes really live. Compare that throughput with a large transfer from a known-near endpoint to separate distance from a registry-side limit.
- Why is comparing against the warm host misleading?The warm host already holds most layers by digest and prints `Already exists` for them, so it transfers and unpacks only the changed layers. It is measuring an incremental pull. The only fair comparison for a fresh node is another fresh node, which is why reproducing cold, by clearing the image store on a scratch host, matters before drawing conclusions.
saying these in an interview costs you the question
- Blames the network without checking the extraction phase
- Raises download concurrency before measuring anything
- Treats the warm-host time as the achievable target
- Measures registry latency from a laptop, not the node
- Ignores that blob requests are usually redirected elsewhere
- Assumes a bigger instance always fixes pull time