skip to content

Pushing a rebuilt version of an 800 MB image often transfers only a few megabytes, and copying that image to a second repository on the same registry can transfer nothing at all. What mechanisms make that possible?

level: middleimportance: should knowfreq 40%

answer

  1. Blob = sha256 of bytes → dedup for free
  2. HEAD blob per layer, upload only 404s
  3. Existence is per repository, not per registry
  4. mount=<digest>&from=<repo> → 201, zero bytes
  5. Recompression changes the digest, kills reuse

basics

~20 s

Layers are content-addressed by SHA-256 digest, so before uploading the client asks the registry whether each digest already exists and skips those it has. Within one registry, a cross-repository blob mount links an existing blob into another repository, transferring zero bytes.

solid answer

~60 s

Two mechanisms, both resting on content addressing. **Existence checks.** Every layer and the config are blobs named by the digest of their bytes. Push begins with `HEAD /v2/<repo>/blobs/<digest>` per blob and uploads only the 404s. A rebuild that changes only application code leaves every base-image layer byte-identical, so only the top layer and a tiny manifest travel. This is also why Dockerfile ordering matters: putting rarely-changing layers below frequently-changing ones maximises what stays identical. **Cross-repository blob mount.** The existence check is scoped per *repository*, so `team/app` cannot see blobs stored under `team/base` even on the same server. The API closes that gap: `POST /v2/team/app/blobs/uploads/?mount=<digest>&from=team/base` returns `201 Created` and links the blob with no upload. Registries typically require the caller to have pull rights on the source repository, and fall back to a normal upload session (`202`) if they refuse. A caveat: identical *content* is not enough if the bytes differ. Re-compressing a layer with different settings yields a different digest and defeats reuse, which is why reproducible build tooling and cache reuse matter.

code

bash · 11 lines
bash
# Second push of a rebuilt image: base layers report as already present
docker push registry.example.com/team/app:1.3
# 5f70bf18a086: Layer already exists
# 8c2e19a1c2b4: Pushed        <- only the app layer
# 1.3: digest: sha256:9b1c... size: 1160

# Server-side copy attempt using a cross-repo mount
curl -si -X POST -H "Authorization: Bearer $TOKEN" \
  "https://registry.example.com/v2/team/app/blobs/uploads/?mount=sha256:ab12...&from=team/base"
# 201 Created            -> linked, nothing transferred
# 202 Accepted + Location -> declined, upload normally

go deeper

for a junior

Know that layers are identified by a hash and that unchanged layers are not re-uploaded.

for a middle

Explain the per-blob HEAD check, why layer ordering in the build affects push size, and what a cross-repository mount does.

for a senior

Cover the per-repository scoping rationale, the permission requirement for mounts, and the practical causes of digest churn that silently kill reuse.

for a principal

Reason about it as a distribution cost model: shared base images, promotion between namespaces, cross-registry replication costs, and build determinism as an infrastructure concern.

## Content addressing is the foundation A registry stores **blobs** keyed by `sha256:<hex>` of their exact bytes. A layer blob is a compressed tar of the filesystem changes that layer introduces. Because the name is derived from the content, two questions become trivial: "do you already have this?" is a lookup, and "is what you sent me what you claimed?" is a hash comparison. Everything below follows from that. Note the distinction between two digests you will hear about: the **blob digest** names the compressed bytes as stored in the registry, while the **diff_id** in the image config names the *uncompressed* tar. Registry deduplication works on the blob digest. ## Mechanism 1 — skip what already exists Before transferring anything, the push loop issues `HEAD /v2/<name>/blobs/<digest>` for each blob. A `200` means present and the client moves on; a `404` triggers an upload session. For a typical service image: - OS base layers: unchanged across every build → skipped - runtime/JDK layer: unchanged for months → skipped - dependency layer: changes when the lockfile changes → usually skipped - application layer: changes every build → uploaded (a few MB) Hence "800 MB image, 4 MB push". The corollary is a build-time discipline: layers whose content changes on every build should sit at the top, and anything you copy in should be split so that volatile files do not invalidate stable ones. Copying the whole source tree before installing dependencies changes the dependency layer's bytes on every commit, and the push (and everyone else's pull) pays for it. The pull side benefits identically: the client fetches only blobs its local store lacks, so two images sharing a base download it once per host. ## Mechanism 2 — cross-repository blob mount A subtlety trips people up: **existence is checked per repository**. A registry stores each blob once in its backend, but each repository maintains its own set of references to blobs — this is what makes per-repository access control and garbage collection tractable. So `HEAD /v2/team/app/blobs/<digest>` returns 404 even when `team/base` in the same registry holds that exact blob. Without a remedy, copying an image from one repository to another would mean downloading and re-uploading everything through the client. The Registry API provides: ``` POST /v2/team/app/blobs/uploads/?mount=sha256:ab12...&from=team/base ``` - `201 Created` — the blob is now referenced by `team/app`; **no bytes were transferred**. - `202 Accepted` with a `Location` — the registry declined to mount (unsupported, or the caller lacks read access to the source) and handed back a normal upload session; the client uploads as usual. Authorization for a mount needs a token whose scope covers both repositories: `push` on the destination and `pull` on the source. That is why an identity restricted to one namespace silently loses the optimization and falls back to full uploads — a real cause of "why does promoting an image take twenty minutes now". Tools that copy images server-side (registry-to-registry copy utilities, promotion pipelines) lean on this heavily, as does re-tagging an image into a release namespace. ## When deduplication fails to kick in - **Different bytes, same files.** Recompressing a layer, or rebuilding it from scratch with a different tar ordering or timestamps, changes the digest. The content may be logically identical; the registry sees a new blob. Build caching and deterministic layer production preserve digests. - **A different registry.** Mounts work within one registry only. Moving between registries transfers every blob the destination lacks. - **A cache-busting instruction early in the build.** An instruction whose output changes each build invalidates every layer beneath it. - **Squashing or flattening images.** Collapsing everything into one layer produces a single huge blob that changes entirely on every rebuild — you trade a small transfer for a full one, and destroy sharing across images. - **Cross-registry replication with re-signing or re-compression** in the pipeline, which again alters bytes. ## Storage versus transfer Be precise about what is saved. Deduplication in the registry backend means the bytes are stored once regardless of how many repositories reference them; garbage collection removes a blob only once no manifest anywhere references it. Transfer savings come from the existence checks and mounts. On the client side, the local image store also shares layers between images, so pulling a second image built on the same base costs only its unique layers on disk and on the wire.

  • Why is the blob existence check scoped per repository rather than per registry?
    Because references are what make per-repository access control and garbage collection possible: if any repository could reach any blob by digest, knowing a digest would grant access to content in repositories you cannot read, and deciding when a blob is unreferenced would be harder. Each repository therefore keeps its own references even though the backend stores the bytes once. The cross-repository mount exists to recover the transfer savings while still requiring read permission on the source.
  • Two builds produce logically identical layers but the push uploads both. What could cause that?
    The blob digest covers the exact compressed bytes, so anything that changes them — a different compression level or implementation, different file ordering or timestamps in the tar, a rebuilt rather than cached layer — produces a new digest. Deterministic build tooling and preserving the build cache keep the bytes identical so the existence check hits.

The registry is a warehouse of numbered crates where the number is a fingerprint of the contents: a new shipment only needs the crates the warehouse lacks, and moving a crate to another aisle is a paperwork change, not a truck.

saying these in an interview costs you the question

  • Claiming the registry deduplicates by comparing file contents rather than blob digests
  • Assuming a blob present in one repository is automatically usable by another on the same registry
  • Believing cross-repository mounts work between two different registries
  • Thinking squashing images into one layer improves push efficiency
  • Saying identical file content guarantees an identical digest regardless of compression

context