skip to content

Push, Pull & the Registry API

docker push uploads every blob the registry is missing and writes the manifest last; docker pull reads that manifest and fetches only the layers a node lacks. Asked to see whether you treat an image as content-addressed blobs rather than one opaque file.

part ofDockeroverview, primer and where to startread it →
on this pageshow

questions

5

Walk through what a container client sends to a registry during an image push: which objects are uploaded, in what order, and why the manifest is written last.

level: middleimportance: must knowfreq 50%

answer

  1. HEAD blob → 404 → POST uploads/ → PUT ?digest
  2. Blobs first, manifest LAST
  3. Registry verifies digest before commit
  4. mount=<digest>&from=<repo> = zero-byte copy
  5. Existence checks scoped per repository

basics

~20 s

For each layer and the config blob, the client HEADs the blob digest; if absent it opens an upload session, sends the bytes, and completes with the digest. Only after every referenced blob exists does it PUT the manifest under the tag — so a tag never resolves to missing content.

solid answer

~50 s

Push is blob-first, manifest-last. 1. **Authorize** for scope `repository:<name>:pull,push` via the token handshake. 2. **For each blob** — every layer plus the image config JSON — send `HEAD /v2/<name>/blobs/<digest>`. A `200` means the registry already has it, so it is skipped; a `404` means upload. 3. **Upload** a missing blob: `POST /v2/<name>/blobs/uploads/` returns `202` with a `Location` upload URL, then either a monolithic `PUT <location>?digest=sha256:...` or chunked `PATCH` calls followed by that `PUT`. The registry verifies the content hashes to the claimed digest before committing it. 4. **Cross-repo shortcut**: if the blob exists elsewhere on the same registry, `POST /v2/<name>/blobs/uploads/?mount=<digest>&from=<other-repo>` returns `201` with no bytes transferred. 5. **PUT /v2/<name>/manifests/<tag>** with the manifest JSON and its media type. The manifest goes last because it is the only thing that makes the image referenceable. Writing it first would publish a tag pointing at blobs that do not exist yet, so a concurrent pull would fail; registries also reject manifests referencing absent blobs.

code

bash · 21 lines
bash
# 1. Does the registry already hold this layer?
curl -sI -H "Authorization: Bearer $TOKEN" \
  https://registry.example.com/v2/team/app/blobs/sha256:ab12...
# 404 -> must upload ; 200 -> skip

# 2. Open an upload session
curl -si -X POST -H "Authorization: Bearer $TOKEN" \
  https://registry.example.com/v2/team/app/blobs/uploads/
# 202 Accepted ; Location: /v2/team/app/blobs/uploads/<uuid>

# 3. Send the bytes and commit under the digest
curl -s -X PUT -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/octet-stream" \
  --data-binary @layer.tar.gz \
  "https://registry.example.com/v2/team/app/blobs/uploads/<uuid>?digest=sha256:ab12..."

# 4. Only now publish the manifest under the tag
curl -s -X PUT -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/vnd.oci.image.manifest.v1+json" \
  --data-binary @manifest.json \
  https://registry.example.com/v2/team/app/manifests/1.2

go deeper

for a junior

Know that layers and the config upload first and the manifest is written last so a tag never points at missing content.

for a middle

Name the endpoints — HEAD blob, POST uploads/, PUT with digest, PUT manifest — and explain that existing blobs are skipped.

for a senior

Add digest verification on commit, cross-repository mounts, parallel uploads, resumable sessions, and how interrupted pushes leave only garbage-collectable blobs.

for a principal

Discuss the design consequences: content addressing enables dedup and integrity, manifest-last gives readers atomic publication, and per-repository scoping shapes the permission model.

## What an image is, in registry terms Three kinds of object live in a repository: - **Blobs** — opaque content-addressed byte streams: each filesystem layer (a compressed tar) and the image **config** JSON (architecture, env, entrypoint, layer `diff_id` list, history). Every blob is named `sha256:<hex>` computed over its exact bytes. - **Manifest** — a small JSON document listing the config blob and the ordered layer blobs by digest and size, with a media type. - **Tag** — a mutable pointer in the repository's namespace that resolves to a manifest. For multi-platform images there is an extra level: an **index** (OCI image index / Docker manifest list) whose entries are per-platform manifests. Pushing one means pushing each platform manifest and its blobs, then the index. ## The wire sequence **Step 0 — authorization.** The first API call gets a `401` with a bearer challenge; the client obtains a token for scope `repository:<name>:pull,push`. Push needs both actions because the existence checks are reads. **Step 1 — existence checks.** For every blob the image references: ``` HEAD /v2/team/app/blobs/sha256:ab12... ``` `200 OK` (with `Docker-Content-Digest` and `Content-Length`) means present — skip. `404` means upload. This single step is why a rebuild that changes only the top layer pushes almost nothing: all base layers already exist in the repository. **Step 2 — upload sessions.** For each missing blob: ``` POST /v2/team/app/blobs/uploads/ → 202 Accepted, Location: /v2/team/app/blobs/uploads/<uuid> ``` Then either: - **Monolithic**: `PUT <location>?digest=sha256:ab12...` with the whole body, or - **Chunked**: repeated `PATCH <location>` with `Content-Range`, finished by `PUT <location>?digest=sha256:ab12...` with an empty body. The registry hashes what it received and compares with the claimed digest. A mismatch is rejected — a client cannot store content under a digest that does not describe it, which is the property the whole trust model rests on. Success is `201 Created` with the canonical blob location. Uploads for different blobs run in parallel (a handful at a time by default), and an interrupted session can be resumed or simply retried. **Step 3 — cross-repository blob mount.** When the same registry already stores a blob in another repository the client may read: ``` POST /v2/team/app/blobs/uploads/?mount=sha256:ab12...&from=team/base ``` `201 Created` means the blob was linked into this repository with zero bytes on the wire. If the registry declines (no permission on the source, or unsupported), it returns `202` with a normal upload session and the client falls back to uploading. This is what makes copying or re-tagging an image between repositories on one registry nearly free. **Step 4 — the manifest.** ``` PUT /v2/team/app/manifests/1.2 Content-Type: application/vnd.oci.image.manifest.v1+json ``` The registry validates that every digest referenced in the manifest exists in this repository, stores the manifest under its own digest, and points the tag at it. The response carries `Docker-Content-Digest` — the image's digest, which is how you get the `name@sha256:...` reference that pins content. For a multi-platform push, each platform manifest is PUT first (usually by digest), then the index referencing them. ## Why order matters The manifest is the only object that makes an image *usable*: a tag pointing at a manifest is what a pull resolves. If it were written first, there would be a window in which the tag existed but its blobs did not, and any concurrent pull would fail with missing blobs. Blob-first, manifest-last makes publication effectively atomic from a reader's point of view: before the PUT the tag has its old value (or none), after it the whole image is complete. Registries enforce this by rejecting manifests whose referenced blobs are absent. The same reasoning explains why an interrupted push leaves no broken image — just orphaned blobs, which the registry's garbage collection later reclaims because nothing references them. ## Things that surprise people - **Existence checks are per repository, not per registry.** The same layer stored under `team/base` is not visible to `team/app`'s HEAD check, even on one server, hence the mount API. - **Push is a write of existing bytes, not a build.** The layers were already produced locally; push never re-compresses in the classic flow, so digests stay stable. - **`403` at the first upload with a successful login** means the token lacked the `push` action for the scope — authentication succeeded, authorization did not. - **Blob downloads on the pull side frequently redirect** to object storage with a signed URL; the push side generally goes to the registry directly. - **The pull is the mirror image**: GET the manifest (with `Accept` headers listing acceptable media types), pick the platform entry if it is an index, GET the config, then GET each missing layer in parallel and unpack it.

  • What happens if a push is interrupted after some layers have uploaded?
    The uploaded blobs stay in the registry but nothing references them, because the manifest was never written. The tag keeps its previous value, so no reader ever sees a half-published image. Re-running the push skips the blobs already present thanks to the HEAD checks, and the orphaned blobs are reclaimed by the registry's garbage collection.
  • Why does the registry re-hash a blob you upload rather than trusting the digest you supply?
    Because the digest is the blob's name and the basis of every integrity guarantee downstream — clients verify pulled layers against the digests in the manifest. If a client could store arbitrary bytes under a chosen digest, it could poison content that other images and other repositories reference. Verifying on commit means a stored blob always hashes to its own name.

saying these in an interview costs you the question

  • Saying the manifest is uploaded first so the registry knows what to expect
  • Believing the registry compresses or rebuilds layers during push
  • Assuming a blob present anywhere on the registry is automatically visible to every repository
  • Thinking a failed push leaves a broken, pullable tag
  • Confusing the manifest (a small JSON document) with the image content itself

context

open as a page

What is the difference between `docker save` / `docker load` and `docker push` / `docker pull`, and when would you reach for the save/load pair?

level: juniorimportance: should knowfreq 44%

basics

~20 s

Push and pull transfer an image to and from a registry over HTTP, uploading only layers the registry lacks. Save writes the whole image — layers, config, tags — into a tar file that load reads back. Use save/load for air-gapped transfer or when no registry is available.

open as a page

After a successful `docker login`, where does the Docker CLI keep the credentials it will send on later pushes and pulls, and what are the security implications of the default behaviour?

level: middleimportance: should knowfreq 40%

basics

~20 s

By default it writes an entry to ~/.docker/config.json under auths, keyed by registry host, holding base64(username:password). That is encoding, not encryption, so anyone who reads the file has the credential. docker logout <host> removes the entry.

open as a page

Pushing a rebuilt version of an 800 MB image often transfers only a few megabytes, and copying that image to a second repository on the same registry can transfer nothing at all. What mechanisms make that possible?

level: middleimportance: should knowfreq 40%

basics

~20 s

Layers are content-addressed by SHA-256 digest, so before uploading the client asks the registry whether each digest already exists and skips those it has. Within one registry, a cross-repository blob mount links an existing blob into another repository, transferring zero bytes.

open as a page

How would you query a registry's HTTP API v2 directly with curl to see which manifest a tag currently resolves to and whether a given layer blob exists, and when is doing that worth the trouble?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Get a bearer token, then GET /v2/<repo>/manifests/<tag> with an Accept header listing acceptable manifest media types — the response body plus the Docker-Content-Digest header give you the manifest and digest. HEAD /v2/<repo>/blobs/<digest> reports whether a blob exists.

open as a page