skip to content

Walk through what a registry stores for a single-platform container image: which fields in the manifest identify the layers, and why does the config blob's rootfs.diff_ids list different sha256 values than the manifest's layer descriptors?

level: seniorimportance: should knowfreq 34%

answer

  1. manifest = config descriptor + layer descriptors
  2. descriptor = mediaType, digest, size
  3. layer digest = compressed; diff_id = uncompressed
  4. chainID = rolling hash of diff_ids
  5. history[] has empty_layer entries; <missing> on pull

basics

~20 s

The manifest lists a config descriptor and layer descriptors — digests of the compressed blobs as transferred. The config blob lists rootfs.diff_ids — digests of the same layers uncompressed. Different bytes hashed, so different values; diff_ids identify layers on disk, layer digests identify them on the wire.

solid answer

~50 s

Three documents. **Manifest**: `mediaType`, a `config` descriptor, and a `layers` array. Each descriptor is `{mediaType, digest, size}` — the digest of the blob *as stored and transferred*, i.e. the gzip-compressed tar. **Config blob**: runtime metadata (`architecture`, `os`, and a `config` object with Env, Cmd, Entrypoint, User, WorkingDir, ExposedPorts, Labels), plus `rootfs.diff_ids` and a `history` array. `diff_ids` are digests of the *uncompressed* layer tars. They differ because compression is not canonical: the same filesystem content gzipped twice can produce different bytes, so a compressed digest is a transport identity while the diff_id is the stable content identity. Local storage chains diff_ids into a chainID naming each layer on disk; that is why re-pulling an image whose layers you already have can skip work even when compressed digests changed. The image ID is `sha256(config blob)`, and since the config contains all diff_ids, it covers the whole filesystem.

code

bash · 4 lines
bash
docker buildx imagetools inspect --raw alpine:3.20
docker image inspect --format '{{json .RootFS.Layers}}' alpine:3.20
docker image inspect --format '{{json .Config}}' alpine:3.20
docker history --no-trunc alpine:3.20

go deeper

for a junior

Know that a manifest points at a config and a list of layers, and that docker inspect and docker history read that metadata.

for a middle

Explain descriptors (mediaType/digest/size) and that diff_ids are the uncompressed layer hashes stored in the config.

for a senior

Explain why the two digests exist — non-canonical compression, transport identity versus content identity — and connect diff_ids to chainIDs and local layer reuse.

for a principal

Reason about the consequences for registry mirroring, deduplication, signature and provenance anchoring, and what your platform standardises on when copying images between environments.

## The three documents For a single-platform image the registry holds: 1. **A manifest** — the entry point. Roughly: ``` { schemaVersion, mediaType, config: { mediaType, digest: sha256:<config>, size }, layers: [ { mediaType, digest: sha256:<blob>, size }, ... ], annotations: { ... } } ``` Every element is a **descriptor**: media type, digest, size. Media types matter — `application/vnd.oci.image.layer.v1.tar+gzip` versus `+zstd` versus uncompressed — because they tell the client how to decode the blob. 2. **A config blob** — a JSON document referenced by the manifest's `config` descriptor: ``` { architecture: amd64, os: linux, config: { Env: [...], Cmd: [...], Entrypoint: [...], User, WorkingDir, ExposedPorts, Labels, Volumes }, rootfs: { type: layers, diff_ids: [ sha256:..., ... ] }, history: [ { created_by: "RUN apt-get update", empty_layer: true|false }, ... ] } ``` 3. **The layer blobs** themselves, each a tar of filesystem changes, normally gzip-compressed. ## Two digests per layer, on purpose A layer therefore has two identities: - **Layer digest** (in the manifest): sha256 of the *compressed* blob. This is what the registry stores it under and what the client verifies on download. It is a transport-level name. - **diff_id** (in the config): sha256 of the *uncompressed* tar. This is the content identity of the filesystem change itself. Why both? Because compression is not deterministic across implementations or settings. Recompressing an identical tar with a different gzip level or library yields different bytes and hence a different layer digest, while the diff_id is unchanged. If images were identified only by compressed digests, any re-compression would look like brand-new content. Splitting the identities lets a client that already holds a layer (by diff_id) recognise it, and lets registries deduplicate blobs by compressed digest independently. The local daemon goes one step further and computes a **chainID** — a rolling hash folding each diff_id together with the chain so far — because a layer's on-disk contents depend on every layer below it, not just on its own tar. That chainID names the directory under the storage driver. ## Why the image ID is the config's digest Since the config embeds every diff_id in order, plus all runtime metadata, hashing the config yields an identifier that changes if any layer content, any environment variable, or the entrypoint changes. That is exactly the property you want from an image ID, and it is why `sha256(config)` was chosen rather than hashing layers directly. Note what the image ID does *not* cover: the manifest's media types, sizes and annotations. Those live one level up, in the manifest, whose digest is the repo digest. ## The history array `history` records one entry per build step, with `created_by` (the instruction text) and `empty_layer: true` for steps such as `ENV` or `LABEL` that changed only metadata. There are therefore usually more history entries than layers, and the mapping between them is positional over the non-empty entries. `docker history` renders exactly this array, which is why layers of a pulled image show `<missing>` as their ID: the intermediate configs were never transferred, only the history text was. ## Inspecting it in practice - `docker buildx imagetools inspect --raw repo:tag` — the manifest (or index) exactly as stored. - `docker image inspect repo:tag` — `Id` (config digest), `RootFS.Layers` (diff_ids), `Config` (Env/Cmd/Entrypoint/User), `RepoDigests`. - `docker history --no-trunc repo:tag` — per-step provenance, including steps that produced no layer. This is how you audit an image without running it: read the entrypoint and user it will start as, confirm no secret was baked into an `ENV`, and see which instruction produced each layer. ## Why a senior cares Understanding the split explains several real behaviours: why the same image can appear under two repo digests, why a `docker save`/`docker load` round trip changes registry-level digests but not the image ID, why cross-registry copy tools can preserve digests only if they copy blobs byte-for-byte rather than re-compressing, and why deduplication decisions differ between registry storage and node storage.

  • Why does `docker history` show `<missing>` for the layer IDs of a pulled image?
    Image IDs exist only for complete images, i.e. for configs. When you build locally the daemon keeps the intermediate configs, so each step has an ID. A pull transfers only the final config plus the layer blobs, so the intermediate configs simply do not exist on that host — the history text survives, the IDs do not.
  • If a mirroring tool copies an image between registries, can the digest stay the same?
    Yes, provided it copies the blobs and the manifest byte-for-byte. The digest is over exact bytes, so any re-compression of layers or re-serialisation of the manifest changes it. Purpose-built copy tools preserve bytes precisely so that pinned digests keep resolving across registries.

saying these in an interview costs you the question

  • Claiming the manifest's layer digests and the config's diff_ids should be equal and a mismatch is corruption
  • Saying the image ID hashes the layers directly rather than the config that lists them
  • Assuming there is one history entry per layer, ignoring metadata-only steps flagged empty_layer
  • Believing `docker save`/`load` preserves registry digests

context