skip to content

questions

5

What does an image manifest list, and what is the configuration blob it names?

level: middleimportance: must knowfreq 62%

answer

  1. a list, not the thing listed
  2. ordered descriptors plus one more
  3. each descriptor is digest and size
  4. one blob is read, not unpacked
  5. defaults and platform live one level down

basics

~20 s

A manifest is a small document listing an image's layer blobs in order, each by digest and size, plus a descriptor for one configuration blob. The configuration blob is the document holding the image's runtime defaults and the identities of its unpacked layers.

solid answer

~40 s

A manifest holds no image bytes at all — it is a list of references. It names an ordered set of **layer descriptors**, each a digest plus a size, and one **configuration blob**, also by digest. The layer order is the stacking order the runtime uses when it merges them into one view. The configuration blob is fetched like any other blob, but it is read rather than unpacked: it declares the platform the image targets, the defaults for the process that will run, and the list of unpacked-layer identities that ties this configuration to those exact layers. The manifest's own digest is the identity of that image for that platform, and it changes if any descriptor inside it changes.

code

json · 11 lines
json
{
  "configuration": {
    "digest": "hash-of-the-configuration-blob",
    "size": 1472
  },
  "layers": [
    { "digest": "hash-of-the-base-layer", "size": 31457280 },
    { "digest": "hash-of-the-dependency-layer", "size": 8912896 },
    { "digest": "hash-of-the-application-layer", "size": 1048576 }
  ]
}

go deeper

for a junior

Recall that a manifest lists layers by digest and names one configuration blob, and that it contains no file contents itself. Knowing the two kinds of thing it names is enough at this stage.

for a middle

Explain why the layer order is load-bearing, why the sizes are transfer sizes rather than on-disk sizes, and what the configuration blob declares beyond defaults.

for a senior

Use the structure diagnostically: inspect a manifest to answer size, coverage and shared-base questions before downloading anything, and reason about what a changed digest does and does not force a fleet to re-fetch.

for a principal

The leverage here is standardising what publishers must put in the configuration blob — a declared non-root user, a declared platform, honest labels — so consumers across teams can gate on it mechanically.

## A document of references, not of content It is worth stating the surprising part first: **a manifest contains none of the image**. It is a few kilobytes of structured text whose entire job is to name things that live elsewhere in the registry as content-addressed blobs. Two kinds of thing are named. - **Layer descriptors**, in order. Each is a digest identifying one layer blob, plus the size of that blob so a client can plan the transfer. - **One configuration descriptor**, naming a single blob that is not a layer at all. The manifest's own digest is computed over the manifest's bytes. Because those bytes contain the digests of everything the image is made of, that one digest transitively names the whole image for one platform. Change a layer and its digest changes, so the manifest's bytes change, so the manifest's digest changes. That is what people mean when they say an image's identity is content-derived. ## The layer list is ordered on purpose The order is not decorative. A runtime unpacks the layers bottom-up and merges them into one view, and where two layers hold the same path the upper one wins. So the manifest's order is what decides the filesystem the process will see, and reordering the list produces a different image even from identical blobs. The descriptors carry the **compressed transfer size** — what will cross the network — not what the layers will occupy once expanded on disk. A host reading a manifest therefore knows what it is about to download, but not, from the manifest alone, what it will cost to store unpacked. ## The configuration blob The configuration blob is fetched by digest like a layer, but it is a document to read rather than an archive to unpack. It carries three kinds of statement: 1. **What platform this image is for** — the processor architecture and operating system the binaries inside were built to run on. This is the image's own claim about itself, made independently of anything an index entry says. 2. **Runtime defaults** — the command and arguments to run, the user to run as, the working directory, default environment entries, and declared ports or volume paths as metadata for whoever starts the workload. 3. **The identities of the unpacked layers**, in the same order as the manifest's list. These are the identities of the layers *after* expansion, which is how a runtime can tell that a layer it already has unpacked is the one this configuration expects. ## Manifest against configuration blob | | manifest | configuration blob | |---|---|---| | what it is for | telling a client what to fetch | telling a runtime what to run and on what | | names blobs | yes, layers and the config | no, it is itself a blob | | read or unpacked | read | read | | carries defaults | no | yes | | declares a platform | no | yes | The split matters. The fetching half of a pull only needs the manifest; the starting half needs the configuration. A registry can serve the manifest to a client that will decide it already has everything and never fetch a single layer. ## What follows from the structure - **Shared layers are free.** Two images built on the same base name the same layer digests, so a host that already holds those blobs fetches nothing for them however many manifests point at them. - **A manifest is cheap to inspect.** Reading what an image is made of — how many layers, how large, which configuration — costs one small request, which is why coverage and size questions are answerable without downloading anything. - **Nothing in the manifest is prose.** There is no field that says what the image is *for*; that lives in whatever labels the publisher put in the configuration blob, and a consumer should not assume any of it is present. - **Deleting a file does not shrink the list.** A later layer that removes a path still ships as a layer, and the earlier layer holding the bytes is still named by the manifest — which is why a manifest's summed sizes are a truer measure of an image than any claim about what was cleaned up. - **Two platforms means two manifests.** One manifest is one platform's image. Covering more platforms means more manifests, gathered by an index above them.

  • If I change only the last layer of an image, what else changes?
    That layer's digest changes, so the manifest's layer list changes, so the manifest's bytes and therefore its digest change. The earlier layers keep their digests and are not re-uploaded or re-downloaded. The configuration blob also changes, because it records the identities of the unpacked layers.
  • Does the manifest tell you how much disk the image will use on the host?
    No. Its descriptors carry compressed transfer sizes, which is what will cross the network. Expanded on-disk cost is larger and depends on how the host stores unpacked layers, and layers already present for another image cost nothing extra.
  • Why is the configuration a separate blob instead of fields inside the manifest?
    Because then it is content-addressed like everything else: the manifest names it by digest, so the defaults are pinned by the same mechanism as the layers, can be cached and shared, and cannot be altered without changing the manifest that names them.

saying these in an interview costs you the question

  • Thinks the manifest contains the layer contents rather than references to them.
  • Says the layer order in the manifest does not matter.
  • Believes the sizes in the manifest are expanded on-disk sizes.
  • Treats the configuration blob as a layer to be unpacked into the filesystem.
  • Assumes one manifest can cover several processor architectures.
open as a page

A telemetry ingester pulls fine on build workstations but fails on fleet machines of another processor architecture — why?

level: middleimportance: must knowfreq 58%

basics

~20 s

The reference almost certainly resolves to a single-platform artifact. It matches the workstations it was built on, and the fleet machines find no manifest for their platform, so the pull is rejected before any layer transfers. The fix is coverage, published under the same reference.

open as a page

What does an image index contain, and how does a host choose which manifest inside it to pull?

level: juniorimportance: should knowfreq 48%

basics

~20 s

An image index is a small document listing one entry per platform, where a platform is a processor-architecture and operating-system pair. Each entry names a manifest by digest. The pulling host matches its own platform against those entries and fetches only that manifest.

open as a page

What runtime defaults does an image's configuration blob carry, and what happens when a workload spec sets the same fields?

level: middleimportance: should knowfreq 45%

basics

~20 s

The configuration blob declares defaults for the process: the command and its arguments, the user, the working directory, environment entries, and declared ports as metadata. A workload spec that sets the same field wins at start-up, and the image is unchanged — same digest, same defaults for the next caller.

open as a page

An index entry claims the fleet's platform, the pull succeeds, yet the process dies instantly — what happened?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

The platform fields on an index entry are a publisher's claim, and nothing verifies them against the blobs. Selection matched a false claim, so the layers that arrived hold binaries the machine cannot execute. The image's own configuration blob declares the truth.

open as a page