What does an image index contain, and how does a host choose which manifest inside it to pull?
answer
- one reference, two possible documents
- the upper one lists, the lower one describes
- one entry per platform pair
- each entry names a manifest by digest
- no match, no pull
basics
~20 sAn image index is a small document listing one entry per platform, where a platform is a processor-architecture and operating-system pair. Each entry names a manifest by digest. The pulling host matches its own platform against those entries and fetches only that manifest.
solid answer
~40 sA reference resolves to one of two documents. It is either a **manifest**, which describes one image for one platform, or an **index**, which describes no image itself and instead lists descriptors — one per platform — each naming a manifest by its digest, with the architecture and operating system that manifest was built for. When a host gets an index, it compares each entry's platform against its own, picks the matching entry, fetches that manifest by the digest in the entry, and only then reads the configuration blob and the layers the manifest lists. If no entry matches, the pull fails before any layer moves. The index itself carries no layers.
code
json · 14 lines{
"entries": [
{
"manifestDigest": "hash-of-the-workstation-manifest",
"size": 1174,
"platform": { "architecture": "workstation-arch", "os": "server-os" }
},
{
"manifestDigest": "hash-of-the-fleet-manifest",
"size": 1172,
"platform": { "architecture": "fleet-arch", "os": "server-os" }
}
]
}go deeper
Recall the two documents and their order: an index lists manifests, a manifest lists layers and a configuration blob. Be able to say that the host matches its own architecture and operating system against the entries.
Explain the selection as a metadata comparison followed by a fetch by digest, and be able to say what the index does not contain: no layers, no runtime defaults, no guarantee that every platform is covered.
Show that you check coverage rather than assume it: resolve a reference, count entries, and treat a missing platform as a missing build in the release, not as a pull problem to retry.
The trade-off you own is the platform matrix itself: every extra architecture is another build, another set of blobs to store and another surface to test, weighed against the machines your fleet is actually allowed to buy.
## One reference, two possible documents When a host asks a registry for an image by reference — a repository name plus a tag, or a repository plus a digest — the first thing that comes back is a small document, not the image. That document is one of two kinds, and telling them apart is most of this subject. A **manifest** describes exactly one image for exactly one platform. It lists the layer blobs in order and names a configuration blob, each by digest and size. A host holding a manifest knows everything it still has to download. An **index** describes no image at all. It is a list of **descriptors**, one per platform, each naming a *manifest* by digest. It carries no layers of its own and no runtime defaults. Its only job is to let one reference stand for several per-platform images at once, so that a single tag in a single deployment description works on machines that do not share a processor architecture. ## What one entry in an index carries Each entry is deliberately thin: - the **digest** of the manifest it points at — the content-derived name of that manifest's bytes; - the **size** of that manifest, so a client knows what it is about to fetch; - the **platform**: at minimum a processor architecture and an operating system, and often a variant field that distinguishes revisions within one architecture family; - nothing about the image's contents. Layers appear one level down, in the manifest the entry points at. Because an entry names its manifest by digest, the set an index describes is fixed the moment the index bytes are written. You cannot swap one platform's image out from under an index without producing a different index. ## How a host chooses 1. Resolve the reference against the registry and read the document that comes back. 2. If it is a manifest, there is no choice to make: that is the image, for whatever platform it was built for. 3. If it is an index, walk its entries and compare each entry's architecture and operating system against the host's own. 4. On a match, fetch that entry's manifest by the digest recorded in the entry. 5. Then proceed exactly as for a single-platform image: read the configuration blob, and fetch the layer blobs the host does not already have. Step 3 is a metadata comparison and nothing more. The host is trusting what the entry says about itself; it is not inspecting the layer contents to confirm the claim. ## Index against manifest | | index | manifest | |---|---|---| | what it lists | one descriptor per platform | ordered layer descriptors plus a configuration blob | | what it points at | other manifests | content blobs | | how many per reference | at most one | one per platform | | what its digest identifies | the whole multi-platform set | one platform's image exactly | | does a host download it whole | yes, it is tiny | yes, then selectively fetches its blobs | ## Why the indirection is worth having - **One name everywhere.** A deployment description, a release note and a change record all name one reference, and heterogeneous machines each resolve it to something they can run. - **No merged mega-image.** Each platform keeps its own separate layers. A host downloads only the layers of the one manifest it selected, so a fleet of one architecture pays nothing for the existence of the others. - **Exactness survives.** Naming the index by its digest is still an exact name: the index bytes are fixed, so the per-platform manifests it lists are fixed too, while each host still ends up with the build meant for it. - **Publishers can extend it.** Adding support for a new architecture means publishing a new index with one more entry; the old entries keep their digests and the hosts already running them are unaffected. ## Where this trips people up - **Not every reference resolves to an index.** A single-platform artifact is published as a bare manifest, and a host of a different platform gets no choice and no match. - **The tag says nothing about coverage.** Nothing in a tag string records which platforms the index behind it covers, so "it is on the shared tag" is not evidence that your machines are covered. - **An index cannot invent a build.** If the release only ever produced one platform's manifest, wrapping it in an index does not make it runnable elsewhere; the missing entry is a missing build. - **The index is not the config.** Runtime defaults — the command, the user, the working directory — live in the configuration blob named by the *selected manifest*, one level below the entry that was matched.
- If a host pins the index digest rather than a tag, does it still get the right image for its own architecture?Yes. The index digest fixes the index bytes, and those bytes list one manifest digest per platform. Every host resolving that same index still performs its own platform match and ends up on its own entry. What is fixed is the *set*; what varies per host is which member of the set it fetches.
- What does a host do when an index contains entries for platforms it does not run?Nothing. It reads the index — a document of a few kilobytes — compares platforms, and fetches only the manifest and blobs of its own entry. Extra entries cost the reader one comparison each and no transfer, which is why publishing broad coverage is cheap for consumers.
- Can the same manifest digest appear under two entries in one index?It can appear, because an index is just a list of descriptors and nothing stops a publisher repeating one. It is almost always a publishing mistake: it means two platform claims resolve to identical bytes, so one of the two platforms will get binaries that were not built for it.
A catalogue page that lists one edition per region, each with its own order number. You do not receive the page; you read the line for your region and order that number.
saying these in an interview costs you the question
- Thinks a multi-platform image is one image containing every architecture's binaries at once.
- Believes the host downloads every platform's layers and picks one at start-up.
- Says the index lists layer blobs directly.
- Assumes every reference resolves to an index.
- Thinks the tag string encodes which architectures are covered.