skip to content

An index entry claims the fleet's platform, the pull succeeds, yet the process dies instantly — what happened?

level: seniorimportance: nice to knowfreq 30%

answer

  1. everything succeeded except the process
  2. layers moved, so selection matched
  3. the entry is a claim, not a proof
  4. the image declares its own platform
  5. compare entry against configuration blob

basics

~20 s

The platform fields on an index entry are a publisher's claim, and nothing verifies them against the blobs. Selection matched a false claim, so the layers that arrived hold binaries the machine cannot execute. The image's own configuration blob declares the truth.

solid answer

~40 s

Selection is a metadata comparison. A host reads the platform fields on each entry, matches its own, and fetches that manifest — it never inspects the layer contents to confirm the architecture. So an entry whose platform fields are wrong produces a pull that succeeds completely and a process that cannot start. The two documents to compare are the **entry** in the index and the **configuration blob** of the manifest it points at, because the configuration is the image's own declaration of what it was built for. When they disagree, the publisher is at fault: an index assembled by hand, a mirroring or promotion step that rewrote entries, or a build that produced one architecture's binaries under another's label. Fix it at publish time and add a check that the two agree.

go deeper

for a junior

Recall the distinction that matters here: a failed pull and a failed start are different problems. If the layers downloaded, the platform was matched, whether or not the match was truthful.

for a middle

Explain that selection compares metadata only, and name the two documents that state the platform — the index entry and the image's configuration blob — and which one is the image's own claim.

for a senior

Drive the diagnosis from evidence: confirm bytes moved, read the selected entry, read the configuration blob it leads to, spot duplicate manifest digests, and put the agreement check into the publish step.

for a principal

What you own is where trust is established in the pipeline: which step is allowed to assemble an index, what it must verify before a release is tagged, and how that evidence is recorded for the teams consuming the reference.

## Two places the platform is written down The same fact — what architecture and operating system this image is for — appears in two documents, written by two different steps, and nothing forces them to agree. - **The index entry.** A descriptor with a manifest digest, a size, and platform fields. Whoever assembled the index wrote those fields. - **The configuration blob.** Named by the manifest the entry points at, it declares the platform the image itself targets, alongside the runtime defaults. Selection reads only the first. The digest in the entry guarantees that the bytes you get are the bytes the publisher pointed at — it does not guarantee that the publisher pointed at the right ones, and it says nothing about whether the platform fields describe those bytes truthfully. ## Why this failure looks so unlike the ordinary one | | no matching platform | mismatched platform claim | |---|---|---| | where it stops | at selection | at process start | | bytes transferred | none | all of them | | what the host reports | no entry for this platform | the process exited immediately | | cause | a build or an entry is missing | an entry's claim is false | | where it is fixed | the build and publish steps | the publish step, plus a check | The second one is worse to live with because every layer of the system reports success: the resolution succeeded, the transfer succeeded, the local copy is present and verifiable against its digest. Only the workload fails, and it fails in a way that looks like an application crash. ## How an entry comes to lie 1. **An index assembled by hand or by a script.** Someone gathers manifest digests and types the platform fields next to them. A copy-paste between two nearly identical entries is all it takes. 2. **A promotion or mirroring step that rewrites.** Moving artifacts between registries sometimes re-assembles the index rather than copying it, and a step that fills platform fields from the wrong source of truth produces a plausible, wrong index. 3. **A build that emulated another architecture and mislabelled the result.** The build ran, produced binaries for the machine it ran on, and the publish step labelled the entry with the architecture it was *supposed* to target. 4. **A single-platform artifact wrapped to look broad.** An index is created with entries for several platforms, all pointing at the same manifest digest, on the theory that any image is better than a failed pull. All four are publishing defects. None of them is visible to a puller until something tries to run. ## Diagnosing it - Confirm the transfer actually happened. If layers moved, this is not a selection failure and the no-matching-platform reasoning does not apply. - Resolve the reference and read the entry your host would select: note its platform fields and its manifest digest. - Fetch that manifest and read its configuration blob. Compare the platform it declares against the entry's fields. - Check whether two entries share one manifest digest. That is the wrapped-single-platform case and it is immediately conclusive. - Compare against an entry known to work. If the working entry's claim agrees with its configuration and the failing one does not, you have located the defect precisely. ## Making it not recur - **Verify at publish, not at pull.** Before a release is tagged, walk every entry, fetch each manifest's configuration, and assert the declared platform equals the entry's fields. It is a handful of small requests over documents of a few kilobytes. - **Prefer generated indexes to assembled ones.** An index produced by the same step that produced the manifests has one source of truth for the platform fields; an index typed by a human has two. - **Never fill coverage with duplicates.** An entry that points at another platform's manifest converts a clear pull failure into an opaque crash, and it will be diagnosed by whoever is on call rather than by whoever published it. - **Record the check's result.** Once a release states which platforms were verified rather than claimed, the next incident starts from evidence instead of from the index's own say-so. The general lesson is about trust boundaries in this model: digests make bytes tamper-evident, and metadata *about* those bytes is only as good as the step that wrote it. Selection runs on the metadata.

  • Does verifying the digest catch this?
    No. A digest proves the bytes you received are the bytes named, so it detects corruption and substitution. It cannot detect that the entry naming those bytes described them wrongly, because the platform fields sit beside the digest rather than inside the content it names.
  • What is the giveaway that an index was padded with duplicate entries?
    Two or more entries carrying different platform fields but the same manifest digest. One image cannot be two platforms, so at most one of those claims is true, and the hosts matching the others will pull successfully and fail to start.
  • Why not have the runtime check the binaries before starting them?
    Some runtimes do compare the image's declared platform against the host before starting, which turns this into an earlier, clearer refusal. But they are comparing another declaration, not inspecting the binaries, so a configuration blob that lies in the same direction still gets through.

saying these in an interview costs you the question

  • Thinks a matching digest proves the image targets that platform.
  • Believes the host inspects layer contents to confirm the architecture.
  • Diagnoses it as an application bug because the process crashed.
  • Pads an index with duplicate entries to avoid failed pulls.
  • Fixes it on the affected hosts rather than in the publishing step.