skip to content

A telemetry ingester pulls fine on build workstations but fails on fleet machines of another processor architecture — why?

level: middleimportance: must knowfreq 58%

answer

  1. nothing downloaded before it failed
  2. every machine of one architecture, none of the other
  3. selection, not authentication
  4. the tag promised coverage it never had
  5. a missing entry is a missing build

basics

~20 s

The reference almost certainly resolves to a single-platform artifact. It matches the workstations it was built on, and the fleet machines find no manifest for their platform, so the pull is rejected before any layer transfers. The fix is coverage, published under the same reference.

solid answer

~40 s

The tag everyone used points at either a bare manifest for one platform or an index with only one entry. On the workstations the platform matched and the pull succeeded, which is why nobody noticed. On the fleet machines the host agent walks the entries, finds no architecture-and-operating-system pair matching its own, and refuses with a no-matching-platform error — a resolution failure, not a network or credential failure. Diagnose it by resolving the reference and listing what platforms it actually covers; do not retry the pull. The remedy is for the release to publish a manifest per platform and an index over them under the same reference. Producing those per-architecture builds is the build pipeline's job, not the puller's.

code

pseudocode · 16 lines
pseudocode
document = registry.resolve(reference)

if document is a manifest:
    selected = document            # one platform, take it or leave it
else:
    selected = null
    for each entry in document.entries:
        if entry.architecture == host.architecture
           and entry.os == host.os:
            selected = registry.fetch(entry.manifestDigest)
            break

if selected == null:
    fail "no entry matches this host's platform"   # no blob has moved yet

fetch_missing_layers(selected.layers)

go deeper

for a junior

Recall the symptom and its meaning: a pull that fails with no matching platform means the artifact has no image for that machine's architecture. Nothing was downloaded, and retrying will not help.

for a middle

Explain the selection step that failed, and distinguish the two shapes — a bare single-platform manifest versus an index missing an entry — because they point at different fixes.

for a senior

Diagnose without guessing: resolve the reference, list the covered platforms, decide whether the gap is a missing build or a missing publish, and move the coverage check into the release so the next architecture change fails at publish time.

for a principal

The call you own is the platform matrix: which architectures the estate is allowed to run, what each extra one costs in build time, storage and test surface, and who is accountable when the fleet's hardware moves ahead of the release pipeline.

## What the two machines actually did differently Both machines ran the same reference. Both got the same document back. The difference is entirely in the **selection step** that follows. The workstations are of one processor architecture; the fleet machines are of another. When a host resolves a reference it receives either a manifest (one platform's image) or an index (a list of per-platform entries). Then it compares platforms against its own and picks. The workstation matched; the fleet machine had nothing to match. That is why this failure has a characteristic shape: - it happens **before any bytes move** — no layers transfer, no disk fills, no progress bars; - it is **identical on every machine of the affected architecture** and absent on every machine of the other, which rules out a flaky host; - it says something about *platform*, not about authentication, not about a missing repository; - **retrying never helps**, because nothing about the outcome is timing-dependent. ## The two shapes the mistake takes 1. **A bare single-platform manifest under the tag.** Nobody published an index at all. The artifact is simply one platform's image, and the tag looks exactly like any other tag. This is the common case when the release was first cut on developer machines. 2. **An index with one entry.** Multi-platform publishing was set up but only one architecture's build actually reached the registry — the second build failed, was skipped for time, or was never added when the fleet changed hardware. The index is now a list with a hole in it, and it is more misleading than the bare manifest because its existence implies coverage. There is a third, nastier variant in which the pull *succeeds* on the fleet machine and the process then dies immediately, because the machine cannot execute the instructions the binaries contain. That happens when platform selection is bypassed or when an entry's platform claim does not match its manifest's contents; it is a different diagnosis, and the giveaway is that layers transferred. ## Working it 1. Resolve the reference and read the document it returns. Is it a manifest or an index? 2. If it is an index, list its entries and their platform pairs. Compare that list against the architectures the fleet actually runs — including the ones it recently started running. 3. If it is a bare manifest, read the platform declared in its configuration blob. That is the one platform the artifact supports. 4. Decide whether the gap is a missing build or a missing publish. A build that never ran and a build that ran and was not added to the index look identical from the puller's side and are fixed in different places. 5. Re-publish the reference over an index covering every platform the fleet runs, and make the coverage check part of the release rather than of the incident. ## Why it survived so long unnoticed | assumption people made | what is actually true | |---|---| | "the tag covers our fleet" | a tag is a movable pointer and records nothing about platform coverage | | "it pulled here, so it is fine" | it pulled on a machine whose platform happened to match | | "the registry would have told us" | a registry stores what it is given; it does not require a platform matrix | | "one build produces one image" | one build produces one platform's manifest; coverage needs one per platform | Heterogeneous fleets make this failure common in a specific way: the day a team's workstations stop sharing an architecture with the machines they deploy to, every single-platform artifact that has worked for a year becomes a failed pull, all at once, with no change to the artifact. ## Preventing the repeat - **Make coverage explicit in the release check**: resolve the published reference and assert an entry exists for each platform the fleet runs. It is one small request and it fails loudly at publish time rather than at rollout time. - **Treat a one-entry index as a defect**, not as multi-platform support. If the pipeline can produce only one architecture today, say so in the release notes rather than implying breadth. - **Keep the gap where it belongs.** When a new architecture arrives in the fleet, the work is upstream — a build per architecture and a republished index — not a change to how machines pull. - **Do not paper over it per host.** Forcing a host to accept a non-matching platform gets you a pull that succeeds and a process that cannot start, which is strictly worse than a failure that names the real cause.

  • How do you tell this apart from a credential or a network failure at a glance?
    By where it stops and how it spreads. A platform mismatch fails at resolution with nothing transferred, is perfectly reproducible, and hits every machine of one architecture while sparing the other. Credential failures hit every machine regardless of architecture; transfer failures happen mid-download and often succeed on retry.
  • Could the same tag have worked on the fleet last month and fail today?
    Yes, in two ways. A tag is a movable pointer, so a later publish may have replaced a multi-platform index with a single-platform manifest. Or the fleet gained machines of a new architecture, and the artifact never covered that one — unchanged artifact, changed fleet.
  • Is forcing the host to pull the non-matching manifest ever a reasonable workaround?
    Only to confirm a diagnosis, never as a fix. You trade a clear resolution failure for a container that starts and dies because the machine cannot execute those binaries, and the real gap — no build for that platform — is now hidden behind a runtime crash.

saying these in an interview costs you the question

  • Blames the registry or the network and retries the pull.
  • Assumes a tag implies coverage for every architecture in the fleet.
  • Thinks the host can translate an image built for another architecture.
  • Treats a one-entry index as multi-platform support.
  • Fixes it per host instead of republishing coverage under the reference.