skip to content

Signature verification passes in one region and fails in two others behind regional caches — same digest. Why?

level: seniorimportance: must knowfreq 40%

answer

  1. identical digest rules out the artifact
  2. one lookup moved, the other did not
  3. an intermediary holds what was asked for
  4. an empty answer is not an error
  5. promote the graph, not the manifest

basics

~20 s

Because the digest travelled and the trust material did not. Signatures are separate objects discovered by a second lookup against the same repository; an intermediary that serves the image a client asked for has no reason to hold the referrers nobody requested.

solid answer

~60 s

Start from the fact that the digest is identical in all three regions: the bytes are provably the same, so this is not an image problem, it is a discoverability problem. A signature is a distinct manifest found by a separate referrers lookup scoped to `repository + digest`, and an intermediary populates itself from what clients ask for. A pull of the image manifest and its layers never touches the referrers endpoint, so the trust material is simply not there when the admission check asks — and depending on the intermediary, the referrers query may be answered locally as empty rather than proxied upstream. Confirm it by querying referrers directly against each endpoint for the same digest and comparing. The durable fixes are to make promotion referrer-aware so trust material is replicated as a first-class part of the artifact graph, or to resolve verification against a source you control rather than whatever served the bytes. Failing open in the two broken regions is the one option that turns a transport bug into a security hole.

go deeper

for a junior

Remember the one fact that anchors this: pulling an image does not fetch its signature, because the signature is a separate object found by a separate request.

for a middle

Explain why an intermediary can return a well-formed empty referrers answer, and how to tell that apart from an unimplemented endpoint returning 404.

for a senior

Show the diagnosis in order — compare digests first, then compare referrers per endpoint — and land on making promotion carry the whole artifact graph rather than exempting the broken regions.

for a principal

Own the standing rule: every hop that moves artifacts is a place verifiability can be lost, so the estate needs one referrer-aware promotion path and a policy on where verification resolves from, not per-incident exceptions.

## Read the symptom precisely The most useful sentence in the report is "same digest". A digest is a hash over the manifest bytes, so if all three regions resolve to the same digest, every region is about to run identical content. Nothing about the artifact differs. What differs is whether the thing that vouches for the artifact can be *found* from where the check is running. That reframing matters because the instinctive diagnosis — "the image in those regions is different" or "the policy is misconfigured" — sends people to the wrong half of the system. The policy is the same policy; it is being handed a different answer to the same question. ## Why an intermediary loses trust material Trust material is not part of the image. It is a separate manifest in the same repository carrying a `subject` field that names the image digest, and it is found by a second request — a referrers lookup keyed on that digest, or on older registries a derived `sha256-<hex>` tag. Two properties combine badly: **Discovery is repository-scoped, not global.** The digest is a universal name for the bytes, but "what refers to this digest" is only ever answered by a specific registry about its own contents. Ask three registries and you can legitimately get three different answers. **An intermediary holds what was requested of it.** A client pulling an image asks for a manifest and some layers. It does not ask for referrers, so an on-demand intermediary has no reason to have fetched them. Worse, when the admission check *does* ask for referrers, an intermediary that serves the request itself — from an index built only over what it happens to hold — can return a perfectly well-formed empty answer. There is no error anywhere. Every component reports success. The verification simply concludes there is nothing to verify. The region that works is the one whose verification path reaches a registry where the signatures were actually pushed. ## How to confirm it in five minutes 1. Resolve the tag to a digest at each of the three endpoints and confirm they match. If they do not, you have a different and simpler problem. 2. Query the referrers of that digest against each endpoint directly, and against the central registry. Expect a populated index centrally and an empty one in the failing regions. 3. If any endpoint answers 404 rather than an empty index, check the derived `sha256-<hex>` tag too — 404 means the endpoint is not implemented, not that there are no referrers, and a client that conflates the two will report "unsigned" incorrectly. 4. Check whether the failing regions can reach the central registry for the referrers request at all. Sometimes the fix is nothing more than pointing the verification lookup at a resolver that is allowed to go upstream. ## Fixing it properly **Make the artifact graph the unit of promotion.** Whatever copies or replicates artifacts between registries has to walk the referrers of each subject and carry them too, including the derived fallback tag where the destination needs it. "Copy the image" is the wrong operation; "copy the image and everything that refers to it" is the right one. This is the fix that keeps working when a fourth region appears. **Or decouple verification from the serving path.** Verify against a registry you control, or resolve the trust material once at admission time from a source of record, and let the intermediary serve only bytes. This has the pleasant property that the intermediary can be as lossy as it likes — it is no longer in the trust path — but it costs you an availability dependency on that source of record, which is exactly what regional caches existed to remove. **Do not fail open.** The tempting mitigation is an exception for the two regions until replication is fixed. That converts a transport bug into a standing hole: any artifact reaching those regions now runs unverified, and exceptions written during an incident outlive the incident. If you must ship before the fix, prefer narrowing — verify against the central registry from those regions, or freeze them on already-verified digests — over disabling the check. ## The adjacent trap worth naming Referring manifests are usually untagged. A registry policy that garbage-collects untagged manifests can reap signatures while leaving every image intact, producing exactly the same symptom on a single registry with no mirror involved: an artifact that verified last month and does not today, with no change to the artifact. When you audit the distribution path for places trust material is dropped, retention policy belongs on the list alongside copies, exports and caches.

  • The same team also exports images to tarballs for an air-gapped site and verification fails there too. Same cause?
    Same class, different transport. An image export carries the manifest, config and layers — the things that make up the image — and nothing that merely refers to it, so the signatures and attestations never board the transfer. The fix is to move an artifact bundle that includes the referring manifests, not an image, and to verify on the far side before the artifact is loaded into the internal registry. Otherwise the inside holds artifacts nothing can prove, which is precisely the audit position an air-gapped regulated environment cannot defend.
  • Verification suddenly fails for an image on a single registry with no mirrors or copies involved. What do you check?
    Whether the trust material was garbage-collected. Referring manifests are typically untagged, so a retention policy that reaps untagged manifests will delete signatures and attestations while leaving the image untouched — the artifact is unchanged, and only its discoverability disappeared. Check the retention rules, and confirm by looking for the referring manifests by digest rather than trusting the referrers index.
  • Why not just let the two failing regions skip verification until replication is fixed?
    Because it removes the control from the exact paths you have just proven are lossy, and an exception created under incident pressure tends to become permanent. Narrower options exist: have those regions perform the referrers lookup against the central registry while continuing to pull bytes locally, or pin them to digests already verified elsewhere. Both keep the check in force while the replication gap is closed.

saying these in an interview costs you the question

  • Assumes the image itself differs between regions despite the same digest
  • Treats an empty referrers response as a policy misconfiguration
  • Proposes disabling verification in the failing regions as the fix
  • Believes copying an image also copies what refers to it
  • Forgets that an unimplemented referrers endpoint answers 404, not empty

context